lc/gau: pulling known URLs from OTX, the Wayback Machine, Common Crawl and URLScan
Fetch known URLs from AlienVault's Open Threat Exchange, the Wayback Machine, and Common Crawl.
At a glance
- What is it?
- gau is a Go command line tool that collects URLs already recorded for a domain by four public archives. It is a reconnaissance input, not a crawler, and its output quality depends entirely on what those archives happened to capture.
- Who is it for?
- gau fits reconnaissance work where you already have a domain list and want historical URLs without touching the target. It does not fit anyone expecting a live crawl, a verified inventory of what is currently served, or a single source of truth: the README lists four providers and the output is only as complete as their archives.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Activity is slowing. The repository last received commits 6 months ago.
- What is it written in?
- Mainly Go, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What gau collects, and what it deliberately does not
gau answers one question: which URLs have already been seen for this domain by somebody else. The README describes it as fetching known URLs from AlienVault's Open Threat Exchange, the Wayback Machine, Common Crawl and URLScan for a given domain, and credits Tomnomnom's waybackurls as the inspiration. That framing matters. gau does not crawl the target, does not send requests to the host you name, and does not confirm that anything it prints still exists. It queries archive and threat exchange APIs and prints what they return.
The intended user is a security engineer or bug bounty hunter who has a list of domains and wants a starting set of endpoints, parameters and paths to look at. Because the URLs come from third party archives, the collection step is passive with respect to the target. You can build a picture of a host's historical surface without generating traffic against it.
The cost of that design is freshness and completeness. A URL that was never indexed by any of the four providers will not appear, no matter how live it is. A URL that was indexed years ago and has since been removed will appear anyway. gau gives you a candidate list, and the filtering flags exist precisely because that list is noisy.
How the provider queries and filtering pipeline fit together
The module path in go.mod is github.com/lc/gau/v2, and the repository splits into cmd/, pkg/ and runner/ directories. The dependency list shows what the network layer looks like: valyala/fasthttp for the HTTP client, json-iterator/go for JSON decoding, deckarep/golang-set/v2 for deduplication, and lynxsecurity forks of pflag and viper for flags and configuration. That is a small, focused dependency set for a tool whose whole job is issuing parallel HTTP requests and merging the responses.
The --providers flag is the control point for data flow. Its documented values are wayback, commoncrawl, otx and urlscan, so you can restrict a run to a single archive. The --threads flag sets the number of workers, --retries and --timeout govern the HTTP client, and --proxy accepts either a socks5:// or http:// URL. Deduplication across providers is implied by the set dependency rather than documented as a flag.
After collection, gau can filter on what the archives recorded. --fc and --mc take comma separated status codes to filter out or match, and --ft and --mt do the same for mime types. --blacklist drops extensions such as ttf, woff, svg or png. --fp removes different parameters of the same endpoint, which collapses near duplicates that differ only in query string. --from and --to accept a YYYYMM date to bound the archive window, and --subs expands the query to subdomains of the target. Output is line oriented by default, or JSON with --json.
One design detail worth flagging: the status code and mime type filters depend on metadata the providers attach to archived URLs. Where a provider does not record that metadata, the filter has nothing to act on. The README does not state how each provider populates those fields, so treat status and mime filtering as best effort rather than a guarantee.
Installing gau with go install, a release binary or Docker
The README gives four installation routes. The shortest is from source with Go, and it installs the latest tagged module:
go install github.com/lc/gau/v2/cmd/gau@latestAfter that, gau should be on your PATH and gau --version prints the build you got. If you prefer to build from a clone, the README's GitHub route clones the repository, changes into the cmd directory, builds, and moves the binary into /usr/local/bin, ending with gau --version to confirm:
git clone https://github.com/lc/gau.git; \
cd gau/cmd; \
go build; \
sudo mv gau /usr/local/bin/; \
gau --version;Pre-built binaries are available from the releases page. The README shows extracting an archive and moving the binary into place:
tar xvf gau_2.0.6_linux_amd64.tar.gz
mv gau /usr/bin/gauDocker is also supported, both from the published image and from the Dockerfile in the repository root:
docker run --rm sxcurity/gau:latest --help
docker build -t gau .
docker run gau example.comThe README adds a warning about the container: piping a domain into it, as in echo "example.com" | gau, will not work with the docker container. Pass the domain as an argument instead.
For a first real run, start with a single domain and write the results to a file:
gau --o example-urls.txt example.comThe file should contain one URL per line. If it is empty, the domain may simply have no coverage in the selected providers, which is a common outcome for hosts that were never indexed. Feeding a list of domains is the same idea, with a worker count:
cat domains.txt | gau --threads 5There is a shell level trap the README calls out directly. ohmyzsh's git plugin defines a gau alias for git add --update, which collides with this binary. The README points to issue 8 in the repository for workarounds.
Configuring gau through .gau.toml
gau reads a configuration file on every run. The README says it looks at $HOME/.gau.toml, or %USERPROFILE%\.gau.toml on Windows, and that --config can point somewhere else, with an example such as $HOME/.config/gau.toml. Flags on the command line override whatever the file sets. The repository root contains a .gau.toml, which the README links as the example configuration.
The behaviour when the file is missing is documented and slightly awkward: gau still runs with a default configuration, but prints a message to stderr. Scripts that treat any stderr output as failure will need to account for that. Anyone running gau in a pipeline should either create the file or expect the message.
The practical use of the config file is to fix the options you always want. If you never want png or woff URLs, or you always want a specific proxy, or you always query only wayback, those belong in the file rather than in every invocation. The README does not enumerate which keys the file accepts beyond saying that options set by flags exist; the linked .gau.toml is the reference, so read it rather than guessing key names.
Where gau gives you the wrong answer
The most important limitation is that gau reports history, not current state. Every URL it prints came from an archive or a threat exchange, and none of it was verified against the target. A path that returns 200 today and a path that was deleted in 2019 look identical in the output. If your workflow needs confirmed live endpoints, gau is the wrong layer and you need a step after it that actually requests the URLs.
Coverage is another failure mode. Four providers is not the same as the web. Domains that were never crawled, internal hostnames, and recently registered domains will often produce nothing at all. An empty result is not evidence that a host has no interesting URLs; it is evidence that the archives have nothing for it. The --subs flag widens the query to subdomains, but it cannot create records that were never made.
The filtering flags inherit the same weakness. --mc and --fc rely on status codes recorded at archive time, and --mt and --ft on recorded mime types. Where a provider stores a URL without that metadata, the URL either passes or fails the filter for reasons the user cannot see. The README does not document per provider metadata behaviour, so a filter that looks precise may be silently dropping or keeping entries.
Finally, there is the version history. The go.mod file carries a retract directive covering v2.0.1, v2.0.2, v2.0.3 and v2.0.7. Those versions are withdrawn from the module system, which is a strong signal to pin something newer when you install. The most recent release listed is v2.2.4 from 2024-10-28, and the last push to the repository was on 2026-03-20.
gau against waybackurls and archive APIs
The README names its ancestor: waybackurls by Tomnomnom. The difference in approach is provider count. waybackurls queries the Wayback Machine. gau queries the Wayback Machine plus AlienVault OTX, Common Crawl and URLScan, and lets you select among them with --providers. If your target has no Wayback coverage but appears in a threat exchange feed, that is a case where gau returns data and a single archive client does not.
The other alternative is calling the archive APIs yourself. The Wayback CDX API and Common Crawl's index are public, and a script against them gives you full control over pagination, rate limits and output shape. What you give up is the merging and filtering layer: deduplication across sources, extension blacklisting, parameter collapsing with --fp, date bounds with --from and --to, and JSON output. Writing that yourself is a real amount of work for a result gau already produces.
The honest comparison is that gau is a convenience wrapper with sensible defaults, not a superior data source. It cannot see anything the underlying archives do not have. If you already have a working script against one archive and only care about that archive, gau adds a dependency without adding coverage. If you want several sources merged in one pass, gau is the shorter path.
Editorial conclusion
gau fits reconnaissance work where you already have a domain list and want historical URLs without touching the target. It does not fit anyone expecting a live crawl, a verified inventory of what is currently served, or a single source of truth: the README lists four providers and the output is only as complete as their archives. Before adopting it, check the .gau.toml example and confirm which providers you actually want, because the default set queries all of them and the retracted v2.0.x versions in go.mod mean you should pin v2.2.4 or later.
Frequently asked questions
What is gau used for?
gau fetches known URLs for a domain from AlienVault's Open Threat Exchange, the Wayback Machine, Common Crawl and URLScan, and prints them one per line or as JSON. It is used to assemble a list of historical endpoints and parameters for reconnaissance, without sending requests to the target itself.
What does gau do?
It queries the four providers named in the README for URLs associated with the domain you give it, then optionally filters the results by status code, mime type, extension or date range. The output is a candidate list drawn from archives, not a verified set of live URLs.
How do I install gau on Linux?
The README gives three routes: go install github.com/lc/gau/v2/cmd/gau@latest, a clone and build from the cmd directory, or a pre-built binary from the releases page extracted and moved onto your PATH. A Docker image at sxcurity/gau:latest is also published.
Where does gau read its configuration file?
gau looks for $HOME/.gau.toml, or %USERPROFILE%\.gau.toml on Windows, and --config can point to another path. If the file is absent, gau still runs with defaults but writes a message to stderr.
Why does gau conflict with an alias in zsh?
ohmyzsh's git plugin defines gau as an alias for git add --update, which collides with this binary. The README links to issue 8 in the repository for workarounds.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/lc-gau)