hakrawler: a stdin-driven Go crawler for endpoint discovery
Simple, fast web crawler designed for easy, quick discovery of endpoints and assets within a web application
At a glance
- What is it?
- hakrawler reads URLs from standard input and prints the links, scripts and form targets it finds, which makes it a small filter in a recon pipeline rather than a full scanner. It is simple, it is GPL-3.0, and its documentation is thin on everything except usage examples.
- Who is it for?
- Adopt hakrawler if your recon workflow already pipes URLs between tools and you want a crawler that fits that shape: it reads stdin, writes stdout, and takes about a dozen flags. Do not adopt it if you need a crawler that manages its own queue, persists state, or explains what it did; the README documents none of that, and the single-file layout at the repository root suggests there is not much under the surface.
- Can I use it commercially?
- Yes, with conditions. GPL-3.0 is a copyleft licence: if you distribute software that includes it, you must release that software's source code under the same licence. Running it internally without distributing it does not trigger that obligation.
- Is it still maintained?
- Yes. The repository last received commits 57 days ago.
- What is it written in?
- Mainly Go, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What hakrawler is for, and who ends up using it
The problem hakrawler addresses is narrow: given a list of hosts that already respond, find the URLs and JavaScript file locations reachable from them. The README describes it as a "Fast golang web crawler for gathering URLs and JavaScript file locations" and says it is "basically a simple implementation of the awesome Gocolly library." That framing matters. This is not a scanner, a fuzzer or a content discovery tool. It does not guess paths. It follows what the pages actually link to.
The intended user is someone doing reconnaissance, which the repository topics confirm: bugbounty, crawling, hacking, osint, pentesting, recon, reconnaissance. In practice that means a penetration tester or bug bounty hunter who has already enumerated subdomains and resolved which ones answer HTTP, and now wants the surface area of each. The README's own tool chain shows exactly this shape: take a domain, pipe it through haktrails to get subdomains, then httpx to find the ones that respond, then hakrawler to pull URLs out of them. hakrawler is the third stage, not the first.
If you are looking for something that takes a single domain and does everything, this is the wrong tool by design. It expects its input on standard input, one URL per line, and it expects you to have done the work of finding those URLs.
How the crawl actually works: stdin in, links out
The data flow is the whole architecture. You pipe URLs in, hakrawler crawls each one to a configurable depth, and it prints discovered URLs to standard output. The README's examples are all one-liners built on that contract: `echo https://google.com | hakrawler` for a single URL, `cat urls.txt | hakrawler` for a list.
Underneath, the go.mod file shows a single direct dependency on `github.com/gocolly/colly/v2`, with goquery, htmlquery, xmlquery, robotstxt and a few others pulled in indirectly. So the HTML parsing and link extraction are Colly's, and hakrawler is the wrapper that reads stdin, configures the collector and formats output. The repository root confirms how small that wrapper is: .gitignore, Dockerfile, LICENSE, README.md, go.mod, go.sum and hakrawler.go. One Go file.
Depth defaults to 2, threads default to 8, and page size defaults to -1, meaning no limit. The `-s` flag reports where a URL was found (href, form, script and so on), `-w` reports which link it was found on, and `-u` prints only unique URLs. The `-i` flag restricts crawling to inside the starting path, which is the difference between crawling a whole host and staying inside one application. The `-subs` flag widens scope to subdomains.
That scope question is where most confusion lives. The README warns that a domain specified as https://example.com may redirect to https://www.example.com, and because the subdomain is not in scope, no URLs are printed. The fix it gives is either to specify the final URL in the redirect chain or to pass `-subs`. Read that note before you conclude the tool is broken.
Installing hakrawler and running a first crawl
The README's normal install path assumes Go is already present. It points at the Go install documentation first, then gives a single command that downloads and compiles the binary.
go install github.com/hakluke/hakrawler@latestThat places the binary at `~/go/bin/hakrawler`. Running it by bare name requires adding that directory to your path, and the README suggests `export PATH="~/go/bin/:$PATH"` for the current shell or the same line in `~/.bashrc` to persist it.
If you would rather not install Go, the README offers a Docker path against the published image, passing `-subs` and running interactively with `--rm`:
echo https://www.google.com | docker run --rm -i hakluke/hakrawler:v2 -subsFor a local build, the README clones the repository, builds the image and runs it with `--help` to confirm it works:
git clone https://github.com/hakluke/hakrawler
cd hakrawler
sudo docker build -t hakluke/hakrawler .
sudo docker run --rm -i hakluke/hakrawler --helpThere is also an apt route on Kali, but the README is explicit that it installs an older version without all the features and "may be buggy," and recommends the other methods instead. Worth taking at face value.
A first real run, on a single target with a five second cap per input line, looks like this:
echo https://example.com | hakrawler -timeout 5 -subsYou should see URLs printed to standard output as they are discovered. If nothing appears, check the redirect note above before anything else.
The silent failure that catches most new users
The redirect-to-subdomain case is the limitation the README itself flags, and it deserves more attention than a parenthetical. A target entered as https://example.com that answers with a redirect to https://www.example.com will produce no output, because the crawl stays inside the original scope. The tool is behaving correctly and the operator reads it as a failure. In a pipeline where hakrawler is one stage among several, that silent empty result is easy to misread as "this host has no endpoints" rather than "this host redirected."
There are other constraints worth naming. There is no documented output for errors, no documented way to resume an interrupted crawl, and no documented rate limiting beyond the thread count. The `-timeout` flag caps the time spent crawling each URL from stdin, defaulting to -1, which the option list presents as no limit; on a large input list with slow hosts, that default is a decision you have to make deliberately. The `-size` flag caps page size in KB, also defaulting to -1.
The dependency situation is a trade-off too. Pinning to a specific Colly pseudo-version, as go.mod does with a dated commit hash, is reproducible but means upstream Colly changes arrive only when someone edits that line. The go.mod declares `go 1.16` while the Dockerfile builds on `golang:1.17`, a small inconsistency that says more about how the project is maintained than about whether it works.
hakrawler compared with Gospider and katana
The obvious alternatives people search for are Gospider and katana, and the difference is mostly about what the tool owns. hakrawler owns almost nothing: it reads stdin, crawls, prints. Gospider is a broader reconnaissance crawler that handles more of the surrounding work such as site listing and sitemap and robots handling, so the pipeline you build around hakrawler is partly built in there. That is a real difference in approach, not a feature checklist. If you want one binary that does the enumeration and the crawling, hakrawler is the wrong shape.
katana, from ProjectDiscovery, is the closer comparison in spirit: a modern crawling framework with a broader feature set and its own ecosystem conventions. Choosing between them is largely about which pipeline you already live in. If your workflow is ProjectDiscovery tools end to end, katana fits without adapter work. If your workflow is hakluke's own tooling, where haktrails and httpx feed a crawler over stdin, hakrawler slots in with no glue at all.
The honest summary is that hakrawler's advantage is its narrowness. Eight threads, depth 2, stdin to stdout. There is very little to configure and very little to go wrong, and the codebase is one file. If you want a crawler you can read in an afternoon, that is the argument for it.
Maintenance, licensing and upgrade cost
The repository is not archived, and the last push was on 2026-08-05, which is recent enough that the project is not abandoned. That said, the release history tells a different story about versioning: 2.1 dates to 2022-05-23, 2.0 to 2021-07-05, and v1.0 beta to the same day as 2.0. Commits continue while tagged releases have not, so if you install via `go install ...@latest` you are tracking the master branch rather than a release. That is the upgrade cost in practice: there is no release cadence to plan around, and no changelog in the repository to read before you pull.
The licence is GPL-3.0, stated in the repository and present as a LICENSE file at the root. For a reconnaissance tool you run as a binary in your own pipeline, that is unremarkable. It matters if you intend to link the code into another program or redistribute a modified version, because GPL-3.0 carries copyleft obligations that permissive licences do not. Whether those obligations apply to your specific use is a question for your own counsel, not for this article.
The apt package on Kali is the one upgrade path the README actively discourages, describing it as an older version without all the features that may be buggy. If you installed that way, the documented recommendation is to switch to `go install` or Docker.
What hakrawler does not do
It does not discover endpoints by guessing. There is no wordlist flag, no fuzzing mode, no way to ask it about a path that is not linked from somewhere it crawled. If a URL is not referenced in the HTML it parses, hakrawler will not find it. That is a fundamental boundary, and it is why the tool pairs with content discovery rather than replacing it.
It does not manage state. There is no documented database, no resume flag, no way to hand it a partially completed crawl. Each invocation starts fresh from whatever is on stdin.
It does not report on itself. The option list has no verbosity flag, no log file, no error output format. You get URLs, and with `-s` you get where each came from, and with `-w` you get which link it came from. That is the entire vocabulary. For a tool meant to be piped into the next stage, that is defensible. For debugging a crawl that returned less than you expected, it leaves you with the redirect note and not much else.
Finally, it does not handle JavaScript execution. The README describes gathering "JavaScript file locations," which is about finding script URLs, not about rendering pages that build their links at runtime. Single-page applications that construct navigation in the browser will look sparse to a crawler that reads the HTML response.
Editorial conclusion
Adopt hakrawler if your recon workflow already pipes URLs between tools and you want a crawler that fits that shape: it reads stdin, writes stdout, and takes about a dozen flags. Do not adopt it if you need a crawler that manages its own queue, persists state, or explains what it did; the README documents none of that, and the single-file layout at the repository root suggests there is not much under the surface. Before relying on it, verify the redirect behaviour yourself with a domain that redirects to a www subdomain, because that is the failure the README calls out by name and the one most likely to make a pipeline look empty when it is working correctly.
Frequently asked questions
How do I install hakrawler?
The README's normal method is `go install github.com/hakluke/hakrawler@latest`, which requires Go to be installed first and puts the binary at `~/go/bin/hakrawler`. There is also a Docker image at hakluke/hakrawler:v2, and a Kali apt package that the README warns is an older, possibly buggy version.
How do I use hakrawler?
You pipe URLs into it on standard input and it prints discovered URLs to standard output, for example `echo https://google.com | hakrawler` for one target or `cat urls.txt | hakrawler` for a list. Flags control depth, threads, timeout, proxy, subdomain scope and output format.
How is hakrawler different from Gospider?
hakrawler is deliberately minimal: it reads URLs from stdin, crawls them with Gocolly, and writes discovered URLs to stdout, with a single Go file in the repository root. Gospider is a broader reconnaissance crawler that handles more of the surrounding enumeration work, which is the trade-off between a small pipeline stage and a larger single tool.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/hakluke-hakrawler)