projectdiscovery/katana: a Go crawler built for automation pipelines
A next-generation crawling and spidering framework.
At a glance
- What is it?
- katana is a crawling and spidering framework from ProjectDiscovery that runs in standard or headless mode and emits URLs, forms and classified endpoints. It suits engineers who already script reconnaissance and want crawl output they can pipe into another tool.
- Who is it for?
- Adopt katana if you already run ProjectDiscovery tooling and want crawl output that stays in a pipeline: URLs on stdout, JSONL with -jsonl, scope narrowed with -e and -fsu. Skip it if you need a graphical crawler for a one-off site audit, or if you cannot install Chrome, because headless mode depends on it and the README's Ubuntu recipe installs google-chrome-stable.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 1 day ago.
- What is it written in?
- Mainly Go, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What katana is for, and who ends up using it
katana is a command line crawler written in Go, published under the MIT licence by ProjectDiscovery. The README describes it as "a fast crawler focused on execution in automation pipelines offering both headless and non-headless crawling". That sentence sets the audience: people who run crawls as one stage of a longer job, not people who want to click through a site map.
The practical use is reconnaissance. You point katana at a host, it walks links, parses JavaScript, and prints the URLs it found. Those URLs are the input to whatever comes next, which in the ProjectDiscovery ecosystem usually means a scanner. The tool also extracts forms, detects technologies, and can classify pages and endpoints through a knowledge base. Everything is flags and stdout, so it composes with shell pipelines and with other binaries.
If you have never scripted a crawl before, the flag list is long. The help output alone runs to dozens of switches across INPUT, CONFIGURATION and beyond. That breadth is the point: the defaults are conservative, and the interesting behaviour is opt-in.
Standard mode, headless mode, and what actually gets parsed
The crawler has two execution paths. Standard mode issues HTTP requests and parses responses, which is fast and cheap. Headless mode drives a real browser through the go-rod library, which is what the repository's go.mod lists as github.com/go-rod/rod. Headless mode is what lets katana see content that only exists after JavaScript runs, and the README's Docker example shows the flag pair: -system-chrome -headless.
The parsing layer is not one parser. HTML goes through goquery, per the same go.mod. JavaScript gets its own treatment: -jc enables endpoint parsing and crawling inside JavaScript files, and -jsl turns on jsluice parsing, which the help text flags as "memory intensive". Those are two different mechanisms with two different costs, and turning both on for a large site is a decision about RAM, not just coverage.
Crawl order is a flag, not a fixed rule. -s takes depth-first or breadth-first, defaulting to depth-first. Depth is capped by -d, default 3. There is also a wall-clock cap, -ct, which accepts s, m, h or d and defaults to seconds. A crawl can therefore end because it ran out of links, hit the depth limit, or hit the duration limit, and the output does not necessarily tell you which.
The knowledge base is the piece worth understanding before you rely on it. -kb enables page-type and form classification, and the README states the model is auto-downloaded. That means the first run with -kb reaches out to fetch a model, and an air-gapped machine will not get one. The related extractors, -kb-secrets, -kb-validate-secrets and -kb-endpoints, sit behind the same switch, and -kb-validate-secrets is explicitly documented as sending live API calls to validate detected secrets. That is a network-visible action against third parties, and it is opt-in for a reason.
Installing katana and running a first crawl
The primary install path is Go. The README states katana requires Go 1.26 or newer, and the go.mod confirms go 1.26. Note the environment variable in the command: CGO must be enabled, which means a C toolchain has to be present on the build machine.
CGO_ENABLED=1 go install github.com/projectdiscovery/katana/cmd/katana@latestAfter that, katana lands in your Go binary directory. Running the help command prints the full flag list, which is the fastest way to confirm the binary works before pointing it at anything.
katana -hIf you would rather not build, the README points at the release page for pre-compiled binaries, and there is an official image. Pulling and running it in standard mode against a URL is two commands, and the entrypoint is already set to katana in the Dockerfile.
docker pull projectdiscovery/katana:latest
docker run projectdiscovery/katana:latest -u https://tesla.comFor headless crawling the README shows the same image with the browser flags. The Dockerfile installs chromium in the runtime stage, so the container has a browser to drive.
docker run projectdiscovery/katana:latest -u https://tesla.com -system-chrome -headlessOn Ubuntu, the README's recommended path installs prerequisites through snap and apt, including google-chrome-stable from Google's repository, then installs katana with the same go install command shown above. That Chrome step is not optional decoration: headless mode needs a browser, and the README supplies one.
A first real run should be narrow. Point katana at a single host, keep the default depth of 3, and let it print to stdout. Add -jsonl when you want structured records instead of bare URLs, and add -e to drop hosts you do not want in the results. Reach for -headless only after the standard run tells you the page content is JavaScript-dependent.
Scope control and filters, where most of the tuning happens
Crawl output is only useful if it stays inside the boundary you meant. katana's scope controls are the part of the tool most likely to need adjustment on a real target.
-exclude, also written -e, drops hosts matching a filter, and the help text names the accepted forms: cdn, private-ips, cidr, ip and regex. That is a set of categories rather than a single pattern, so excluding a CIDR and excluding a CDN are the same flag with different arguments.
Two filters address a specific failure of naive crawlers. -ignore-query-params stops katana treating the same path with different query values as separate pages. -filter-similar goes further and skips URLs that look structurally alike, with the help text giving /users/123 and /users/456 as the example. The threshold is tunable through -fst, which defaults to 10 distinct values before a path position is treated as a parameter. That default is a judgement call: a site with fewer than ten users will not trip it, and a site with thousands will.
There is a separate mechanism for page content rather than URLs. -page-content-similar, aliased as -similarity-deduplication, compares pages using simhash, tfidf or bm25, selected by -pcsm and defaulting to simhash. The tuning surface is real: -pcsd sets the simhash hamming distance (default 3), -pcst sets the tfidf and bm25 score floor (default 0.85), and -pcsn limits how many pages are fully processed per similarity cluster (default 1). Read those three together, because they interact. A tighter distance and a higher score floor will suppress more pages, and -pcsn decides how much work each surviving cluster still costs.
-knowledge-base sits outside this group. It classifies rather than filters, and it depends on a downloaded model, so it is the wrong switch to enable when you need a crawl to run with no outbound traffic beyond the target.
Where katana is the wrong tool
The README is thin on failure modes, so the constraints have to be read off the flags and the build files. Several are worth stating plainly.
CGO is the first. The documented install command sets CGO_ENABLED=1, and the Dockerfile installs gcc and musl-dev in the build stage for that reason. Cross-compiling a static binary without a C toolchain is not the documented path, and the Makefile shows the project working around this itself: it only applies -extldflags "-static" when the host is not darwin.
Headless mode is the second. It needs a browser, and the README's Ubuntu recipe installs google-chrome-stable system-wide. On a machine where you cannot install Chrome, headless crawling is unavailable, and the fallback is standard mode with -jc, which will miss anything rendered only at runtime.
The knowledge base is the third. It auto-downloads a model, which the README does not describe in terms of size, hosting or offline behaviour. If your environment blocks that fetch, -kb and its extractors are unusable, and the README does not document a local model path.
Memory is the fourth. -jsl is labelled memory intensive in the tool's own help output, and the knowledge base adds model inference on top. On a large JavaScript-heavy site, the combination is where a crawl is most likely to fall over.
Finally, katana writes URLs and records. It does not tell you whether a finding matters. If what you want is a vulnerability verdict rather than a surface map, katana is a stage in that pipeline, not the pipeline.
How katana differs from a general purpose crawler
The obvious comparison is a general purpose web crawler such as Scrapy, and the difference is not speed claims. It is where the boundary of the tool sits.
Scrapy is a framework you write against. You define spiders, item pipelines and middleware in Python, and the crawl logic is yours. katana is a binary you invoke. The crawl logic is fixed and exposed through flags, and the extension point is the output format rather than the code. If your crawl needs custom extraction rules, pagination handling that no flag covers, or state carried across requests, Scrapy gives you a place to put that code and katana does not.
What katana gives in exchange is a preset reconnaissance posture. Scope exclusion by CIDR or private-IP range, similar-URL filtering with a tunable threshold, page-content similarity across three algorithms, JavaScript endpoint parsing, technology detection via wappalyzergo, and a knowledge base that classifies pages and endpoints are all switches rather than modules you assemble. For someone whose job is to point a crawler at a host and hand the URL list to a scanner, that is less code to maintain.
The trade-off is visible in the flag list. When the behaviour you need is not a flag, you are not extending katana, you are writing a different tool. The repository layout reflects this: cmd/ holds the entry point, internal/ and pkg/ hold the implementation, and integration_tests/ holds a shell-driven test runner rather than a plugin surface.
Maintenance, upgrade cost and licence
The repository is not archived, and the last push was on 2026-09-21. Releases are regular rather than constant: v1.7.0 on 2026-08-05, v1.6.1 on 2026-05-05, and v1.6.0 on 2026-05-04. The default branch is dev, which is worth knowing if you build from source rather than from a release tag.
Upgrade cost is dominated by the Go version floor. The README states Go 1.26 or newer, and go.mod agrees, so a machine pinned to an older toolchain cannot build katana from source at all. The Dockerfile pins golang:1.27.1-alpine and alpine:3.24.1, so container users inherit the project's own build environment and avoid that problem. The runtime image also installs chromium, which is the largest single component in it.
Dependency churn is the other factor. go.mod pulls in a long list of ProjectDiscovery modules, including fastdialer, retryablehttp-go, goflags, gologger and utils, plus go-rod and goquery. Upgrading katana upgrades that set together, and the transitive tree is large. There is a Makefile target for tidying modules and one for golangci-lint, so the project does run its own checks, but that does not reduce the surface you inherit.
The licence is MIT, which is permissive and places few conditions on redistribution or commercial use. One caveat that is not a licence question but sits next to it: the knowledge base model is auto-downloaded at runtime, and the README does not state the terms attached to that model or where it is hosted. If your organisation reviews downloaded artifacts, that fetch is the thing to look at, not the MIT header on the source.
Editorial conclusion
Adopt katana if you already run ProjectDiscovery tooling and want crawl output that stays in a pipeline: URLs on stdout, JSONL with -jsonl, scope narrowed with -e and -fsu. Skip it if you need a graphical crawler for a one-off site audit, or if you cannot install Chrome, because headless mode depends on it and the README's Ubuntu recipe installs google-chrome-stable. Before committing, verify that the knowledge base model downloads on first use, that Go 1.26 or newer is present, and that CGO_ENABLED=1 works on your build machine. The MIT licence permits commercial use, but the knowledge base model is downloaded at runtime and its own terms are not stated in the README.
Frequently asked questions
How do I install projectdiscovery/katana?
The README's primary path is a Go install with CGO enabled, requiring Go 1.26 or newer. Pre-compiled binaries are available from the release page, and there is an official Docker image at projectdiscovery/katana. On Ubuntu the README also walks through installing prerequisites, including google-chrome-stable, before the Go install.
Can I run projectdiscovery/katana on Kali Linux?
The README does not give Kali-specific instructions. It documents a Go install, Docker, and an Ubuntu recipe that installs prerequisites through snap and apt. On a Debian-derived system like Kali, the Ubuntu steps are the closest documented starting point, but the README does not confirm they work unchanged.
How do I use projectdiscovery/katana in headless mode?
Pass -headless together with -system-chrome. The README's Docker example is docker run projectdiscovery/katana:latest -u https://tesla.com -system-chrome -headless. Headless mode drives a real browser through go-rod, and the project's Dockerfile installs chromium in the runtime image.
What do the -jc and -jsl flags do in projectdiscovery/katana?
-jc enables endpoint parsing and crawling inside JavaScript files, while -jsl enables jsluice parsing in JavaScript files. The tool's own help output marks -jsl as memory intensive, so the two flags carry different resource costs.
Does projectdiscovery/katana need Chrome installed?
Only for headless crawling. Standard mode issues HTTP requests and does not need a browser. The README's Ubuntu instructions install google-chrome-stable, and the Dockerfile installs chromium, so both documented headless paths bring a browser with them.
What is the knowledge base in projectdiscovery/katana?
It is an ML page-type and form classification feature enabled with -kb. The README states the model is auto-downloaded, and related extractors for secrets and endpoints sit behind the same switch, including -kb-validate-secrets, which the help output says sends live API calls to validate detected secrets.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/projectdiscovery-katana)