URLFinder: extracting JS, URLs and secrets from a page with one Go binary
一款快速、全面、易用的页面信息提取工具,可快速发现和提取页面中的JS、URL和敏感信息。
At a glance
- What is it?
- URLFinder is a Go command line tool that crawls a page, follows its JavaScript, and pulls out URLs, JS files and sensitive strings. It is built for reconnaissance, and it trades precision for coverage on purpose.
- Who is it for?
- Adopt URLFinder if you do web reconnaissance and want a single binary that walks JavaScript and reports URLs, JS files and sensitive strings with status codes attached. Do not adopt it if you need a low-noise asset inventory, if you cannot send traffic to the target, or if you need a maintained library rather than a CLI.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 106 days ago.
- What is it written in?
- Mainly Go, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What URLFinder extracts, and the reconnaissance gap it fills
URLFinder is a page information extraction tool written in Go. The README describes it as analysing the JavaScript and URLs on a page to find sensitive information or unauthorised API endpoints hidden inside them. The intended reader is someone doing web reconnaissance: a pentester, a bug bounty hunter, or an engineer mapping an application's front end before an assessment.
The author states the motivation directly. The project was started because JSFinder often returned empty results or incomplete link lists, and its author had stopped fixing bugs. URLFinder is the response to that: same broad idea, but with deeper JavaScript traversal, status code verification, and export formats.
The scope is deliberately narrow. URLFinder does not scan for vulnerabilities, does not run a template engine, and does not fingerprint servers. It reads pages and JavaScript, applies regular expressions, verifies what it finds with HTTP requests, and writes the result out. If your problem is "I have one host and I want to know what endpoints and secrets its front end mentions", that is the problem this solves.
Crawl modes, depth and the fuzz pass
The execution model is a crawl with configurable depth. The -m flag selects the mode: 1 is normal crawling, 2 is deep crawling, and 3 is "safe deep crawling". The README explains the difference between 2 and 3 as filtering out sensitive routes such as delete and remove. The specific keywords live in the risks key of the configuration file, so the safe mode is only as good as that list.
Depth is split by target type. The config keys urlSteps and jsSteps control how many layers the crawler descends into URLs and JavaScript respectively. The README notes that deep mode goes one layer into URLs and three into JavaScript, with the stated goal of preventing the crawl from drifting off target.
There is a separate fuzz stage triggered by -z. According to the README, it takes the 404 directories and paths the crawl already found, treats them as a dictionary, and recombines them to hit valid paths. The point is to recover from broken path concatenation in the front end. It only applies to links under the main domain, requires -s to be set, and offers three levels: 1 for directory decrement, 2 for two-level combination, 3 for three-level combination. The README itself suggests level 3 only for a small number of links, which is a fair hint that the combination count grows quickly.
One design decision deserves attention. The README states that to improve compatibility and avoid missing links, the project gave up a low false positive rate: wrong links increase, missed links decrease. Filtering with -s 200 is offered as the remedy, but the README explicitly does not recommend looking only at 200. That is an honest description of a trade-off, and it should shape how you read the output.
Installing URLFinder and running a first scan
The repository ships prebuilt binaries for Windows, Linux and macOS, produced by GoReleaser from a pushed tag. Downloading the binary for your platform and running it is the shortest path. The README's own quick start examples use the Windows executable name, so on Linux or macOS substitute the binary you downloaded.
A single URL scan with all status codes shown looks like this. The -s all flag means every status code is displayed, and -m 3 selects safe deep crawling.
URLFinder.exe -u http://www.baidu.com -s all -m 3To restrict output to specific status codes, pass them comma separated. The README gives 200 and 403 as the example pair.
URLFinder.exe -u http://www.baidu.com -s 200,403 -m 3For a batch, put one URL per line in a text file and pass it with -f, then choose an output directory with -o. A single dot means the current directory. The -ff variant treats every crawled URL as one result set, printing and writing only one output instead of one per target.
URLFinder.exe -s all -m 3 -f url.txt -o .
URLFinder.exe -s all -m 3 -ff url.txt -o .If you need to change request headers, extraction regexes or the risk keyword list, run with -i. The README states that if config.yaml does not exist in the current directory, the program creates a default one and exits, so the first run with -i is a setup step rather than a scan. The commonly edited keys are proxy, timeout, thread, urlSteps, jsSteps, max, headers, jsFind, urlFind, infoFind, jsFiler, urlFiler, risks, jsFuzzPath. Two rules matter when you edit them: extraction regexes (jsFind, urlFind, infoFind) must contain at least one capture group and the first group is used as the result, while filter regexes (jsFiler, urlFiler) need no capture group. The README notes that infoFind replaces the older infoFiler key, which is still read for compatibility. Startup validates the regexes and runtime parameters, so thread, timeout and max must be greater than zero.
If you prefer to build it, the README recommends Go 1.26.4 or a later security patch release. The local verification sequence is go mod tidy, go test ./..., go vet ./..., go build ./..., go test -race ./..., and govulncheck. Cross-compilation is documented with CGO_ENABLED=0 and a GOOS/GOARCH pair, for example:
SET CGO_ENABLED=0
SET GOOS=linux
SET GOARCH=amd64
go build -ldflags "-s -w -X github.com/pingc0y/URLFinder/cmd.Update=dev" -o ./URLFinder-linux-amd64Where URLFinder produces noise instead of answers
The false positive trade-off is the first limitation, and it is documented rather than incidental. Expect dead links in the output. The README's suggested filter, -s 200, removes them, but the same README warns against relying on 200 alone, because a 403 or a 302 can still be a real endpoint. You are left choosing between a clean list that may hide something and a noisy list you have to triage by hand.
The fuzz mode has a narrower failure mode. It only operates on links under the main domain, so a target that spreads its front end across several hostnames will not get the same treatment everywhere. Level 3 recombination is explicitly recommended only for small link sets, which means the feature does not scale to a large crawl.
Deep crawling is also not free at the network level. The default thread count is 50 and the default timeout is 5 seconds, and the tool verifies status codes with real requests. Pointed at a production host without coordination, that is a lot of traffic from one binary. The default is aggressive; the configuration file is where you lower it.
Finally, URLFinder is not a general crawler. It reads HTML and JavaScript looking for patterns. Content rendered entirely client side after the initial response, or endpoints that only appear in a WebSocket handshake, are outside what a regex-over-response-body approach can see. The README does not claim otherwise, but the tool's name invites the wrong expectation.
URLFinder compared with a general purpose crawler
The closest comparison in the same problem space is a general purpose web crawler that also parses JavaScript, such as a headless browser driven crawler. The difference is architectural. A headless browser executes the page and observes network requests, so it sees endpoints that are constructed at runtime. URLFinder does not execute anything: it fetches responses and applies regular expressions, which is why it is a single static Go binary with no browser dependency, and why it is fast enough to run across a batch of hosts from a shell loop.
The cost of that choice is exactly the coverage gap described above. A browser-driven crawler will find dynamically assembled API paths that URLFinder's regexes never see, and it will not report dead links as often, because it only records requests the page actually made. In exchange it needs a browser runtime, is slower per host, and is harder to script into a tight loop.
There is a second axis: URLFinder versus the tool it was written to replace, JSFinder. The author's stated reason for building URLFinder was incomplete or empty results from JSFinder and the absence of upstream fixes. The differences visible in the README are the deeper JavaScript traversal, the status code verification pass, the fuzz stage, and the CSV, JSON and HTML export. If JSFinder already works for you, the upgrade case rests on those additions rather than on a different extraction philosophy.
Maintenance, licensing and the cost of upgrading
The repository is not archived, and the last push was on 2026-06-17. The most recent release is tagged 2026.6.16, published on 2026-06-16. The two releases before it are from September 2023, which is worth noting: the project had a long quiet period and then a burst of maintenance work in June 2026.
That burst is substantive. The release notes for 2026/6/17 list new GitHub Actions CI covering test, vet, build, race and govulncheck, a GoReleaser v2 release configuration that injects the version number from the tag, capture group validation for configuration regexes, runtime parameter validation, exact status code matching, timeout control when auto-detecting the protocol, and a set of fixes including fuzz result deletion, CSV ID grouping headers, and an array out of bounds caused by an abnormal source. The 2026/6/16 entry adds a response body size limit to reduce memory use on unusually large responses, plus fixes for non-HTTP references being concatenated into the target URL, a hang when the status verification channel had zero capacity under a small thread count, and HTML output escaping.
Upgrade cost is low in normal use. The binary is self-contained and flag driven. The one migration detail the README calls out is the configuration: new setups should use infoFind, while the older infoFiler key is still read for compatibility. If you carry a config.yaml forward from an older version, check that every extraction regex still has a capture group, because the 2026/6/17 release added validation that will reject configs which previously loaded. Status code filtering also changed to exact matching, so a filter that used to match loosely may now behave differently.
URLFinder is MIT licensed. That is a permissive licence, and the practical implication is that you can use it internally, modify it, and redistribute it, provided the copyright notice and licence text are preserved. This is a description of the licence, not legal advice; check the LICENSE file in the repository and your own organisation's rules before shipping it inside a product.
Editorial conclusion
Adopt URLFinder if you do web reconnaissance and want a single binary that walks JavaScript and reports URLs, JS files and sensitive strings with status codes attached. Do not adopt it if you need a low-noise asset inventory, if you cannot send traffic to the target, or if you need a maintained library rather than a CLI. Before trusting it, run a small batch with -s 200,403 and check the generated config.yaml against your own regex rules, because the extraction rules and the risk keyword list are the part you will end up editing.
Frequently asked questions
How do I install URLFinder?
The repository publishes prebuilt binaries for Windows, Linux and macOS through GoReleaser, so downloading the binary for your platform is the shortest path. Building from source is also documented, with Go 1.26.4 or a later security patch release recommended.
What does the -m 3 safe deep crawl mode do in URLFinder?
It performs a deep crawl while skipping dangerous paths. The README describes it as filtering routes such as delete and remove, and the specific keywords come from the risks key in config.yaml.
Why does URLFinder return so many dead links?
The README states that to improve compatibility and reduce missed links, the project gave up a low false positive rate, so incorrect links increase. It suggests filtering with -s 200 but does not recommend looking only at 200 status codes.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/pingc0y-urlfinder)