raviqqe/muffet: a recursive link checker for whole sites
Fast website link checker in Go
At a glance
- What is it?
- Muffet is a Go command line tool that crawls a website, extracts links from HTML tags and reports broken ones. It is fast, multi-format on output, and thin on documentation beyond its install and usage pages.
- Who is it for?
- Adopt muffet if you want a single Go binary that walks a site recursively and fails a CI job on broken links, and if text, JSON or JUnit XML output is enough for your pipeline. Do not adopt it if you need a documented crawl-delay policy, a JavaScript renderer, or per-link retry semantics, because the README does not describe any of those.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 9 days ago.
- What is it written in?
- Mainly Go, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 24, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem muffet solves, and who actually needs it
A site with a few hundred pages accumulates dead links quietly. A page moves, an external reference 404s, an image path changes case, and nothing in the build notices. Muffet exists for that gap. Its README describes it as a website link checker which scrapes and inspects all pages in a website recursively, so the unit of work is the site, not a single URL you paste in.
The intended user is someone who owns a deployed site and can run a command against it: a docs maintainer, a static site author, a platform engineer who wants a link check in a pipeline. The README lists different tag support for a, img, link, script and others, which matters because a checker that only reads anchor hrefs misses broken stylesheets and images. It also lists multiple output formats: text, JSON, and JUnit XML. The JUnit XML option is the one that maps onto CI test reporters, and it is the reason this tool can sit in a build rather than in a manual checklist.
How the crawler is put together in the repository
The top-level file list is unusually readable as an architecture sketch. html_page_parser.go and html_page.go handle turning a fetched document into links. link_finder.go and link_finder_test.go sit next to link_filterer.go, so finding and filtering are separate steps: the finder extracts candidates, the filterer decides which are worth checking. link_fetcher.go and link_fetcher_options.go then do the requests, and link_validator.go judges the responses.
Two files say a lot about the design. concurrent_string_set.go is a visited-set built for parallel access, which is what lets the crawl fan out without revisiting pages. host_throttler.go and host_throttler_pool.go apply per-host limits, so a single slow or rate-limited origin does not get hammered by the full concurrency of the crawl. The HTTP layer is fasthttp_http_client_factory.go and fasthttp_http_client.go, built on valyala/fasthttp rather than net/http. That is the source of the speed claim in the README: fasthttp is a lower-level client, and the trade-off is that it is not the standard library client, so behaviour around redirects, TLS and headers is whatever that library does.
The dependency list confirms the rest. github.com/temoto/robotstxt is present, so robots.txt parsing is part of the picture. github.com/oxffaa/gopher-parse-sitemap handles sitemaps, which gives the crawler a way to seed from a published sitemap rather than only from links. go.uber.org/ratelimit provides the rate limiting primitive behind the throttler. github.com/breml/rootcerts and andybalholm/brotli cover certificate roots and Brotli decoding. cache.go, daemon_manager.go and checked_http_client.go round out the runtime: caching of responses, a daemon manager, and a client wrapper that tracks what has been checked.
Installing muffet and running a first crawl
The README gives one install command, using the Go toolchain. The module path carries the v2 major version, so the command is:
go install github.com/raviqqe/muffet/v2@latestAfter that, the binary is in your Go binary directory. The README points to an install page for anything beyond this command, and the repository also carries a Dockerfile, so a container image is an intended distribution path. That Dockerfile builds with CGO_ENABLED=0 GOOS=linux and copies the binary into a scratch image, with /muffet as the entry point.
The README's usage example is a single positional argument, the site root:
muffet https://shady.bakery.hotlandWhat you should see is a crawl that starts at that URL, follows links across the site, and prints a report of the links it could not resolve. The README states that output can be text, JSON or JUnit XML, and points to a usage page for the flags that select a format. For a build that should fail on a broken link, the JUnit XML form is the one to look at, because your CI runner can consume it as a test report. The README does not document the exact flag names for format selection, so read the usage page before writing the pipeline step.
For a container run, the Dockerfile produces an image whose entry point is the checker itself, so the site URL is passed as the container argument rather than as a shell command. The repository layout also includes goreleaser.yaml, which indicates prebuilt release artifacts are produced by GoReleaser; the README does not describe them, so treat the release page as the source of truth for what is published.
Where muffet is the wrong tool
Muffet checks links. It does not render pages. The HTML parser reads the markup the server returns, so any link produced by client-side JavaScript after load is invisible to it. On a single-page application whose navigation is built at runtime, a crawl will see the shell document and little else. That is a structural limit of the approach, not a bug, and no flag in the README changes it.
The second limit is documentation depth. The README is short and defers to an install page and a usage page. It does not document rollback behaviour, retry counts, timeout defaults, or how the host throttler decides its limits. For a tool you run by hand that is fine. For a tool you put in a shared pipeline where a transient 503 on one external host can fail a release, the absence of documented retry and timeout semantics is a real gap you have to close by reading the source or by testing against a staging host.
Third, a recursive crawl of a large site makes a lot of requests, and the repository's own host throttling exists precisely because that is a problem. If your target is a third-party site you do not control, a recursive checker is the wrong instrument regardless of how polite it tries to be; robots.txt handling being present does not make an aggressive crawl appropriate. Point it at sites you own or have permission to scan.
Alternatives and how their approach differs
The obvious comparison is a link checker built on a headless browser. Those tools load each page, execute scripts, and then inspect the resulting DOM, so they catch links injected by JavaScript and can report on client-side routing. The cost is weight: a browser engine per worker, slower runs, and a much larger install. Muffet takes the opposite position. It parses HTML directly with github.com/yhat/scrape and fetches with fasthttp, which is why the README can lead with massive speed. If your site is server-rendered or static, that trade is clearly in muffet's favour. If it is a JavaScript application, the browser-based approach is the only one that will see the links you care about.
A second alternative is the link checker embedded in a static site generator or a docs framework. Those run at build time against the generated output and know the internal link graph exactly, so they can resolve relative links without any HTTP request at all. Muffet cannot do that: it works over HTTP against a running site, which means it also validates the deployed result, external links included. The two are complementary rather than competing. Build-time checks catch internal breakage before deploy; muffet catches what the deploy did to those links and what third parties have done to theirs since.
Maintenance, licence and upgrade cost
The repository is not archived, and the last push was on 2026-09-22. The most recent release in the list is v2.11.5, dated 2026-06-09, with v2.11.4 on 2026-05-21 and v2.11.3 on 2026-04-26 before it. That is a steady minor release cadence, and the version numbers suggest the project is past its breaking-change phase: v2 has been the major line across all three releases listed.
The module path is versioned, github.com/raviqqe/muffet/v2, and go.mod declares go 1.26.0. Upgrading within v2 is therefore a matter of pulling a new tag; the Go module system will not silently move you to a v3. The build toolchain requirement is the practical upgrade constraint: a machine with an older Go than the go directive states will not build the module, and the Dockerfile pins golang:1.27.1 for its build stage. If you build the container yourself, that base image is what you inherit.
The licence is MIT, which is permissive: it allows use, modification and redistribution with the licence text retained. That is a statement about the licence identifier in the repository, not advice about your situation; if licence terms matter to your organisation, read the LICENSE file and talk to whoever handles that.
Editorial conclusion
Adopt muffet if you want a single Go binary that walks a site recursively and fails a CI job on broken links, and if text, JSON or JUnit XML output is enough for your pipeline. Do not adopt it if you need a documented crawl-delay policy, a JavaScript renderer, or per-link retry semantics, because the README does not describe any of those. Before wiring it into CI, run it against a staging host and check the exit code behaviour and the host throttling defaults, then read the usage page for the Docker and GitHub Actions forms.
Frequently asked questions
What is muffet, the Go link checker?
Muffet is a website link checker which scrapes and inspects all pages in a website recursively, written in Go and distributed under the MIT licence. It supports several HTML tags including a, img, link and script, and can emit text, JSON or JUnit XML.
How do I install muffet?
The README gives a single command, go install github.com/raviqqe/muffet/v2@latest, and points to an install page for more detail. The repository also contains a Dockerfile that builds a static binary into a scratch image.
How do I run a first crawl with muffet?
The README's usage example passes the site root as the only argument, muffet https://shady.bakery.hotland. The crawl starts there, follows links across the site, and reports the links it cannot resolve.
Does muffet check links generated by JavaScript?
Nothing in the README or the repository layout indicates a browser engine or script execution. The parser works on the HTML the server returns, so links created at runtime by client-side JavaScript are not seen.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/raviqqe-muffet)