goclone: mirroring a website to a local directory with one Go command
Website Cloner - Utilizes powerful Go routines to clone websites to your computer within seconds.
At a glance
- What is it?
- goclone is a small Go CLI that downloads a site's HTML, CSS, JavaScript and images and rewrites the links so the copy browses from disk, with flags for cookies, proxies and a built in local server.
- Who is it for?
- goclone is a good fit when you want a browsable copy of a small site for offline reading, for archiving a documentation set, or for reviewing a page's assets without a hosted preview. It is the wrong tool for a site built by a JavaScript framework, for anything behind a login, or for scraping at scale, because it follows the HTML it is given and does not execute a browser.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 126 days ago.
- What is it written in?
- Mainly Go, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Installing the binary with brew or go install
There are two documented install paths and both are short. The Homebrew route taps a custom tap first:
# tap
brew tap goclone-dev/goclone
# install tool
brew install gocloneThe manual route is a single `go install` with a stated minimum of Go 1.20:
# Go version >= 1.20
go install github.com/goclone-dev/goclone/cmd/goclone@latestThat version floor in the README is out of step with the repository's own module file, which declares Go 1.25.0. If you build from source on an older toolchain, the module declaration is the line that decides whether the build works.
Building by hand from a clone is also documented, with the binary moved onto your PATH as an optional last step:
# Clone the repository
git clone https://github.com/goclone-dev/goclone.git
cd goclone
# Build and run
go build -o goclone cmd/goclone/main.go
# Move binary to a directory in your PATH (optional)
mv goclone /usr/local/bin/The command entry point lives under `cmd/` and the library code under `pkg/`, which is the layout that makes the `go install` path work at all.
Running a clone is one command and a URL
The whole example in the README is a URL and nothing else:
# goclone <url>
goclone https://configtree.coWhat that produces is a local directory named after the site, containing the HTML, CSS, JavaScript, images and other files fetched from the server. The claim worth examining is the link rewriting: the README says goclone arranges the original site's relative link structure, so that opening a page of the mirrored site in a browser lets you browse from link to link as if you were online. That is the difference between a mirror and a pile of downloads, and it is implemented by rewriting hrefs and srcs to point at local paths as the crawler walks.
Every option the tool has is on the flag list, and it is short enough to read in full:
Usage:
goclone <url> [flags]
Flags:
-C, --cookie strings Pre-set these cookies
-o, --open Automatically open project in default browser
-p, --proxy_string string Proxy connection string. Support http and socks5
-s, --serve Serve the generated files using Echo.
-P, --servePort int Serve port number. (default 5000)
-u, --user_agent string Custom User AgentThe three that change behaviour most are `--cookie` for a site that gates content behind a session, `--user_agent` when the target blocks unknown agents, and `--serve` when you want the copy available over HTTP on port 5000 rather than opened as local files.
What the crawler is built on: colly, goquery and Echo
The module file makes the architecture legible because the dependency list is short. The crawler itself is `github.com/gocolly/colly/v2 v2.3.0`, which handles the request queue, the callback model and link discovery. HTML manipulation is `github.com/PuerkitoBio/goquery v1.12.0`, which is the jQuery-like selector library colly itself is built on, and that combination is how link rewriting is done without a full HTML parser of its own.
Serving the mirrored copy uses `github.com/labstack/echo/v4 v4.15.1`, which is why the `--serve` flag description names Echo specifically. The command line is `github.com/spf13/cobra v1.10.2`. Two smaller dependencies stand out: `github.com/torden/go-strutil` for string handling and `github.com/yosssi/gohtml` for HTML parsing.
Two indirect dependencies hint at behaviour the README does not mention. `github.com/temoto/robotstxt` arrives through colly, so the crawler carries a robots.txt parser, and `github.com/common-nighthawk/go-figure` is an ASCII art banner library, so the tool prints a text logo on startup.
The v1.2.0 release notes from 2021-11-26 show the design settling: an entry named "Removing SSL-Check", another that ignores base64 image encoded downloads, one that returns errors and uses `context.Context` for cancellation while allowing cookies, one that sets a custom user agent, and one that sets a proxy server while updating the colly version. Those five commits are the clearest description of the crawler's actual boundaries.
What the mirror will not contain
goclone walks HTML. It does not run a browser, so anything a JavaScript framework assembles after page load is absent from the copy. A single page application whose routes exist only in a bundle will mirror as its landing shell, and following a link will produce a page with no content in it. That is not a bug in the tool, it is what a crawler is.
The commit history shows other boundaries being discovered and patched rather than designed around. A change named "ignore base64 image encoded download" says data URI images are skipped, so an inlined image does not appear as a file. Another notes that `mime.ParseMediaType` does not keep case and should not be used for parsing cookies, which is the kind of detail that turns into a real bug when a site sends `Content-Type` parameters with unusual casing.
The "Removing SSL-Check" entry is worth naming without dramatising it. Whatever certificate validation that check performed in an earlier version is no longer part of the tool. If you point goclone at a site with an incomplete certificate chain, expect the failure to surface from the HTTP client rather than from a dedicated check, and treat any downloaded page as untrusted input rather than something to open casually on a machine that matters.
Proxy support is limited to what the README states: an http or socks5 connection string, passed to colly's `SetProxy`. There is no flag for authentication beyond cookies, no rate limiting, and no depth or page count limit.
Release history and the project's current shape
Three releases are on record. v1.2.0 on 2021-11-26 is the substantive one, a long list of refactors and feature commits including making `cmd/goclone` installable with `go get`, removing the Makefile, supporting blob URLs, and cleaning up clone parameters. v1.2.1 on 2023-06-19 is mostly maintenance: TravisCI removed in favour of GitHub Actions, the browser-open command updated per operating system, Go moved to 1.20.1, and a series of dependabot bumps.
v1.2.2 on 2025-08-20 continues that pattern. It is dominated by dependency bumps, including golang.org/x/crypto from 0.1.0 to 0.31.0 and protobuf from 1.24.0 to 1.33.0, a gorelease config update, a dependabot configuration, and one behavioural change described as improving server handling with context and goroutines. The tag also adds release information to the command's own output.
The last push to master was on 2026-06-02, and the repository is MIT licensed with the license file also rendered as a badge in the README. The repository tree is deliberately small: `cmd/`, `pkg/`, `docs/`, `testutils/`, a `.goreleaser.yml` for building binaries, and no Makefile, which matches the v1.2.0 decision to drop one. The project is narrow on purpose, and the topics list, cloning, crawler, website-cloner and website-scraper, says exactly what it is.
Editorial conclusion
goclone is a good fit when you want a browsable copy of a small site for offline reading, for archiving a documentation set, or for reviewing a page's assets without a hosted preview. It is the wrong tool for a site built by a JavaScript framework, for anything behind a login, or for scraping at scale, because it follows the HTML it is given and does not execute a browser. The three flags that matter in practice are --cookie for an authenticated area, --serve if you want the mirrored copy served over HTTP rather than opened from the filesystem, and --proxy_string if the target is only reachable through one. Start from the v1.2.2 binary on 2025-08-20, run it once against a site you control, and read the written output before trusting it with anything larger.
Frequently asked questions
What is the best web cloner for offline use?
For a small static or server rendered site, goclone does the job in one command and rewrites links so the copy browses from disk. It does not run JavaScript, so a framework built site will mirror as its shell, and it has no page limit or rate limiting, so it is not a scraping tool.
How do I clone a site that requires a login?
Use the -C or --cookie flag to pre-set cookies before the crawl starts, and the -u flag if the site blocks unknown user agents. Cookies are set through colly before the crawl, which is what the v1.2.0 change using context.Context for cancellation and allowing cookies established.
Can goclone serve the mirrored site over HTTP?
Yes. The -s or --serve flag serves the generated files using Echo, and -P sets the port, which defaults to 5000. Without it, the -o flag opens the project in your default browser as local files instead.
Does goclone work behind a proxy?
The -p or --proxy_string flag takes a proxy connection string and supports http and socks5, passed through to colly's SetProxy. No other network configuration is exposed, so there is no separate setting for headers, retries or timeouts.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/goclone-dev-goclone)