CLI tool
xroche/httrack avatar
xroche/httrack

HTTrack Website Copier: Mirroring a Site with the httrack Command Line

HTTrack Website Copier, copy websites to your computer (Official repository)

4,771 stars783 forksCGPL-3.0

At a glance

What is it?
HTTrack copies a website to a local directory and rebuilds its relative link structure so you can browse it offline. The repository is the development home of the C tool, its WinHTTrack and WebHTTrack front ends, and the httrack CLI.
Who is it for?
HTTrack fits engineers who need a browsable local copy of a site, whether for archival, migration checks or offline reference, and who are willing to work from the httrack command line or the WinHTTrack and WebHTTrack front ends. It is the wrong tool if you need JavaScript-rendered pages or if the site's robots rules forbid mirroring.
Can I use it commercially?
Yes, with conditions. GPL-3.0 is a copyleft licence: if you distribute software that includes it, you must release that software's source code under the same licence. Running it internally without distributing it does not trigger that obligation.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly C, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What HTTrack actually does with a website

HTTrack is an offline browser. Its job is to download a World Wide website from the Internet into a local directory and to build all directories recursively, pulling html, images and other files from the server. The part that distinguishes it from a bulk downloader is what it does afterwards: it arranges the original site's relative link structure, so opening one page of the mirrored site in a browser lets you follow links as if you were online.

The intended user is someone who wants a site to keep working after the network is gone or the content changes. That covers archival of documentation, checking a site you are migrating away from, and offline reference. It is not a scraping framework and it does not promise to extract structured data. It reproduces pages.

Two secondary behaviours matter in practice. HTTrack can update an existing mirrored site rather than starting over, and it can resume interrupted downloads. For a large site that fails halfway, the resume path is the difference between a usable mirror and a restart.

The httrack command line underneath WinHTTrack and WebHTTrack

The repository ships three faces of the same engine. WinHTTrack is the Windows front end, WebHTTrack is the one for Linux, BSD and macOS, and underneath all of them sits the httrack command line. The README also notes an Android app and a `brew install httrack` path for macOS, plus a release DMG.

That layering explains most of what you see in the tree. The top-level entries include `src/`, `man/`, `lang/`, `templates/`, `html/` and `doc/`, alongside `configure.ac`, `Makefile.am`, `bootstrap` and `build.sh`. The engine is C, the build is autotools, and the front ends are separate presentation layers over the same crawl logic. The README states that HTTrack is fully configurable and has an integrated help system, which is the honest description of a tool whose option surface is wide enough that the GUI exists mainly to hide it.

The practical consequence: anything the GUI can do, the CLI can do, and the CLI is the only interface available on a headless server. If you plan to run mirrors on a schedule from a build machine, you are writing `httrack` invocations, not clicking through WinHTTrack.

Installing HTTrack and running a first mirror

There are two build paths and the README is explicit about the difference. A git checkout ships only the autotools sources, so `./bootstrap`, which runs `autoreconf`, has to regenerate `configure` first; that step needs autoconf, automake and libtool. Released tarballs already include `configure`, so building from a tarball skips `./bootstrap`.

For a checkout, the README gives this sequence:

bash
git clone https://github.com/xroche/httrack.git --recurse-submodules
cd httrack
./bootstrap
./configure --prefix=$HOME/usr && make -j8 && make install

The `--recurse-submodules` flag is not decorative; the repository has a `.gitmodules` file, so a plain clone can leave submodule directories empty and break the build. After `make install`, the `httrack` binary lands under the prefix you chose, which here is `$HOME/usr`.

There is also a one-shot wrapper that runs bootstrap, configure and make together and forwards its arguments to `configure`:

bash
./build.sh --prefix=$HOME/usr

On macOS the README points to `brew install httrack` and to the release DMG as alternatives to compiling. Once the binary is on your PATH, a first mirror is a single invocation naming the site and a destination directory. Open a page from that directory in a browser afterwards and the relative links should resolve locally.

Where HTTrack stops being the right tool

The mirror is a copy of what the server returns, not of what a browser renders. A site that builds its pages in JavaScript, or that loads content through an API after the initial HTML arrives, will come down as the shell HTML and whatever assets the crawler can reach by following links. The README describes downloading html, images and other files and arranging relative link structure; it does not describe executing page scripts. If your target is a single-page application, expect gaps and budget time to find them.

Crawling also has an etiquette dimension the tool cannot settle for you. Mirroring a site you do not control may conflict with its robots rules or terms, and the README does not document a policy layer that decides this on your behalf. The recursion that makes HTTrack useful is the same recursion that can pull far more of a site than you intended if you point it at the wrong host or leave depth settings loose.

Finally, the front ends are not interchangeable in maintenance terms. The README describes WinHTTrack and WebHTTrack as separate front ends over the same command line, and the repository is C with autotools. If your environment cannot build autotools projects and has no package for HTTrack, you are relying on the platform-specific distributions the README mentions rather than on a self-contained build.

HTTrack against wget --mirror

The obvious alternative for anyone already on a Unix system is `wget --mirror`, and the difference is the output rather than the download. Both walk a site and fetch linked resources. HTTrack's stated goal is that you open a mirrored page and browse from link to link as if online, which means it rewrites and arranges the original site's relative link structure so local navigation works. A `wget` mirror preserves the directory tree but does not attempt the same offline-browsing arrangement, and its option set is aimed at retrieval rather than at producing something a browser can walk.

The trade-off runs the other way too. `wget` is already present on most Linux and BSD systems, so there is nothing to build, whereas HTTrack from a git checkout needs autoconf, automake and libtool before `./bootstrap` will run. If you only need the files on disk and do not care about link rewriting, the tool you already have is the cheaper choice. If you want a directory you can hand to someone and say "open index.html", HTTrack's link arrangement is the reason to install it.

Maintenance, licence and upgrade cost

The repository is not archived, and the last push was on 2026-09-22. Releases have been coming steadily: 3.50.1 on 2026-09-02, 3.50.2 on 2026-09-11 and 3.50.3 on 2026-09-18. That cadence suggests the project is being maintained, but the shape of the codebase sets the upgrade cost. This is a C project built with autotools, and the README's own build instructions assume a working toolchain: autoconf, automake and libtool for a checkout, nothing extra for a released tarball.

For most users the upgrade path is a package manager or a new release tarball, not a source rebuild. The `brew install httrack` route on macOS and the release DMG are the low-friction options. If you build from a checkout, upgrading means pulling and re-running `./bootstrap`, `./configure` and `make`, which is cheap once the toolchain is in place and annoying if it is not.

The licence is GPL-3.0, and the repository carries both `COPYING` and `license.txt` at the top level. GPL-3.0 is a copyleft licence, so if you intend to redistribute HTTrack or a modified version, the terms in `COPYING` govern what you must pass on. That is a distribution question, not a usage question, and it is worth reading the file rather than assuming. Nothing here is legal advice; the licence text is the authority.

Editorial conclusion

HTTrack fits engineers who need a browsable local copy of a site, whether for archival, migration checks or offline reference, and who are willing to work from the httrack command line or the WinHTTrack and WebHTTrack front ends. It is the wrong tool if you need JavaScript-rendered pages or if the site's robots rules forbid mirroring. Before committing, verify the licence terms in COPYING, confirm the build path for your platform (released tarball versus git checkout) and check whether the site's robots.txt permits the crawl.

Frequently asked questions

Is HTTrack safe to use?

The README describes HTTrack as an offline browser that downloads a website into a local directory, and the repository carries a SECURITY.md file and a GPL-3.0 licence. Whether a given crawl is appropriate depends on the site's rules and on what you point it at, which the documentation does not decide for you.

What is HTTrack used for?

It downloads a website from the Internet to a local directory, building directories recursively and fetching html, images and other files, then arranges the original relative link structure so you can browse the mirror offline. It can also update an existing mirrored site and resume interrupted downloads.

Is HTTrack free?

The repository is licensed GPL-3.0, with COPYING and license.txt at the top level. The README does not describe a paid tier.

How do I use the httrack command line?

The README states that the httrack command line sits underneath WinHTTrack and WebHTTrack, so anything the front ends do is available from the CLI. Build it from a tarball or a checkout, then invoke the installed binary against a site and a destination directory.

How do I install HTTrack on macOS?

The README says WebHTTrack arrives with `brew install httrack` and in the release DMG. Building from source is the other documented path, but a git checkout requires autoconf, automake and libtool for `./bootstrap`.

How do I install HTTrack on Ubuntu?

The README does not give an apt command. The documented source path is cloning the repository and running `./bootstrap`, `./configure --prefix=$HOME/usr`, `make -j8` and `make install`, or using `./build.sh --prefix=$HOME/usr`.

Official sources

  1. License: GPL-3.0
  2. Project website
  3. README
  4. Releases
  5. xroche/httrack on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/xroche-httrack.svg)](https://hysenlabs.com/projects/xroche-httrack)