Self-hosted service
internetarchive/heritrix3 avatar
internetarchive/heritrix3

Heritrix 3: the Internet Archive crawler, and what changed in 3.16 to 3.17

Heritrix is the Internet Archive's open-source, extensible, web-scale, archival-quality web crawler project.

3,321 stars794 forksJavaNOASSERTION

At a glance

What is it?
A Java web crawler built for archival fidelity rather than speed, shipping through Maven Central and Docker Hub, with a release cadence in 2026 that has been busy fixing how crawls report failures and recognising bot walls.
Who is it for?
Heritrix is the crawler to reach for when the output has to be defensible years later, because WARC plus a crawl log plus a configuration file is a record you can replay, and most general purpose crawlers are not built that way. What the recent releases show is a project in maintenance mode with a specific agenda: better failure reporting, honest bot-block detection, and a packaging migration forced by Maven Central size limits.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 14 days ago.
What is it written in?
Mainly Java, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 23, 2026, and from our analysis. They are not legal advice.

Editorial analysis

An archival crawler, which is a different job from crawling

Heritrix is described in the README as the Internet Archive's open-source, extensible, web-scale, archival-quality web crawler project. The two words that matter are archival and web-scale, and they pull in opposite directions. Archival quality means you cannot lose a response to a transient error, which rules out aggressive timeouts and aggressive retry limits. Web-scale means you still have to cover a large seed set in finite time.

The name gets an unusual paragraph of explanation in the README. Heritrix, sometimes spelled heretrix and misspelled several other ways, is an archaic word for an heiress, a woman who inherits. The stated reasoning is that the crawler seeks to collect and preserve the digital artifacts of our culture for the benefit of future researchers and generations, so the name seemed apt. That framing is the project's whole thesis in one paragraph: the output is an inheritance, and the crawler's job is to make it complete enough to be worth passing on.

The topic tags are minimal and descriptive: heritrix, java, warc, and webcrawling. The project is written in Java, has roughly 3,300 stars and about 790 forks, and is not archived.

Crawler operators get a lecture on politeness

One section of the README is addressed directly to what it calls Crawl Operators, and its tone is closer to a code of conduct than to a feature list. Heritrix is designed to respect robots.txt exclusion directives and META nofollow tags. The README asks you to consider the load your crawl will place on seed sites and set politeness policies accordingly.

The second instruction is the one that gets skipped and should not be: always identify your crawl with contact information in the `User-Agent` so that sites which may be adversely affected by your crawl can contact you or adapt their server behaviour. In other words, the project treats being a good citizen on someone else's server as an operational requirement, not a courtesy, and it asks you to configure it.

That framing is a reasonable indicator of how the rest of the tool is built. Politeness policies are configuration, which means they belong to the same layer as scope, filters and revisit intervals rather than being hardcoded. If you are evaluating crawlers for a large or hostile corpus, this is a useful signal, because the operator burden is visible in the documentation rather than discovered in production.

Maven artifacts, a distribution zip, and a Docker image

Heritrix ships three ways, and the split matters when you script a deployment. Individual modules are published to Maven Central under the `org.archive.heritrix` group, with a version badge at the top of the README. The JavaDoc is published per module, which tells you the artifact layout: `heritrix-engine`, `heritrix-modules`, `heritrix-commons` and `heritrix-contrib` are documented separately, so the engine is separate from the crawler behaviors you select.

A Docker image exists on Docker Hub under the `iipc` organisation, with a version badge in the README and per-release layers linked from each release page. That is the shortest path for someone who wants to run a crawl without managing a JVM.

Third is the distribution archive, a zip or tar.gz containing the `lib` directory of bundled dependencies. The README explains why that archive exists: Heritrix is distributed with the libraries it depends upon, found under `lib` in the release distribution, used under their respective licenses, which are included alongside them. Since version 3.16.0 those archives are only on the GitHub release page, because recently introduced size limits at Maven Central prevent publishing the `-dist.zip` and `-dist.tar.gz` packages. Module binaries, Javadoc and source JARs still go to Maven Central. If you have automation pinned to Maven Central for the full distribution, that is a breaking change to absorb.

The repository layout reflects a modular crawler core

The tree at the repository root is short and tells you where the work happens. `engine/` is the crawler engine. `modules/` holds the behaviors that configure what a crawl does. `commons/` is shared code. `contrib/` is where outside contributions sit. Alongside those sit `docker/` for the image definition, `docs/` for the Read the Docs source, `docgen/`, and a `dist/` directory that is presumably where the assembled distribution is produced.

`pom.xml` at the root confirms a Maven multi-module build rather than a Gradle one, which matches the Maven Central publishing described above.

The documentation set is split into operator-facing and developer-facing halves, and the split is useful when deciding what you need to read. Operator documentation covers Getting Started, Operating Heritrix, Configuring Crawl Jobs and a Bean Reference, plus a GitHub wiki. Developer documentation is a separate developer manual hosted at crawler.archive.org, a REST API page, and the per-module JavaDoc. If you are scripting crawls rather than running them by hand, the REST API page and the developer manual are the two that matter, and neither is reachable from the JavaDoc index.

Three releases in ten weeks, mostly about reporting honestly

The recent release history is worth reading closely because it shows what the maintainers think is broken in web crawling today.

Version 3.17.0, published 2026-08-25, added an `includeResCode` property to `SourceTagsReport` so reports can include HTTP response codes, and extended `ResponseCodeReport` to include Heritrix's own negative error codes rather than only HTTP statuses. The same release added `BotBlockDetector`, a processor that detects responses from common bot-blocking services including Akamai, Anubis, Cloudflare, DataDome and Incapsula, then annotates affected URIs with a `botblock:<service>` tag.

That detector is the single most interesting item in the batch. A crawler that silently receives a challenge page records a 200 response and a short HTML body, and a WARC built from that is not an archive of the page at all. Annotating the URI so a later operator can tell a bot wall from a thin page is exactly the distinction an archival workflow needs.

Version 3.17.1, published 2026-09-10, made SSL handshake failures use a dedicated `S_SSL_ERROR` fetch status with value -9 instead of `S_CONNECT_FAILED`, which reduces unlikely retries, and extended `SourceTagsReport` to include Heritrix status codes for failed, retried and disregarded URIs. It also bumped freemarker from 2.3.34 to 2.3.35 and groovy-bom from 5.1.0 to 5.1.1. Version 3.16.0, published 2026-07-03, added `PaginationBehavior`, a `BrowserProcessor` behavior that repeatedly clicks the next-page link and extracts links so client-side paginated sites can be crawled.

Licensing is Apache 2.0 with a file-level caveat

The README states that Heritrix is free software that you can redistribute or modify under the terms of the Apache License, Version 2.0, and the badge at the top of the page says the same. There is a caveat worth carrying into any compliance review, though: some individual source code files are subject to or offered under other licenses, and the README directs you to the included `LICENSE` file for more information rather than claiming the whole tree is uniformly Apache 2.0.

That is a normal and honest formulation for a project that has absorbed contributed code, and it means the safe process is to read the `LICENSE` file rather than assume a blanket grant. The GitHub metadata for the repository records no standard license identifier, so if your tooling reads the API field rather than the text, it will tell you nothing useful.

On activity: the last push to this repository is 2026-09-22, and three tagged releases landed between July and September 2026, so this is a project receiving work on a normal cadence. There are 36 open issues and a `SECURITY.md` alongside `CHANGELOG.md` and `RELEASING.md` at the root, which is a mature repository convention set.

Editorial conclusion

Heritrix is the crawler to reach for when the output has to be defensible years later, because WARC plus a crawl log plus a configuration file is a record you can replay, and most general purpose crawlers are not built that way. What the recent releases show is a project in maintenance mode with a specific agenda: better failure reporting, honest bot-block detection, and a packaging migration forced by Maven Central size limits. The three changes worth knowing about are the GitHub-hosted distribution archives from 3.16.0, the `BotBlockDetector` processor in 3.17.0, and the SSL handshake status split in 3.17.1, which stops a TLS negotiation failure from being retried as if it were a refused connection. Start with the Getting Started page, then Configuring Crawl Jobs and the Bean Reference, since the crawler is configured as a graph of beans and that is the concept you need before any job runs.

Frequently asked questions

What is Heritrix used for?

It is the Internet Archive's open-source web crawler, aimed at collecting WARC records for long term preservation rather than at search indexing or scraping. The README frames the goal as preserving the digital artifacts of culture for future researchers and generations, which is why it favours completeness and honest failure reporting over raw throughput.

Is Heritrix free and what license does it use?

Yes. The README states it is free software you can redistribute and modify under the Apache License, Version 2.0, and ships a LICENSE file. It also notes that some individual source files are subject to or offered under other licenses, so the LICENSE file is the authoritative document rather than a blanket grant across the tree.

How do I install or run Heritrix?

Individual modules are on Maven Central under org.archive.heritrix, and a Docker image is published under the iipc organisation on Docker Hub, which is the shortest route for running a crawl without managing a JVM. The full distribution zip or tar.gz has been hosted on the GitHub release page since version 3.16.0, because new Maven Central size limits prevent publishing it there.

How does Heritrix handle sites that block crawlers?

Since version 3.17.0 a processor called BotBlockDetector recognises responses from Akamai, Anubis, Cloudflare, DataDome and Incapsula, and annotates affected URIs with a botblock tag naming the service. That matters for archival work because a challenge page returns a normal status code and would otherwise be recorded in the WARC as if it were the real page.

What is the newest version of Heritrix?

Version 3.17.1, published on 2026-09-10, preceded 3.17.0 on 2026-08-25 and 3.16.0 on 2026-07-03. The repository itself was last pushed on 2026-09-22. The wiki carries a latest releases section, and each release page links its changelog, JavaDoc, Maven Central entry and Docker Hub layers.

How do I configure a crawl job?

Heritrix is configured as a graph of beans rather than through a single options file, so the Bean Reference in the documentation is the page to read first. Operator documentation is split into Getting Started, Operating Heritrix and Configuring Crawl Jobs, with the REST API page and developer manual covering scripted operation.

Official sources

  1. internetarchive/heritrix3 on GitHub
  2. Issues
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/internetarchive-heritrix3.svg)](https://hysenlabs.com/projects/internetarchive-heritrix3)