Open-source project
apache/nutch avatar
apache/nutch

Apache Nutch: a Hadoop-based crawler you build with Ant

Apache Nutch is an extensible and scalable web crawler

3,293 stars1,281 forksJavaApache-2.0

At a glance

What is it?
Apache Nutch is an extensible web crawler written in Java and built on the Hadoop batch model. This review covers how its crawl cycle works, how to build and run it from the repository, and where its batch design stops fitting.
Who is it for?
Adopt Nutch if you need a crawl that runs as a batch job over Hadoop and you are willing to maintain a Java build with Ant, Ivy, and a configured nutch-site.xml. Do not adopt it if you need a single-process crawler you can start in one command, or if you cannot run the Hadoop components it depends on.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 1 day ago.
What is it written in?
Mainly Java, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Nutch solves, and who it is actually for

Nutch is a crawler for people who want the crawl itself to be a batch job. The repository describes it as an extensible and scalable web crawler, written in Java, with Hadoop among its topics. That combination is the whole proposition: fetch pages, parse them, follow links, and write the results out as part of a Hadoop-style pipeline rather than as a long-running service you babysit.

The audience follows from that. If you already run Hadoop and you need a crawl stage that fits next to other batch stages, Nutch is aimed at you. The README is written for contributors and for developers wiring the project into an IDE, not for someone who wants a hosted crawler. It links to the Nutch Tutorial on the project wiki for getting started, which is a signal about where the project expects the introductory material to live.

It is a poor fit for small jobs. A crawler that needs a Java toolchain, an Ant build, Ivy dependency resolution, and a configured nutch-site.xml before it will run is not the tool for fetching twenty pages on a laptop. The project is also not a scraping framework: it does not hand you a selector API for extracting typed fields from pages. It hands you a crawl cycle and a plugin system.

The crawl cycle: inject, fetch, parse, index

Nutch's architecture is visible in the class names the README names. The Eclipse instructions tell you to create a Java Application configuration and choose org.apache.nutch.crawl.Injector, with two arguments: first the crawldb directory, second the URL directory the injector reads from. That is the entry point of the cycle. URLs go into a crawldb, and the crawl proceeds in segments.

The README names two more classes in passing. org.apache.nutch.indexer.IndexingJob is given as an example of a class with a main function, and the indexing instructions mention a -deleteGone argument alongside crawldb and segment paths. So the shape is: inject seed URLs into the crawldb, fetch them into segments, parse and score what came back, then run an indexing job that takes the crawldb and the segment paths and can drop records for pages that have gone away.

Everything that touches the network, the parser, or the output format sits behind the plugin system, which is why plugin.folders matters so much. The README states that plugin.folders normally points to <project_root>/build/plugins. Plugins are not pulled from a registry at runtime; they are built and loaded from a directory on disk. That is a deliberate design choice, and it is also the single most common thing to get wrong.

Building Nutch from the repository

The README does not give a plain command-line build for running a crawl. It gives the contributor path and the IDE path. The contributor instructions start with cloning the repository and creating a branch named after a JIRA issue id of the form NUTCH-xxxx.

bash
git clone https://github.com/apache/nutch.git
cd nutch
git checkout -b NUTCH-xxxx

The branch name is not decoration. The commit message convention in the README is "fix for NUTCH-xxx contributed by <your username>", so the issue id is expected to appear in both places. For anyone who only wants to run a crawl, this tells you the project is organized around patch contribution, and that the runnable instructions live on the wiki tutorial rather than in the README.

The build system is Ant, and the README gives one command for generating IDE project files:

bash
ant eclipse

That produces the .classpath and .project files. Two details matter here. First, the README warns that you must manually trigger a build through Ant to pick up changes when running in IntelliJ IDEA, because the Ant build and the IDE build are separate. Second, the README notes that if you see "No plugins found on paths of property plugin.folders=plugins", the quick fix of editing plugin.folders in nutch-default.xml should not be used. That is an unusually direct warning, and it tells you the correct fix is to make the build produce plugins where the configuration expects them.

Configuring nutch-site.xml before the first run

The README is explicit that you must configure nutch-site.xml before running, and that two properties have to be present: http.agent.name and plugin.folders. The first is the crawler's user agent identity; the second is where plugins are loaded from. The README states plugin.folders normally points to <project_root>/build/plugins.

The Eclipse route to a first run is: run ant eclipse, configure those two properties, create a Java Application configuration, set the main class to org.apache.nutch.crawl.Injector, and pass the crawldb directory and the URL directory as arguments in that order. The README does not document what a seed URL file should contain, nor the exact output of the injector. It also does not document rollback if a run goes wrong. Those gaps are real, and they are why the wiki tutorial is the place to start rather than the README.

For IntelliJ IDEA the README requires the IvyIDEA plugin before running ant eclipse, then importing the project from existing sources as an Eclipse project. The project SDK is Java 17 or newer. On macOS with a Homebrew OpenJDK, the README says to use the directory under libexec: <openjdk_directory>/libexec/openjdk.jdk/Contents/Home. That is a specific, easy-to-miss path detail, and it is the kind of thing that costs an afternoon if you skip it.

Running Nutch from an IDE, and the Ant rebuild trap

Once the project is imported, the README's IntelliJ instructions for running are more involved than they first look. You create an Application configuration, set the main class to something with a main function such as org.apache.nutch.indexer.IndexingJob, and set the program arguments to what you would pass on the command line. The README suggests getting those arguments by running the crawl executable for your job, and using fully qualified paths for the crawldb and segment paths, plus -deleteGone where appropriate.

The working directory should be your nutch runtime/local path. Then you modify the classpath to add the config directory for that working directory, for example runtime/local/conf, and add VM options matching the crawl executable, with -Xmx4096m and a -Dhadoop.log.dir setting given as examples. Every one of those pieces has to line up: main class, arguments, working directory, config directory, VM options.

The trap is stated plainly: you must manually trigger an Ant build to get the latest changes when running, because the Ant build system is separate from the IDE one. Editing a plugin in IntelliJ and pressing run will not necessarily exercise the code you just changed. That is a workflow cost that recurs on every iteration, and it is the strongest argument for doing day-to-day crawling from the command line and reserving the IDE for debugging specific classes.

Where Nutch is the wrong tool

The batch model is the limitation. A Nutch crawl is a pipeline of jobs over a crawldb and segments, and the README's own examples treat it that way: an injector run, then a crawl, then an indexing job that takes the crawldb and segment paths and can delete records for pages that have gone. There is no documented single-command "crawl this site and give me the pages" flow in the README.

That has consequences. If you need pages immediately after fetching them, or you need to react to a page's content mid-crawl, the inject-fetch-parse-index cycle is the wrong shape. If you have no Hadoop components and no intention of running them, the setup cost is hard to justify for a small crawl. And if your extraction needs are the hard part rather than the fetching, Nutch gives you a plugin system and a Java codebase, not a scripting surface.

The plugin loading model compounds this. Plugins come from a directory on disk that the build is expected to populate. The README's warning against patching plugin.folders in nutch-default.xml implies that misconfigured plugin paths are common enough to warrant a callout. A crawler whose network behavior, parsing, and scoring all live in plugins is powerful, but it means a broken plugin path can look like a broken crawl.

How Nutch differs from Scrapy and StormCrawler

Scrapy is a Python scraping framework. The difference is not the language, it is the unit of work. Scrapy gives you a spider class with callbacks and an item pipeline, and you run it as a process. Nutch gives you a crawl cycle that runs as batch jobs over a crawldb and segments, and you run it through Hadoop. If your team writes Python and wants to iterate on extraction logic quickly, Scrapy's model is a better match. If your crawl has to be one stage in a larger Hadoop batch, Nutch's model is the one that fits.

StormCrawler is the closer comparison, because it also addresses crawling at scale, but it is built around stream processing rather than the batch cycle Nutch uses. That changes the operational shape: streaming crawlers keep state in a running topology, while Nutch materializes crawldb and segment data on disk between stages. Which is better depends on whether you want a long-running topology or a job you schedule and inspect. The README does not make this comparison, and it does not benchmark against either project.

One practical difference favors Nutch for regulated environments: it is an Apache project under the Apache License 2.0, with LICENSE.txt, NOTICE.txt, and a THREAT_MODEL.md at the top level of the repository. The presence of a threat model document is not something every crawler ships.

Maintenance, licence, and what a fork costs

The repository is not archived, and the last push was on 2026-09-23. The README shows active build infrastructure: a GitHub Actions workflow for the master branch, a Jenkins smoke test job described as a single-node Hadoop cluster test, and SonarCloud quality gate badges. CI uses Java 17. The project also runs Apache Yetus test-patch on pull requests for style and reporting checks, with baselines under .yetus/ such as .yetus/excludes.txt and .yetus/detsecrets-ignored-hashes.txt.

That infrastructure is a maintenance signal in both directions. It means there is a working build and a smoke test you can lean on. It also means contributing a change requires passing Yetus, following a two-space indent code format defined by eclipse-codeformat.xml, and filing a JIRA issue first. The README describes the local test-patch invocation with --build-tool=nobuild and a plugin list that excludes jira, gitlab, unit, and compile. For a team running an internal fork, that is the cost of staying close to upstream: every patch you carry has to survive the same checks.

The licence is Apache-2.0. The repository carries LICENSE.txt and NOTICE.txt along with LICENSE-binary and NOTICE-binary, which indicates binary distributions have their own notices. If you redistribute a Nutch build, those binary notice files are the ones to read. This is not legal advice; confirm obligations with your own counsel.

Editorial conclusion

Adopt Nutch if you need a crawl that runs as a batch job over Hadoop and you are willing to maintain a Java build with Ant, Ivy, and a configured nutch-site.xml. Do not adopt it if you need a single-process crawler you can start in one command, or if you cannot run the Hadoop components it depends on. Before committing, verify two things: that your nutch-site.xml sets http.agent.name and plugin.folders, and that the plugin.folders path resolves to <project_root>/build/plugins after an Ant build. The README points to the Nutch Tutorial on the project wiki for getting started, and that tutorial, not the README, is where the end-to-end crawl walkthrough lives.

Frequently asked questions

What is Apache Nutch?

Apache Nutch is an extensible and scalable web crawler written in Java, with Hadoop among its listed topics. It is licensed under Apache-2.0 and its homepage is nutch.apache.org.

How do you use Apache Nutch?

You build it with Ant, configure nutch-site.xml with http.agent.name and plugin.folders, then run the crawl cycle starting from org.apache.nutch.crawl.Injector with a crawldb directory and a URL directory as arguments. The README points to the Nutch Tutorial on the project wiki for getting started.

How does Apache Nutch compare with Scrapy?

Scrapy is a Python scraping framework you run as a process, while Nutch runs its crawl as a batch cycle over a crawldb and segments on Hadoop. Nutch's README does not compare the two or benchmark against Scrapy.

How does Apache Nutch compare with StormCrawler?

StormCrawler is built around stream processing, while Nutch's cycle materializes crawldb and segment data on disk between batch stages. The Nutch README does not discuss StormCrawler or make performance claims about it.

What are the alternatives to Apache Nutch?

Scrapy and StormCrawler are the two alternatives that come up in searches. The practical difference is the execution model: a Python spider process, a streaming topology, or Nutch's Hadoop batch cycle.

Official sources

  1. apache/nutch on GitHub
  2. Issues
  3. License: Apache-2.0
  4. Project website
  5. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/apache-nutch.svg)](https://hysenlabs.com/projects/apache-nutch)