Open-source project
andeya/pholcus avatar
andeya/pholcus

Pholcus: Distributed web crawling in Go with three downloader engines

Pholcus is a distributed high-concurrency crawler software written in pure golang

7,585 stars1,665 forksGoApache-2.0

At a glance

What is it?
A Go framework for building web crawlers that run standalone, across a cluster, or as a client receiving tasks. It offers multiple interfaces and flexible rule authoring.
Who is it for?
Pholcus suits teams building large-scale crawlers where concurrency and rule flexibility matter. The choice between static and dynamic rules trades off startup overhead for ease of updates.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 37 days ago.
What is it written in?
Mainly Go, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

A distributed crawler framework with flexible rule authoring

Pholcus solves the problem of collecting data from many websites at once, whether on a single machine or spread across a cluster. It is written in pure Go, which allows it to start tens of thousands of concurrent HTTP requests without spawning threads. The README lists 30+ built-in crawling rules for sites like Baidu, JD.com, Taobao, and Zhihu. Engineers who need a web scraper that can be deployed as a service and reconfigured without recompiling will find it useful. The software is under Apache-2.0 license. The last push was on 2026-08-24, showing active maintenance. The framework supports three operational modes: single-machine for standalone use, server mode for distributing tasks across a network, and client mode for receiving tasks from a server. The approach trades off simplicity for power; instead of a drag-and-drop interface, you write rules directly in Go or XML, gaining full control over crawling behavior.

How Pholcus distributes crawling tasks

Pholcus works in three modes: single-machine mode runs on one computer, server mode distributes tasks across a network, and client mode runs on machines that accept tasks from a server. All three modes use the same core engine, defined in app/crawler/. The scheduler in app/scheduler/ decides which URLs to fetch next and handles deduplication and retry logic to avoid re-crawling the same page. The spider rule engine in app/spider/ extracts data from responses and filters unwanted content. Data flows through app/pipeline/ to multiple output backends: MySQL, MongoDB, Kafka, Beanstalkd, CSV, Excel, or raw file download. The distributed communication layer in app/distribute/ uses a full-duplex socket framework to coordinate between master and slave nodes, allowing work to be split across a cluster. Request records persist automatically, which means you can resume a crawl from where it stopped if the process crashes or is paused. This distributed architecture makes Pholcus suitable for large-scale operations where no single machine can handle the load. The modular design lets you swap output backends without touching the crawling logic.

Installing and running your first crawl

Get the source and build Pholcus with Go 1.18 or later. The README recommends Go 1.22 or above for best compatibility.

bash
git clone https://github.com/andeya/pholcus.git
cd pholcus

Create a main.go file that imports the rule library and calls exec.DefaultRun. The sample/main.go shows the pattern:

go
package main

import (
    "github.com/andeya/pholcus/exec"
    _ "github.com/andeya/pholcus/sample/static_rules"
)

func main() {
    exec.DefaultRun("web")
}

Compile the binary:

bash
go build -o pholcus ./sample/

Run pholcus and access the web interface at http://localhost:2015 to select a spider, set parameters, and start the crawl. On Windows, hide the console window with a build flag:

bash
go build -ldflags="-H=windowsgui -linkmode=internal" -o pholcus.exe ./sample/

The web UI lets you choose from the 30+ built-in rules, configure parameters like concurrency and pause time, and monitor progress in real time. Alternatively, use the GUI (Windows only) or the command-line interface for batch crawling in cron jobs. The framework automatically opens the browser after startup if you use web mode.

Three download engines: Surf, PhantomJS, and Chrome

Pholcus lets you choose which engine fetches each URL. The Surf engine uses pure Go HTTP, handling high concurrency but not executing JavaScript. PhantomJS is a headless browser that can run JavaScript but has lower concurrency and is no longer maintained. The Chrome engine uses chromedp to drive Chromium or Google Chrome, executing JavaScript and handling security redirects, making it suitable for sites that load content client-side. In static rules (Go code), specify the engine with the DownloaderID field: pass 0 for Surf (the default), 1 for PhantomJS, or 2 for Chrome. In dynamic rules (XML), set the same field as a number. Chrome requires that you have Chromium or Google Chrome installed on the machine running the crawler. The README documents each engine's use case: Surf for static pages, Chrome for SPA sites and security-protected content.

Static rules versus dynamic rules and when each applies

Static rules are Go code compiled into the binary, offering maximum performance and control. You write rules in sample/static_rules/, then rebuild. Dynamic rules are XML files stored in dyn_rules/, loaded at startup without recompiling. The README shows that dynamic rules use Script tags with JavaScript to define URL queues and extraction logic, while supporting an older .pholcus.html format as well. The tradeoff is clear: static rules are faster and less flexible; dynamic rules let you add crawlers without rebuilding, but at the cost of startup overhead and needing to understand the XML and JavaScript template syntax. The repository includes 30+ static rule examples covering major Chinese e-commerce and search sites. This dual approach lets you choose: performance for heavy crawling, or agility for quick rule updates.

Cookie management, proxy rotation, and throttling

Pholcus manages cookies automatically: it keeps a fixed User-Agent and saves cookies across requests, or it rotates the User-Agent and disables cookies per request. You can define an HTTP proxy pool, and the crawler will rotate IPs at a frequency you specify. The app/aid/ module handles historical records and proxy IP tracking. A built-in random pause mechanism mimics human behavior. You control the concurrency level (number of goroutines), the batch size for each round, and the crawl limit (total URLs to fetch). The scheduler deduplicates requests automatically and retries failures. Success and failure records are persisted, allowing you to resume a crawl from where it stopped. The command-line mode supports flags like -a_thread (concurrency), -a_batchcap (batch size), -a_pause (pause time), and -a_limit (crawl limit). These controls make Pholcus suitable for large-scale crawling where you need to respect server load and avoid detection.

Surf cannot execute JavaScript, Chrome adds overhead, and documentation is in Chinese

The Surf downloader cannot execute JavaScript, so any site that renders content client-side will not work. The Chrome engine adds complexity because it requires a running browser process and increases memory use. The documentation in the README is in Chinese, and the software itself gives no English documentation for rule authoring. The command line interface accepts many flags; the README does not explain what each one does in English, so you will need to read the source or rely on the -h output. PhantomJS is deprecated, which means it should not be used for new crawlers. The project's last push was on 2026-08-24, about a month ago, so maintenance is active but infrequent. The sample rules and configuration examples are all in Chinese.

Comparing Pholcus to Colly and Scrapy

Colly is a Go scraping library with simpler APIs; Pholcus is a full-featured framework including multiple interfaces (Web, GUI, CLI), distributed execution, and multiple output backends. Colly is lower-level and less opinionated, focusing on the HTTP layer. Scrapy is a Python framework with rich middleware, pluggable components, and strong documentation. Pholcus has no Python binding and no equivalent middleware system; you write rules directly instead. Scrapy is more mature and has a larger ecosystem, but Pholcus compiles to a single binary and runs without Python installed. Pholcus's distributed mode is more complete than either Colly or Scrapy, offering server and client components out of the box. For teams building bespoke crawlers where distribution matters and Python is not available, Pholcus offers the most integrated solution.

Editorial conclusion

Pholcus suits teams building large-scale crawlers where concurrency and rule flexibility matter. The choice between static and dynamic rules trades off startup overhead for ease of updates. Before committing to it, check whether your target sites work with the Surf HTTP downloader or require the Chrome engine, since Chrome adds operational complexity. Also verify that the command line flags and rule syntax in the samples match your build version.

Frequently asked questions

Can Pholcus crawl JavaScript-heavy websites?

Only if you use the Chrome downloader engine (DownloaderID: 2), which drives Chromium and executes JavaScript. The default Surf engine is pure HTTP and cannot execute JavaScript.

Can I write crawl rules without recompiling the binary?

Yes. Dynamic rules are XML files placed in dyn_rules/ and loaded at startup. Static rules (Go code) require a rebuild but offer better performance.

What does the Surf downloader do?

Surf is a pure Go HTTP client that handles high concurrency, automatic cookie and User-Agent management, and proxy rotation. It cannot execute JavaScript.

Can Pholcus run as a distributed crawler?

Yes. Pholcus has server mode and client mode. The server distributes tasks to clients over a socket framework, allowing crawling to be spread across multiple machines.

Official sources

  1. andeya/pholcus on GitHub
  2. Issues
  3. License: Apache-2.0
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/andeya-pholcus.svg)](https://hysenlabs.com/projects/andeya-pholcus)