Framework
gocolly/colly avatar
gocolly/colly

gocolly/colly: a Go scraper framework for people who want callbacks, not pipelines

Elegant Scraper and Crawler Framework for Golang

25,536 stars1,862 forksGoApache-2.0

At a glance

What is it?
Colly is a Go library for writing crawlers and scrapers through an event-callback API. It suits Go services that need to fetch and parse pages in-process, and it is a poor fit for anyone who wants a declarative, config-driven crawler.
Who is it for?
Adopt Colly if your team already ships Go and the crawler is one component of a larger service, because the callback API keeps fetching, parsing and storage in one binary. Do not adopt it if you want a declarative spider format, a scheduler or a scraping UI, since Colly provides none of those.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 2 days ago.
What is it written in?
Mainly Go, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What problem Colly solves, and for whom

Colly gives Go developers a collector object that fetches URLs and dispatches the results to callbacks. The README describes it as a clean interface for writing "any kind of crawler/scraper/spider", with structured extraction aimed at data mining, data processing or archiving. The intended reader is a Go programmer who wants crawling inside an existing program rather than as a separate process. That matters because the usual alternative in other languages is a framework you configure and run; Colly is a package you import. The README's own example is a complete program: create a collector, register an OnHTML handler for a[href], call Visit. There is no spider class, no settings file and no scheduler. If the crawler needs to share a database connection, a logger or a context with the rest of your service, that fits. If the crawling logic is something a non-Go colleague should be able to edit, it does not.

How the collector dispatches requests and callbacks

The mechanism is a callback registry on a Collector. The README example registers OnHTML with a CSS selector and OnRequest, then calls Visit with a starting URL. Each handler receives a typed argument: OnRequest gets a *colly.Request, OnHTML gets a *colly.HTMLElement. New URLs are queued by calling e.Request.Visit from inside a handler, which is how the sample walks every link on a page. The repository layout shows the pieces this rests on: colly.go holds the collector, request.go and response.go the request and response types, htmlelement.go the HTML element passed to handlers, and xmlelement.go an XML equivalent. Parsing is not written from scratch. go.mod lists github.com/PuerkitoBio/goquery, github.com/antchfx/htmlquery and github.com/antchfx/xmlquery, so selectors and XPath come from those libraries. Robots.txt handling comes from github.com/temoto/robotstxt, character-set detection from github.com/saintfish/chardet, and URL parsing from github.com/nlnwa/whatwg-url. The queue/ and storage/ directories hold the interfaces behind the in-memory defaults, and extensions/ holds optional add-ons. The README also lists what the collector manages for you: request delays and maximum concurrency per domain, cookies and sessions, caching, robots.txt, and configuration via environment variables. Those are properties of the collector, not of your handlers, which is why the callback style stays short.

Installing Colly and scraping one page

Installation is a single go get against the v2 module path. The README states it plainly, and go.mod confirms the module is github.com/gocolly/colly/v2 with go 1.24.0 and toolchain go1.24.9, so a Go toolchain of that generation is what the module expects.

bash
go get github.com/gocolly/colly/v2

A first program follows the README example. It creates a collector, prints each URL as it is requested, and follows every link it finds.

go
c := colly.NewCollector()

c.OnHTML("a[href]", func(e *colly.HTMLElement) {
	e.Request.Visit(e.Attr("href"))
})

c.OnRequest(func(r *colly.Request) {
	fmt.Println("Visiting", r.URL)
})

c.Visit("http://go-colly.org/")

Running it prints one Visiting line per request, starting with the seed URL. Note what the sample does not do: it never checks the error returned by Visit, and it sets no domain restriction, so link following is unbounded. The README points to the _examples directory for more detailed examples, and that directory is where you should look before writing anything that hits a live site. For a single page rather than a crawl, drop the OnHTML handler and read the response in OnResponse instead.

Where Colly gets in your way

The callback style has a cost. State that lives across pages has to be threaded through closures or a shared struct, and the collector gives you no built-in place to put it. The README lists caching and distributed scraping as features but does not document a rollback story, a resume-from-checkpoint story or a deduplication guarantee beyond what the queue interface implies; the queue/ directory is where you would have to look. Rate limiting is per-domain, which is what you want for politeness but not what you want if a single site has to be crawled from many machines without coordination. The release history is the other constraint. v2.0.0 landed in November 2019, v2.1.0 in June 2020, and v2.2.0 in March 2025, so the major line has been stable for years with long gaps between tagged releases. The last push to the repository was on 2026-09-16, so development has not stopped, but anyone who needs a predictable release cadence should read CHANGELOG.md and the commit history rather than assuming one. Colly is also the wrong tool when the target is a JavaScript-rendered page, because it fetches HTML over HTTP and the README describes no browser engine; and it is the wrong tool when the task is a one-off fetch of a single URL, where net/http plus goquery is less machinery.

Colly versus goquery, and versus Scrapy

The comparison that comes up most is Colly against goquery, and it is not really a contest because Colly depends on goquery. goquery wraps net/http and gives you a jQuery-like selection API over a parsed document. It has no collector, no queue, no per-domain delay, no robots.txt handling and no session management. Colly adds those and keeps goquery underneath for the selector work. If you are fetching one page and pulling three fields out of it, goquery alone is the smaller dependency. The other comparison is Scrapy. Scrapy is a Python framework with a project layout, a spider class, a settings module, a scheduler and a command-line runner. Colly has none of that: no spider format, no settings file, no separate process. That difference cuts both ways. Scrapy gives a team a shared structure and a runner they can point at a project directory; Colly gives a Go program a library, which means the deployment, the logging and the data sink are whatever the surrounding service already uses. Choosing between them is mostly choosing the language and the deployment shape, not the feature list.

Licence and the cost of staying on v2

Colly is licensed under Apache-2.0, and the repository carries LICENSE.txt at the top level. Apache-2.0 is a permissive licence with an explicit patent grant and a requirement to preserve notices, which matters if you vendor the source or ship a binary that includes it. This is a description of the licence file, not legal advice; if the crawler is part of a distributed product, have counsel read LICENSE.txt and the notices of the dependencies in go.mod, several of which carry their own terms. The practical upgrade cost is low but not zero. The module path is versioned as /v2, so v1 and v2 can coexist in one build graph, and go.mod even lists github.com/gocolly/colly v1.2.0 as a dependency of v2 itself. Moving from v1 to v2 is a path change in imports plus whatever the v2 release notes describe. Moving between v2 minor versions is the cheaper operation, and the CHANGELOG.md file at the repository root is the place to check before bumping. Because the minor releases are years apart, pinning an exact version and reading its changelog entry is more useful than tracking master.

Editorial conclusion

Adopt Colly if your team already ships Go and the crawler is one component of a larger service, because the callback API keeps fetching, parsing and storage in one binary. Do not adopt it if you want a declarative spider format, a scheduler or a scraping UI, since Colly provides none of those. Before writing production code, read the _examples directory and the CHANGELOG.md entry for v2.2.0, then confirm how OnError and OnResponse behave in the version you pin.

Frequently asked questions

How do you use gocolly/colly to crawl a site?

Create a collector with colly.NewCollector(), register handlers such as OnHTML and OnRequest, then call Visit with a starting URL. Inside an OnHTML handler you can call e.Request.Visit on a link to queue the next page, which is how the README example follows every a[href] it finds.

What is gocolly/colly?

It is a scraping and crawling framework for Go, described in its README as providing a clean interface for writing any kind of crawler, scraper or spider. It is distributed as an importable package rather than a standalone application.

How does gocolly/colly compare with Scrapy?

Scrapy is a Python framework with its own spider format, settings module and command-line runner. Colly is a Go library with no spider format and no separate runner, so the surrounding Go program supplies the deployment, logging and storage.

How does gocolly/colly compare with goquery?

Colly depends on goquery, which is listed in the module's go.mod, and uses it for document selection. goquery alone gives you HTTP fetching plus a selection API, while Colly adds the collector, queue, per-domain delays, cookies, sessions and robots.txt handling.

Official sources

  1. gocolly/colly on GitHub
  2. License: Apache-2.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/gocolly-colly.svg)](https://hysenlabs.com/projects/gocolly-colly)