Framework
geziyor/geziyor avatar
geziyor/geziyor

Geziyor: a Go web crawling framework with JS rendering and built-in exporters

Geziyor, blazing fast web crawling & scraping framework for Go. Supports JS rendering.

2,778 stars155 forksGoMPL-2.0

At a glance

What is it?
Geziyor is a Go library for concurrent crawling and structured extraction, with optional Chrome rendering, caching and Prometheus metrics. It suits Go teams that want a crawler inside their own binary rather than a separate service.
Who is it for?
Adopt Geziyor if your team already writes Go and wants crawling, extraction and export inside one binary, with per-domain concurrency limits and a Prometheus endpoint you control. Do not adopt it if you need a scheduler, a dashboard or a retry queue out of the box, or if you cannot guarantee a Chrome binary on the host when pages require rendering.
Can I use it commercially?
Yes, with conditions. MPL-2.0 is a weak copyleft licence: you can use it inside commercial and closed-source software, but if you distribute changes to its own files, you must publish those changes under the same licence.
Is it still maintained?
Yes. The repository last received commits 90 days ago.
What is it written in?
Mainly Go, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 24, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Geziyor solves, and who ends up using it

Most scraping stacks split into two halves: a crawler process that fetches pages and a script that parses them. Geziyor collapses both into a Go library. You declare start URLs, a parse callback and an exporter, then call Start(). The crawl loop, the HTTP client, the HTML parser and the output writer all live in the same process you compile.

The README frames the intended uses as data mining, monitoring and automated testing. That list matters, because it tells you the project assumes a developer who is comfortable writing Go callbacks rather than configuring a YAML pipeline. If you want a spider definition in a config file and a web UI to watch it, this is not that. If you want to call a crawler from inside an existing Go service, or ship a single static binary that scrapes a supplier's catalogue on a cron, the shape fits.

The dependency list in go.mod confirms the design. goquery supplies the jQuery-style selector API, chromedp drives Chrome, goleveldb and diskv back the cache, prometheus/client_golang provides metrics, and temoto/robotstxt handles robots.txt parsing. Nothing in that list is a crawler framework. Geziyor is the orchestration layer over general-purpose libraries.

How the crawl loop, callbacks and exporters fit together

The mechanism is callback-driven. Initial requests come from StartURLs, or from StartRequestsFunc if you want to construct them by hand. Geziyor makes concurrent requests to those URLs. When a response arrives, ParseFunc runs with the crawler and the response. Inside ParseFunc you either extract data, or enqueue more URLs with g.Get, g.Head or g.Do, passing the same callback or a different one.

That recursion is the whole crawl. There is no separate frontier object to manage. Pagination in the README example is a single line: find the next link, call g.Get with the callback again, and the loop continues until no new links are found. The trade-off is that crawl state lives in your callback closures, so resuming an interrupted crawl means persisting that state yourself. The README does not document a resume mechanism.

Extraction reads from r.HTMLDoc, which is a goquery Document. The README is explicit that HTMLDoc is nil when the response is not HTML, so a parse function that touches it unconditionally will panic on a JSON or image response. Guarding on content type is your responsibility, not the framework's.

Output goes through a channel. You send maps or structs to g.Exports, and each configured Exporter consumes them. The README lists JSON, CSV and custom exporters, with the exporter package linked from godoc. Because export is channel-based, the writer runs concurrently with the crawl rather than buffering everything until the end.

Installing Geziyor and running a first crawl

The README recommends Go modules and gives a single install command. Run it inside a module directory:

bash
go get -u github.com/geziyor/geziyor

If you intend to use GetRendered, the README states you must have Chrome installed. Geziyor launches local Chrome through its CLI by default, or connects to an existing instance if you set the BrowserEndpoint option, for example "ws://localhost:3000".

A minimal crawl that prints response bodies looks like this. StartURLs seeds the queue and ParseFunc receives every response:

go
geziyor.NewGeziyor(&geziyor.Options{
    StartURLs: []string{"http://api.ipify.org"},
    ParseFunc: func(g *geziyor.Geziyor, r *client.Response) {
        fmt.Println(string(r.Body))
    },
}).Start()

To extract structured records and write them to a file, send each record to the exports channel and register an exporter. The README's quotes example does exactly this against quotes.toscrape.com:

go
func main() {
    geziyor.NewGeziyor(&geziyor.Options{
        StartURLs: []string{"http://quotes.toscrape.com/"},
        ParseFunc: quotesParse,
        Exporters: []export.Exporter{&export.JSON{}},
    }).Start()
}

func quotesParse(g *geziyor.Geziyor, r *client.Response) {
    r.HTMLDoc.Find("div.quote").Each(func(i int, s *goquery.Selection) {
        g.Exports <- map[string]interface{}{
            "text":   s.Find("span.text").Text(),
            "author": s.Find("small.author").Text(),
        }
    })
}

Running that program writes a JSON file containing one object per quote. If you need rendered pages instead, swap the enqueue call for g.GetRendered and the same callback receives the post-JavaScript DOM.

Chrome rendering, descriptors and the limits you actually hit

JS rendering is the feature that separates Geziyor from a plain HTTP scraper, and it is also the heaviest part of the dependency tree. chromedp and cdproto pull in a Chrome DevTools Protocol client. Every rendered request costs a browser round trip, so throughput on rendered pages is bounded by Chrome, not by Go's scheduler. The 5.000+ Requests/Sec figure in the README's feature list describes the non-rendered path; the README does not attach that number to rendering.

The README carries an explicit operational warning for macOS: the platform limits open file descriptors, and concurrent requests above 256 require raising that limit. That is not a Geziyor bug, it is a property of the host, but it is the first thing that will look like a crawler defect when requests start failing under load.

Concurrency is configurable globally and per domain, and request delays can be constant or randomized. Those two options are the polite-crawling controls. There is no mention of a retry policy, a backoff schedule or a dead-letter queue in the README, so transient 5xx responses and network timeouts are yours to handle in the parse callback or in a middleware. This is a genuine gap for long unattended crawls.

Proxy support is per-request and covers HTTP, HTTPS and SOCKS5. One caveat is called out directly: an http-scheme proxy is used for http requests and not for https requests, so a mixed-target crawl needs both schemes listed.

When Geziyor is the wrong tool

If your crawl needs to survive a restart and continue where it stopped, Geziyor gives you no documented mechanism for that. The queue is in memory, the state is in your closures, and a killed process loses both. A crawler that must run for days across millions of URLs needs a persistent frontier, which means either building one around Geziyor or choosing a framework that ships one.

If the people operating the crawler are not Go developers, the callback model is a barrier. There is no configuration file format described in the README, no spider DSL, and no control panel. Every change to what gets extracted is a code change and a rebuild.

If you only need to fetch a handful of pages once, the framework is overhead. Adding goquery, chromedp, goleveldb, diskv and the Prometheus client to a small program to parse one page is a poor trade against net/http and goquery alone.

And if the target site forbids crawling, none of the above matters. Geziyor parses robots.txt, but parsing rules is not the same as having permission. The README does not discuss legal or terms-of-service considerations, and the MPL-2.0 licence on the code says nothing about the sites you point it at.

How Geziyor differs from Colly

Colly is the other well-known Go scraping library, and the difference is where each puts the browser. Colly is built around an HTTP-first collector with its own request/response abstractions and a storage interface for visited-URL tracking; rendering is an add-on through a separate module. Geziyor treats rendering as a first-class call, g.GetRendered, sitting beside g.Get in the same API, with chromedp already in the main dependency list.

That choice has a cost. Geziyor's go.mod always pulls chromedp, cdproto and the Prometheus client, whether or not you render anything or expose metrics. Colly's core is smaller because those concerns live outside it. If your crawl never touches JavaScript and you care about binary size or dependency surface, Colly's split is the more economical arrangement.

The second difference is output. Geziyor ships exporters and a metrics interface in the core repository, so a crawl that writes CSV and exposes Prometheus counters is a few lines of options. With Colly you would wire the export and instrumentation yourself. Neither approach is better in the abstract; Geziyor optimises for getting a complete, observable crawler running quickly, and pays for it in dependencies.

Maintenance, upgrading and what the licence means for your build

The repository is not archived, and the last push was on 2026-07-02. There are no tagged releases for this project, so versioning is by commit and by the module path github.com/geziyor/geziyor. That matters for reproducibility: with no tagged releases, pinning means pinning a pseudo-version or a commit hash in go.mod rather than a semantic version. The README's own advice to use Go modules points in the same direction.

The module declares go 1.15. That is the language version floor, not a statement about the toolchain you must run, but it does mean the project targets an older Go baseline than current releases. Upgrading your own toolchain is unlikely to break it; upgrading Geziyor itself is the risk, because a pseudo-version bump can move several transitive dependencies at once. chromedp and cdproto are the ones most likely to shift behaviour, since they track Chrome's DevTools Protocol.

Licensing is MPL-2.0, a file-level copyleft licence. Modifying Geziyor's own source files carries obligations on those files; importing it as a module dependency does not relicense your application. This is a general description of the licence family, not legal advice, and the LICENSE.txt in the repository is the authoritative text. Note that the dependency licences differ from the project licence, and your own compliance review should cover goquery, chromedp, goleveldb, diskv and the Prometheus client separately.

Editorial conclusion

Adopt Geziyor if your team already writes Go and wants crawling, extraction and export inside one binary, with per-domain concurrency limits and a Prometheus endpoint you control. Do not adopt it if you need a scheduler, a dashboard or a retry queue out of the box, or if you cannot guarantee a Chrome binary on the host when pages require rendering. Before committing, verify three things against your own targets: that the sites you crawl allow it under robots.txt and their terms, that your file descriptor limit is raised above 256 if you plan on more than that many concurrent requests, and that the exporter you pick writes the fields your downstream pipeline expects.

Frequently asked questions

Does Geziyor require Chrome to be installed?

Only for rendered requests. The README states that if you want to make JS rendered requests, you should make sure you have Chrome installed, because Geziyor uses the local Chrome application CLI by default. You can instead point the BrowserEndpoint option at a different Chrome instance.

Why does Geziyor fail with many concurrent requests on macOS?

The README notes that macOS limits the maximum number of open file descriptors, and that concurrent requests above 256 require raising that limit. The README links to an external article on maximum limits rather than giving the commands inline.

Can Geziyor export scraped data to JSON or CSV?

Yes. You send data to the Geziyor.Exports channel and register exporters in the Options, and the README lists JSON, CSV and custom exporters. The quotes example registers export.JSON and writes one object per quote.

Does Geziyor support proxies?

It supports HTTP, HTTPS and SOCKS5 proxies. You can set the HTTP_PROXY and HTTPS_PROXY environment values for a single proxy, or set the ProxyFunc option to client.RoundRobinProxy for in-order selection per request. The README warns that an http-scheme proxy is used for http requests and not for https requests.

Official sources

  1. geziyor/geziyor on GitHub
  2. Issues
  3. License: MPL-2.0
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/geziyor-geziyor.svg)](https://hysenlabs.com/projects/geziyor-geziyor)