# EasySpider 0.6.5: JSON tasks, an AGPL licence, and a thesis as the specification

> EasySpider is a visual, no-code crawler where you design a task by right clicking page elements and the software matches the rest, then save the result as a JSON document. Three things about it are not obvious from the front page: the licence is AGPL-3.0 while the README advertises commercial use, the feature inventory is a set of screenshots rather than text, and the most detailed description of the design model is a chapter of the author's master's thesis.

**NaiboWang/EasySpider** — A visual no-code/code-free web crawler/spider易采集：一个可视化浏览器自动化测试/数据采集/网页爬虫软件，可以无代码图形化的设计和执行爬虫任务。别名：ServiceWrapper面向Web应用的智能化服务封装系统。

- Repository: https://github.com/NaiboWang/EasySpider
- Website: https://www.easyspider.net
- Stars: 44,614 · Forks: 5,390
- Language: JavaScript
- License: AGPL-3.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/naibowang-easyspider

## The readme file is Readme.md and the feature list is a set of screenshots

Two small facts shape how much of this project a reader actually sees. The top-level file is named Readme.md rather than README.md, so tooling and conventions that look for the uppercase name do not find it on a case-sensitive filesystem. And the feature section is not text. The line reads that more features please scroll to the bottom of this page to view, and what is at the bottom is a set of images in the media/ directory. Consequence: the project's own inventory of what it does is screenshots. A screen reader user, a text-only reader, a search engine, or a tool reading the repository gets no feature list at all, and the only machine-readable description of the capabilities is the two worked examples further up the page. This is not a criticism of the software, it is a documentation decision, but it means the README cannot answer the question what can this do beyond the two examples it walks through.

## AGPL-3.0 and free for commercial use are both true, and the second is the narrower claim

The repository is licensed AGPL-3.0. The README describes the software as completely free, including for commercial use and secondary development, and says development and maintenance are done by the author. Both statements hold at once, because AGPL-3.0 charges nothing; what it constrains is distribution. The README also says the software can be executed separately in the command line so that it can be easily embedded into other systems, and that is precisely the scenario the network clause of AGPL was written about: modify the program, let other people interact with it remotely, and the corresponding source has to be offered to them. Consequence: the question to answer before embedding is not the price, which is zero either way, but whether your deployment model creates an obligation. An unmodified internal tool is a non-issue. A hosted service running a modified copy is a different conversation, and the licence file is where you find that out.

## A saved task is a JSON document, and the escape hatch from no-code is a JavaScript condition

The Examples/ directory is the most instructive part of the repository, and its contents are data rather than scripts. There are files for a paginated list collection, for looping into each detail page without opening a new tab, for collecting car news, for collecting a listing count, one named Exec and Eval, and one named as a JavaScript condition example, plus a directory of sample tasks written with Python and a readme. Consequence, positive side: a task is a document you can diff, copy, review and put in version control, which is the real difference from a script-based crawler where the crawler is a program. Cost side: the flow is a sequence of interface actions rather than a program, so there is no code representation of the same task to unit test or read in a pull request, and the one place you can express logic is a JavaScript condition embedded in a task. That is where the code-free claim has its boundary. Also worth noting for an international audience, most example filenames are in Chinese.

## The design step needs something in the browser, not just the desktop window

Both worked examples begin the same way, with a right click on the page. Select a large product block and the software detects the other blocks of the same type, then Select All, then Select Child Elements, then Collect Data, and the results come back split into sub-fields. In the second example you right click a product title, the same-type titles are matched automatically, and if you then choose Loop-click every element it opens each product's detail page so you can set up collection from there. Two directories at the top of the tree explain how that works: ElectronJS/ for the desktop shell, which is why the repository's primary language is JavaScript, and Extension/ for a component that lives in the browser. Consequence: setting up a machine means installing a desktop application and a browser-side piece, and the two have to agree. The automatic matching is also what makes the tool approachable and what makes it fragile, since it matches on appearance, so a listing page whose repeated elements change shape partway down produces a selection that is wrong on some pages.

## The sponsor block is proxy and CAPTCHA vendors, and none of that ships in the box

The sponsorship section is unusual and worth reading as documentation of the intended workload. It lists Bright Data as a proxy network with 150 million or more addresses worldwide offering residential IPs and a web unlocker, Webshare with more than 80 million residential, datacenter and ISP proxies across 195 countries starting at 1.40 dollars per gigabyte, Thordata with residential, ISP and datacenter proxies plus SERP and Web Scraper APIs across 190 or more regions, and CapSolver as a CAPTCHA-solving service covering reCAPTCHA, image CAPTCHAs, Cloudflare and AWS WAF. Each comes with a referral code, ESN, SPIDER20, or EasySpider10, and each has a China-domestic signup route. Consequence: the target workload is sites that resist automation, and the project's answer to that is paid third-party services rather than a built-in capability. A plan that assumes the crawler handles blocking on its own will not work, and the referral codes are worth knowing about when you read the free and ad-free claims.

## The thorough specification is chapters three and five of the author's thesis

Documentation is pointed at three places and they are not equivalent. The GitHub wiki is the entry point. An eBay sample walkthrough lives on a CSDN blog. And the Docs/ directory contains the author's master's thesis as a PDF, with the README telling you to mainly read chapters three and five, and the file name translating to a thesis on intelligent service wrapping systems for web applications. Consequence: the design model is specified thoroughly in an academic document, which means it is not searchable, not linkable to a heading, not diffable, and not readable by any tool. A reader who needs one specific behaviour ends up in a PDF. The project description carries the same ambiguity, calling the software a visual crawler in one clause and a service wrapper system in the next, under the alias ServiceWrapper. Resolving what the project is, crawler or service wrapper, is worth doing before you commit, and the thesis is where the answer is.

## The three newest tags span two years and the numbering skips 0.6.4

The release list does not read like a version history. v0.6.2 is dated 2024-04-21, v0.6.3 is dated 2025-01-01, and v0.6.5 is dated 2026-08-19, with 0.6.4 absent from the three most recent. The default branch was pushed on 2026-09-17. Consequence: the newest three tags span about two years and four months, and the highest number is not reliably the newest thing in the repository, so pin an exact version and read the release notes rather than assuming a higher minor means a fresher build. Maintenance is a single-author effort by the project's own account, funded through GitHub Sponsors, Alipay, WeChat and PayPal, and there are two official sites, a Chinese one and an international one, with a separate China-domestic download mirror offered when the GitHub releases page is slow. For downloads, the README points at the GitHub releases page for the latest version and does not document a package manager install or a container image, despite searching for one being common.

## Conclusion

EasySpider suits a reader who collects a defined set of fields from a handful of sites, who wants the task to be a reviewable data file rather than a script, and who is comfortable with a Chinese-first README and a wiki. It does not suit a reader who needs a text-readable feature list, who needs an English-first specification, or who intends to modify it and host the result without planning for the AGPL network clause. Before you build a pipeline on it, check five things: whether the licence suits your deployment, because the repository is AGPL-3.0 and the README invites command line embedding into other systems; whether the automatic element matching holds on every page shape you need, since that is what the whole design rests on; whether your blocking problem is solved, because the README's sponsors are paid proxy and CAPTCHA vendors and none of that is built in; which exact tag you are on, because the three newest releases span about two years and the numbering skips 0.6.4; and where the detail lives, because the thorough version of the design model is chapters three and five of a thesis PDF in the repository rather than the wiki.

## FAQ

### How do web crawlers work?

In EasySpider the flow is a right click on the page element you care about, and the software detects the other elements of the same type automatically. You then choose Select All, optionally Select Child Elements or Loop-click every element to open each detail page, and finally Collect Data, which saves results split into sub-fields. A finished task is stored as a JSON document and can also be run from the command line.

### What does it mean when a web search engine is crawling?

That describes a search engine building an index, which is a different job from the one this project does. EasySpider collects a defined set of fields from pages you navigate yourself, and its documentation is the GitHub wiki, with the design model described at greater length in the author's thesis, where the README suggests chapters three and five.

### Is web development easy for beginners?

The repository does not address that question, and it shows the inverse case instead: a crawler configured without writing code, where both worked examples are a sequence of interface actions ending in Collect Data. The one place logic can be written is a JavaScript condition inside a task, and the examples directory includes a file demonstrating exactly that.

### Does WebCrawler still exist?

The project does not mention a product under that name, so there is nothing there to confirm. The crawler here is EasySpider, a visual no-code browser automation and data collection tool under AGPL-3.0, with a desktop application built on Electron, downloads from the GitHub releases page, and a China-domestic mirror offered for slow connections.

## Sources

- [License: AGPL-3.0](https://github.com/NaiboWang/EasySpider/blob/master/LICENSE)
- [NaiboWang/EasySpider on GitHub](https://github.com/NaiboWang/EasySpider)
- [Project website](https://www.easyspider.net)
- [README](https://github.com/NaiboWang/EasySpider/blob/master/README.md)
- [Releases](https://github.com/NaiboWang/EasySpider/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/naibowang-easyspider
