InfoSpider: a GUI data-portability crawler for 24+ Chinese and Western accounts
INFO-SPIDER 是一个集众多数据源于一身的爬虫工具箱🧰,旨在安全快捷的帮助用户拿回自己的数据,工具代码开源,流程透明。支持数据源包括GitHub、QQ邮箱、网易邮箱、阿里邮箱、新浪邮箱、Hotmail邮箱、Outlook邮箱、京东、淘宝、支付宝、中国移动、中国联通、中国电信、知乎、哔哩哔哩、网易云音乐、QQ好友、QQ群、生成朋友圈相册、浏览器浏览历史、12306、博客园、CSDN博客、开源中国博客、简书。
At a glance
- What is it?
- InfoSpider is an open source Python toolbox that logs into your own accounts with Selenium and exports the data as JSON, with charts for four blog platforms. It is Windows-only, pinned to old libraries, and the README itself says the crawlers need continuous upkeep.
- Who is it for?
- InfoSpider is for Windows users who want a local, inspectable copy of data held by Chinese and Western platforms and who accept that the crawlers break when sites change. It is not for anyone needing unattended, cross-platform or long-running collection: the README states v1.0 was only tested on Windows with Python 3.7, and that the crawler approach creates a timeliness problem requiring continuous maintenance.
- Can I use it commercially?
- Yes, with conditions. GPL-3.0 is a copyleft licence: if you distribute software that includes it, you must release that software's source code under the same licence. Running it internally without distributing it does not trigger that obligation.
- Is it still maintained?
- Yes. The repository last received commits 162 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What InfoSpider solves, and for whom
Your account history is spread across mail providers, e-commerce, telecom operators, music services and blogging platforms. Each one keeps its own copy, and none of them hands you a merged file. InfoSpider's stated goal is to reverse that: the README describes it as a crawler toolbox that helps users "拿回自己的数据" (take back their own data), with open code and a transparent process. The intended user is an individual, not a data team. The README's developer memoir is written in the second person about ordinary browsing, and the QuickStart assumes a person clicking buttons in a window. Supported sources listed in the README include GitHub, QQ Mail, NetEase Mail, Alibaba Mail, Sina Mail, Hotmail, Outlook, JD, Taobao, Alipay, China Mobile, China Unicom, China Telecom, Zhihu, Bilibili, NetEase Cloud Music, QQ friends, QQ groups, browser history, 12306, Blog Garden, CSDN, Oschina and Jianshu. That list is the product. If your data lives elsewhere, the toolbox has nothing to offer you.
How the crawling actually works: Selenium, a GUI, and one folder per source
There is no API layer and no server. Each source is a separate script under the Spiders/ directory, and the README states plainly that all crawler scripts live there and are independent of one another. The runtime is Selenium 3.141.0 driving a real Chrome browser, so the flow is: you launch the GUI, click a source button, choose a save path, and a browser window opens where you type your own username and password. The README says the crawl starts automatically after login and the browser closes when it finishes. Output is uniform: everything is written as JSON, which the README calls out as a deliberate choice so that later analysis is easier. Charts are a separate, thinner layer. The README lists data analysis as supported for only four sources: Blog Garden, CSDN, Oschina and Jianshu, and describes the feature as "目前仅部分支持" (currently only partially supported). The GUI itself is wxPython 4.0.7, per requirements.txt. That architecture explains both the appeal and the failure mode: because the browser is real, anti-bot measures see a real browser, but because the browser is real, every site redesign is a code change.
Installing InfoSpider and exporting your first data source
The README gives a short dependency list before anything runs. Python 3 and Chrome are prerequisites, and the chromedriver build must match the installed Chrome version. Note that the README's own instructions are written for Python 3.7 on Windows, which is what the author says v1.0 was tested against. The repository also ships install_deps.sh at the top level, though the README's documented path is pip.
pip install -r requirements.txtThat installs the pinned set, including selenium==3.141.0, wxPython==4.0.7, pyecharts==1.7.1 and pandas==1.0.1. Several of those pins are old, so expect to fight your Python version rather than the code. The README acknowledges this friction directly: if the dependency step causes trouble, it points to a paid pre-packaged build instead.
With dependencies in place, the documented launch is two steps. Enter the tools directory, then run the GUI entry point.
cd tools
python3 main.pyA window opens. You click the button for the source you want, then pick where the output should go. A Chrome window appears; you enter your credentials there, and per the README the crawl begins on its own and the browser closes when it is done. What you should see afterwards, in the directory you chose, is a .json file with the scraped records. For the four blog platforms that have analysis support, you should also see an .html chart file. If no browser window appears at all, the usual cause is a chromedriver that does not match your Chrome build, since the README requires them to be the same version.
The maintenance problem the README admits to
The developer notes are unusually candid on this point: because the project collects data by crawling, it has a timeliness problem and needs continuous maintenance to track site updates. That is not a caveat bolted on at the end, it is the central cost of the design. Every source is a hand-written scraper against a login flow, so a changed form field, a new captcha step or a shifted DOM breaks that source until someone edits its script. The repository's last push was on 2026-04-21, and the only release listed is v1.0 from 2020-08-18. Whatever state the individual spiders are in, the version number has not moved in years. There is also a platform limit stated outright: v1.0 was tested only on Windows, with Python 3.7, and is not adapted to multiple platforms. The README's plan section lists a web interface for multi-platform use as future work, which confirms it does not exist yet. The data analysis side is narrower still, covering four blogging sites out of more than twenty sources.
The paid build, and what the open repository does not include
InfoSpider has an unusual split that a prospective user should understand before cloning. The README advertises a purchasable package that bundles the latest maintained version, broader personal data analysis, a pre-packaged binary that runs without installing dependencies, a guide to packaging InfoSpider yourself, one-to-one developer support, and free access to a planned 2.0 release. The free repository is the GPL-3.0 source. Practically, that means the open code is the reference implementation while the packaged build is the path of least resistance for a non-developer. It also means the README's claim of a "最新维护版本" (latest maintained version) sits behind a paywall, so you cannot judge from the repository alone how current the maintained spiders are. Nothing here is a licence problem: the code is GPL-3.0 and the paid item is a distribution, not a different licence. But if you need the maintained spiders, the free tree is not where they live.
Where InfoSpider is the wrong tool
If you need scheduled, headless, repeatable exports, this is the wrong shape. The documented flow is a person clicking a button and typing a password into a browser window, which does not fit a cron job or an unattended pipeline. If you are on macOS or Linux, the README says v1.0 is Windows-only and not adapted to other platforms, so you are on your own. If you need a stable interface for a downstream system, the JSON schema is per-spider and the README does not document a versioned contract. And if you want ongoing collection rather than a one-time archive, the maintenance admission matters more than the feature list: you are adopting a set of scrapers whose correctness decays with each site update. The honest framing is that InfoSpider is a snapshot tool with a GUI, not an integration platform.
How it differs from generic scraping frameworks
The obvious comparison is Scrapy, and the difference is not quality but shape. Scrapy is a framework: you write spiders, manage scheduling, retries and pipelines, and you own the target-specific logic. InfoSpider inverts that. The target-specific logic is already written for 24+ named services, and it ships a wxPython desktop interface so a non-programmer can run it. What you give up is control: you cannot easily schedule it, you inherit its pinned Selenium 3.x and Python 3.7 assumptions, and you depend on upstream maintenance for each site. For a developer who already knows Selenium, the reusable part is the Spiders/ directory, which the README explicitly says is structured so it can be ported into your own program. For someone who just wants their Zhihu and JD history on disk, a framework is the wrong amount of work.
Licence and upgrade cost
The repository is GPL-3.0, and the README's badge and License section agree. If you only run InfoSpider locally to export your own data, the copyleft obligations do not attach to your data. If you copy code out of Spiders/ into another program that you distribute, GPL-3.0 terms apply to that distribution; that is a question for a lawyer, not for this article. On upgrades, the picture is simple and not encouraging: requirements.txt pins exact versions (selenium==3.141.0, wxPython==4.0.7, pandas==1.0.1, numpy==1.22.0), so moving to a current Python means unpinning and retesting the GUI and every spider. The README's own escape hatch for that work is the packaged build. The only release tag in the repository is v1.0 from 2020, and the README's plan section describes a 2.0 refactor toward a web interface; until that ships, budget for dependency archaeology rather than a smooth upgrade path.
Editorial conclusion
InfoSpider is for Windows users who want a local, inspectable copy of data held by Chinese and Western platforms and who accept that the crawlers break when sites change. It is not for anyone needing unattended, cross-platform or long-running collection: the README states v1.0 was only tested on Windows with Python 3.7, and that the crawler approach creates a timeliness problem requiring continuous maintenance. Before adopting it, check whether each target site still matches its script under Spiders/, confirm the chromedriver build matches your installed Chrome, and read the GPL-3.0 terms if you intend to reuse the spider code inside another product.
Frequently asked questions
Does InfoSpider run on macOS or Linux?
The README states that v1.0 was tested only on Windows with Python 3.7 and is not adapted to multiple platforms. A multi-platform web interface is listed in the project's plan section as future work, not as a shipped feature.
Which data sources does InfoSpider support?
The README lists GitHub, QQ Mail, NetEase Mail, Alibaba Mail, Sina Mail, Hotmail, Outlook, JD, Taobao, Alipay, China Mobile, China Unicom, China Telecom, Zhihu, Bilibili, NetEase Cloud Music, QQ friends, QQ groups, browser history, 12306, Blog Garden, CSDN, Oschina and Jianshu. Chart generation is listed as supported for only Blog Garden, CSDN, Oschina and Jianshu.
What does InfoSpider output after a crawl finishes?
The README says all scraped data is stored as JSON so that later analysis is easier, and that you choose the save path before the crawl starts. For the four blog sources with analysis support, an HTML chart file is produced alongside the JSON.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/kangvcar-infospider)