Library / SDK
Boris-code/feapder avatar
Boris-code/feapder

feapder: a Python crawler framework with four spider types and a management platform

🚀🚀🚀feapder is an easy to use, powerful crawler framework | feapder是一款上手简单,功能强大的Python爬虫框架。内置AirSpider、Spider、TaskSpider、BatchSpider四种爬虫解决不同场景的需求。且支持断点续爬、监控报警、浏览器渲染、海量数据去重等功能。更有功能强大的爬虫管理系统feaplat为其提供方便的部署及调度

3,738 stars551 forksPythonNOASSERTION

At a glance

What is it?
feapder ships AirSpider, Spider, TaskSpider and BatchSpider for different crawling scenarios, plus resumable crawling and a separate deployment platform called feaplat. The four spider types are also where the learning curve starts.
Who is it for?
Adopt feapder if you need distributed or batch crawling with resumable tasks and you are willing to provision Redis and MySQL, because TaskSpider and BatchSpider depend on external stores. Do not adopt it for a handful of pages that requests plus parsel would cover, and do not expect the render and mongo paths to work from the slim install.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 46 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 6, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What feapder solves, and which crawler you actually need

The README describes feapder as an easy to use Python crawler framework with four built-in spider types: AirSpider, Spider, TaskSpider and BatchSpider. That list is the project's real organising idea. Each type targets a different scenario, and the choice determines how much infrastructure you have to run.

AirSpider is the entry point. The README's first example subclasses feapder.AirSpider, yields a request from start_requests, and prints the response in parse. There is no database in that path. Spider adds the pieces a production crawl usually needs. TaskSpider and BatchSpider move the task queue out of the process, which is what makes distributed and batch collection possible, and the README lists resumable crawling, monitoring and alerting, browser rendering and large-scale deduplication as features of the framework as a whole.

The intended audience is a Python engineer who already writes scrapers and wants scheduling, deduplication and restart behaviour handled by a framework rather than by hand. The README is not aimed at people who have never written a parser. The four-type split asks you to classify your job before you write code, and picking wrong means either running infrastructure you do not need or hitting a wall when the crawl has to survive a restart.

How the four spider types differ in mechanism

The architecture visible in the README is a producer and a parser. start_requests produces tasks, parse consumes responses. AirSpider runs that loop in-process, and the sample log output ends with the line "无任务,爬虫结束", which is the framework reporting that the task queue has drained and the spider is stopping. That is the whole lifecycle for the simple case: no external queue, no scheduler.

TaskSpider and BatchSpider are the types that make the framework distributed. The README's package metadata describes feapder as supporting distributed collection, batch collection, data loss prevention and rich alerting, and the install extras point at the supporting stores: redis is a base dependency, PyMySQL is a base dependency, and pymongo only appears in the all extra. So the distributed and batch paths lean on Redis for the task queue and on a relational database for the data, while Mongo storage is an opt-in piece you get from the full install.

The deduplication story follows the same split. The README's three install variants state that the slim version does not support in-memory deduplication and does not support writing to Mongo, and that the browser render version does not support in-memory deduplication either. Only the full version supports everything. That is a concrete design boundary worth reading twice: if you install the slim package and then look for memory-based deduplication, the feature is not there, and the README says so up front rather than in a troubleshooting note.

Installing feapder and running a first spider

The README gives three install commands and states the environment requirement as Python 3.6.0+ on Linux, Windows and macOS. Note that setup.py is stricter than the README: it raises SystemExit with the message "Sorry! feapder requires python 3.9.0 or later." when the interpreter is below 3.9.0, and it sets python_requires to >=3.9. For a new project, trust the packaging metadata over the badge.

The slim install is the smallest footprint and pulls only the base dependencies listed in setup.py, which include requests, parsel, redis, PyMySQL and loguru.

bash
pip install feapder

If you need browser rendering, the render extra adds webdriver-manager, playwright and selenium. The README warns that the full version can fail to install and points at an installation troubleshooting document, so try the smaller variant first if you do not need Mongo storage.

bash
pip install "feapder[render]"

The full version adds bitarray, PyExecJS and pymongo on top of the render dependencies, and the README says it is the only variant that supports all features.

bash
pip install "feapder[all]"

Once installed, the CLI scaffolds a spider file. The README shows the command and the generated class, which is an AirSpider that fetches a URL and prints the response object.

bash
feapder create -s first_spider
python
import feapder


class FirstSpider(feapder.AirSpider):
    def start_requests(self):
        yield feapder.Request("https://www.baidu.com")

    def parse(self, request, response):
        print(response)


if __name__ == "__main__":
    FirstSpider().start()

Running that file should print a request debug block followed by a response object, and then a line reporting that there are no tasks left and the spider has ended. If the process exits immediately with an import error, the likely cause is the Python version check in setup.py rather than a missing dependency.

Where feapder is the wrong tool

The framework's cost is proportional to its feature set, and the README makes that visible in the install matrix. If your job is a few hundred pages fetched once, AirSpider works, but so does requests plus parsel, which are already dependencies of feapder. You would be adopting a framework to avoid writing a loop.

The sharper limitation is that the distributed and batch spider types are not self-contained. TaskSpider and BatchSpider exist to solve queue and restart problems, and the dependencies that make that possible are Redis and a database. If you cannot run Redis, or your environment forbids an external task store, those two spider types are not available to you regardless of how well the code is written. The README does not document a fallback queue for that case.

There is also a documentation gap around failure handling. The README lists data loss prevention and alerting as features but does not describe the retry policy, the alert transport, or what happens to in-flight tasks when a worker dies mid-crawl. Those details live in the documentation site and the changelog rather than in the README. If your acceptance criteria include a precise restart guarantee, the README alone will not let you verify it.

Finally, the README does not document rollback or downgrade steps between the 1.9.x releases. Treat an upgrade as something to test in a staging environment first, not as a routine pip command.

feapder compared with Scrapy's approach

The repository topics list scrapy alongside feapder, which invites the comparison. The difference is where the queue lives and how many spider shapes you get out of the box.

Scrapy's model is a single engine with a scheduler and a downloader, and the framework is designed around that engine from the start. feapder instead exposes four named spider classes, and the README's framing is that each one solves a different scenario. AirSpider is deliberately simpler than a Scrapy project, with no settings module and no middleware stack to configure before the first request goes out. TaskSpider and BatchSpider are the answer to distribution, and they push the queue into Redis rather than keeping it in the engine process.

That means the migration path differs. Moving a feapder AirSpider to a distributed setup is a change of base class plus the infrastructure to back it, whereas in Scrapy the distribution story is typically an add-on around the existing engine. Neither is strictly better. feapder's four-class design makes the trade-off explicit at the top of your file, and that is also its cost: the class you inherit from is a commitment about what you will run in production.

The second difference is the management platform. feapder's README points at feaplat for deployment and scheduling, a separate system that the framework is designed to work with. Scrapy does not ship an equivalent in the same repository, and operators usually assemble that layer themselves. If you want the framework and the deployment console to come from one project, feapder is the one making that offer.

Maintenance, licence and the cost of staying current

The repository is not archived, and the last push was on 2026-08-21. The most recent release listed is v1.9.3 on 2025-12-16, preceded by v1.9.2 on 2025-02-14 and v1.9.0 on 2024-03-19. The release cadence is therefore uneven: roughly a year between the 1.9.0 and 1.9.2 tags, then about ten months to 1.9.3. Plan upgrades as occasional events that need testing, not as a steady stream of small patches.

The licence field is NOASSERTION, which means the automated classifier could not map the LICENSE file to a known identifier. setup.py, however, declares license="MIT" in the setuptools metadata. Those two signals disagree, and the LICENSE file is the authoritative one. Read it before you ship feapder inside a commercial product; this is a fact to check, not legal advice.

Upgrade cost has a second component beyond the package itself. The install extras pull in playwright, selenium and webdriver-manager, all of which track browser binaries and driver versions independently of feapder. A feapder upgrade that is compatible with your spider code can still break because the browser tooling underneath moved. Pinning the extras in your own requirements file is the practical way to keep those two upgrade cycles separate.

Editorial conclusion

Adopt feapder if you need distributed or batch crawling with resumable tasks and you are willing to provision Redis and MySQL, because TaskSpider and BatchSpider depend on external stores. Do not adopt it for a handful of pages that requests plus parsel would cover, and do not expect the render and mongo paths to work from the slim install. Before committing, verify that your Python is 3.9.0 or later, that the extras you need install cleanly, and that the deployment story you want is feaplat rather than something the framework does on its own.

Frequently asked questions

How do I install feapder, and which install variant should I pick?

The README gives three commands: pip install feapder for the slim version, pip install "feapder[render]" for browser rendering, and pip install "feapder[all]" for everything. The slim version does not support browser rendering, in-memory deduplication or writing to Mongo, so pick the full variant if you need any of those.

What Python version does feapder require?

The README badge says Python 3.6, but setup.py raises SystemExit with the message "Sorry! feapder requires python 3.9.0 or later." below 3.9.0 and sets python_requires to >=3.9. Use 3.9.0 or later.

What is the difference between AirSpider, Spider, TaskSpider and BatchSpider in feapder?

The README states that feapder has four built-in spider types to cover different scenarios, and AirSpider is the one shown in the first example, where start_requests produces tasks and parse consumes responses in-process. TaskSpider and BatchSpider are the distributed and batch types, which is why the framework depends on Redis and PyMySQL.

Does feapder support resumable crawling and deduplication?

The README lists resumable crawling and large-scale deduplication among the framework's features. Note that in-memory deduplication is not available in the slim or render install variants; only the full version supports it.

What licence is feapder released under?

setup.py declares license="MIT" in the setuptools metadata, but the repository's licence field is NOASSERTION, meaning the classifier could not map the LICENSE file to a known identifier. Check the LICENSE file itself before relying on either signal.

Official sources

  1. Boris-code/feapder on GitHub
  2. Issues
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/boris-code-feapder.svg)](https://hysenlabs.com/projects/boris-code-feapder)