Self-hosted service
dotnetcore/DotnetSpider avatar
dotnetcore/DotnetSpider

DotnetSpider: a .NET Standard crawling framework built around entity classes

DotnetSpider, a .NET standard web crawling library. It is lightweight, efficient and fast high-level web crawling & scraping framework

4,140 stars1,051 forksC#MIT

At a glance

What is it?
DotnetSpider is an MIT-licensed C# crawling library that turns HTML pages into typed entities through attributes. It suits .NET teams already running MySQL, Redis or RabbitMQ, and it asks for more infrastructure than a scripted scraper does.
Who is it for?
Adopt DotnetSpider if your team already writes C# and you want page data landing in typed classes backed by MySQL, SQL Server, PostgreSQL or MongoDB, with Redis and RabbitMQ available when you scale out. Do not adopt it if you need a browser-rendered page today, since the README lists the Puppeteer downloader as coming soon, or if you cannot run the storage and queue services the samples assume.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 180 days ago.
What is it written in?
Mainly C#, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem DotnetSpider solves for C# teams

Most scraping code starts as a loop: fetch a URL, parse the HTML, write rows somewhere. That loop gets rewritten for every site, and the parsing rules end up scattered across methods. DotnetSpider's answer is to make the target shape of the data the centre of the program. You declare a class, decorate it with selectors, and the framework handles fetching, parsing, deduplication, storage and, if you configure it, distribution across machines.

The intended user is a .NET developer who is comfortable with dependency injection, attributes and entity framework style configuration. The README's sample spider is not a script; it is a class with a constructor taking IOptions<SpiderOptions>, DependenceServices and ILogger<Spider>. If that shape looks natural to you, the framework will feel consistent. If you wanted three lines of Python and a CSV file, it will not.

The README opens with a disclaimer in Chinese stating the framework exists to simplify development and that users are responsible for complying with their own country's laws. That is a normal framing for a crawling library, and it is worth reading as a statement about intended use rather than boilerplate.

How the entity, selector and data flow model fits together

The mechanism visible in the sample is a chain. A Spider subclass registers data flows in InitializeAsync. The sample calls AddDataFlow<DataParser<CnblogsEntry>>() and then AddDataFlow(GetDefaultStorage), which means parsing happens before storage and the storage step is the default one the base class supplies. Requests are added with AddRequestsAsync, and each Request carries an environment dictionary. In the sample that dictionary holds the key 网站 with the value 博客园, and later selectors read it back with Type = SelectorType.Environment. Environment values travel with the request, so a field like WebSite or Category can be filled without parsing it out of the page.

The entity class carries the extraction rules. Schema("cnblogs", "news") names the target table. EntitySelector with an XPath expression of .//div[@class='news_block'] marks the repeating block, so one page yields many entities. GlobalValueSelector pulls page-level values once, such as the title element, and stores them under a name that ValueSelector can reference by that name with SelectorType.Environment. Field-level ValueSelector attributes then read either XPath expressions or those named environment values. Formatters run after extraction: ReplaceFormatter strips the string " - 博客园" from the title, and TrimFormatter cleans whitespace from PlainText.

Indexes are declared in code too. The Configure method calls HasIndex on Title and a composite unique index on WebSite plus Guid. That is a design decision with consequences: schema and uniqueness live in the entity, so a schema change means a code change and a redeploy, not a migration script you can review separately.

Deduplication and scheduling sit behind this. The repository includes docker-compose and a dockerfile directory, and the README's development environment section lists MySQL, Redis, SQL Server, PostgreSQL, MongoDB, RabbitMQ and HBase as services you start yourself. Redis is marked optional in that list, but the distributed spider page is where scheduling across nodes is described. The README does not document how request deduplication behaves when Redis is absent.

Installing DotnetSpider and running a first spider

The README does not give a NuGet install command, but the package badge links to the DotnetSpider package on nuget.org, so installation goes through the standard NuGet flow. The README does show a MyGet feed for beta packages, which you would add only if you want pre-release builds.

bash
dotnet add package DotnetSpider

The development environment section expects .NET Core 2.2 or later and a database. MySQL is the first listed and the sample schema targets it, so start there. The README gives this exact command, including the root password:

bash
docker run --name mysql -d -p 3306:3306 --restart always -e MYSQL_ROOT_PASSWORD=1qazZAQ! mysql:5.7

If you intend to run the spider across more than one process, the README also gives a Redis container and notes that Redis is optional:

bash
docker run --name redis -d -p 6379:6379 --restart always redis

With the services up, the smallest useful program follows the sample in the README: subclass Spider, register a parser and a storage flow, add a request, and declare an entity with selectors. The README points to src/DotnetSpider.Sample/samples/BaseUsageSpider.cs for the minimal version and EntitySpider.cs for the configurable entity version. The sample builder sets options.Speed = 1, calls UseSerilog(), calls IgnoreServerCertificateError(), and then awaits builder.Build().RunAsync(). The README does not state what Speed counts per second, so treat 1 as a value to raise once you have watched the log output.

One configuration note applies if you use the Redis scheduler. The README says to update your Redis config with two lines:

code
timeout 0
tcp-keepalive 60

Without that change, a long crawl can lose its connection to Redis. The README states the requirement but does not explain the failure it prevents.

Where DotnetSpider is the wrong tool

The most concrete limitation is in the README itself: under Puppeteer downloader it says "Coming soon". If your target pages render their content with JavaScript, a plain HTTP downloader will see an empty shell, and the framework as documented does not give you a browser-backed downloader. You would have to supply your own downloader implementation or pick a different tool.

The second constraint is infrastructure weight. The development environment section lists eleven services, and while several are marked optional, the sample spider stores into a database and the distributed mode depends on a message queue. A single-page scrape that writes a JSON file does not need MySQL, RabbitMQ and HBase, and pulling them in to satisfy the framework's defaults is a real cost in setup and operations.

The third is that the entity attribute model assumes the data has a stable shape. Selectors are compiled into the class, so a site redesign means editing and redeploying the entity. There is no documented way to override a selector at runtime from configuration. For a site you scrape once, that rigidity buys nothing.

Finally, the README does not document rollback, retry policy, or how a partially completed crawl resumes. Those may exist in the code, but the README is silent, and you should read the source before depending on them.

DotnetSpider compared with a general purpose scraping stack

The obvious alternative for a .NET team is a general purpose HTTP and HTML library combination, such as HttpClient plus HtmlAgilityPack, with your own scheduling and storage code. HtmlAgilityPack is itself in DotnetSpider's dependency table under the MIT licence, so the two are not in conflict: DotnetSpider uses it for parsing and adds the surrounding machinery.

The difference in approach is where the abstraction sits. With HttpClient and HtmlAgilityPack you write the control flow and the parsing separately, and you decide when to parallelise and where to persist. With DotnetSpider the control flow is the framework's, and your code is mostly declarations: which requests to add, which parser to register, which selectors to attach. That is a gain when you have many similar sites, because each new site is a new entity class rather than a new program. It is a loss when you have one site with unusual pagination or login handling, because you are working against the framework's assumptions instead of writing the loop you already understand.

If your team is not on .NET, the choice is different. A Python stack gives you a browser automation library that already works today, which matters given the "Coming soon" note on Puppeteer support here. The trade-off is language familiarity and deployment, not crawling capability.

Maintenance, licensing and the cost of upgrading

The repository is not archived, and the last push was on 2026-04-03. The release history is uneven: 5.0.3 shipped on 2021-02-18, 5.1.5 on 2024-06-05, and the most recent push is more recent than the most recent release. That pattern suggests development continues between releases, but it also means the released NuGet package may lag the master branch. Pin your package version and check the release notes rather than assuming master behaviour.

The licence is MIT, which is permissive and places few conditions on commercial use. The dependency table matters more than the project licence for compliance review, because it lists Apache 2.0, MIT, BSD 3-Clause and the PostgreSQL License among the packages DotnetSpider pulls in. If your organisation has an allow-list, that table is the starting point for the review. This is not legal advice; the dependency list is the document to hand to whoever does that review.

Upgrade cost is dominated by the entity model. Because schema names, indexes and selectors live in attributes, a framework upgrade that changes selector or schema semantics touches every entity class you have written. The README does not publish a migration guide for the 5.0 to 5.1 step, so budget time to read the diff.

Editorial conclusion

Adopt DotnetSpider if your team already writes C# and you want page data landing in typed classes backed by MySQL, SQL Server, PostgreSQL or MongoDB, with Redis and RabbitMQ available when you scale out. Do not adopt it if you need a browser-rendered page today, since the README lists the Puppeteer downloader as coming soon, or if you cannot run the storage and queue services the samples assume. Before writing your own spider, verify three things: that the DotnetSpider NuGet package version you install matches the 5.1.5 release line, that your MySQL instance accepts the schema and index attributes you declare, and that your Redis configuration sets timeout 0 and tcp-keepalive 60 if you use the Redis scheduler.

Frequently asked questions

What is DotnetSpider?

DotnetSpider is a .NET Standard web crawling and scraping library written in C#. The README describes it as lightweight, efficient and fast, and its sample shows pages being mapped to entity classes through attributes.

How do I install DotnetSpider?

The README links to the DotnetSpider package on nuget.org, so it installs through the normal NuGet flow. The README also lists a MyGet feed for beta packages if you want pre-release builds.

Does DotnetSpider need a database?

The development environment section lists MySQL, SQL Server, PostgreSQL, MongoDB and HBase, and the sample spider registers a storage data flow. The README does not document running the sample without a storage backend.

Can DotnetSpider render JavaScript pages?

The README's Puppeteer downloader section says "Coming soon", so browser rendering is not documented as available. Pages whose content is produced by JavaScript will not be visible to a plain HTTP downloader.

Official sources

  1. dotnetcore/DotnetSpider on GitHub
  2. Issues
  3. License: MIT
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/dotnetcore-dotnetspider.svg)](https://hysenlabs.com/projects/dotnetcore-dotnetspider)