Abot: an event-driven C# crawler that hands you pages instead of managing threads
Cross Platform C# web crawler framework built for speed and flexibility. Please star this project! +1.
At a glance
- What is it?
- A .NET web crawling framework built around events and pluggable core interfaces, targeting .NET Standard 2.0, with JavaScript rendering left to a companion project.
- Who is it for?
- Abot is a good fit when you want a crawler that will not fight you, and a poor fit if you need JavaScript rendering out of the box. The event model is genuinely the simplest part of the design: register handlers, receive a CrawledPage with parsed HTML already attached, and never think about the worker pool.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 81 days ago.
- What is it written in?
- Mainly C#, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 7, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The quick start is two demos, not one
Installation is a single NuGet package, and the quick start deliberately shows two different levels of use rather than one blessed path:
PM> Install-Package AbotThe first demo configures a polite crawler with a page budget and a per-domain delay, subscribes to one event, and starts it:
var config = new CrawlConfiguration
{
MaxPagesToCrawl = 10, //Only crawl 10 pages
MinCrawlDelayPerDomainMilliSeconds = 3000 //Wait this many millisecs between requests
};
var crawler = new PoliteWebCrawler(config);
crawler.PageCrawlCompleted += PageCrawlCompleted;//Several events available...
var crawlResult = await crawler.CrawlAsync(new Uri("http://!!!!!!!!YOURSITEHERE!!!!!!!!!.com"));The second demo bypasses the crawler entirely and makes a single request, which is the API you want when you only need one page:
var pageRequester = new PageRequester(new CrawlConfiguration(), new WebContentExtractor());
var crawledPage = await pageRequester.MakeRequestAsync(new Uri("http://google.com"));Note that the requester takes its configuration and its content extractor as constructor arguments. That is the pluggability claim made concrete at the smallest possible scale.
One thing to plan for: the sample imports Serilog and configures a console logger before anything else runs, so logging is treated as a first-class concern rather than an afterthought.
Events are the whole API surface you need for most crawls
You register handlers and Abot calls them. Four subscriptions cover most cases:
crawler.PageCrawlStarting += crawler_ProcessPageCrawlStarting;
crawler.PageCrawlCompleted += crawler_ProcessPageCrawlCompleted;
crawler.PageCrawlDisallowed += crawler_PageCrawlDisallowed;
crawler.PageLinksCrawlDisallowed += crawler_PageLinksCrawlDisallowed;The completed handler is where the value is. The event argument carries a CrawledPage whose HttpResponseMessage gives you the status code, whose Content.Text gives you the raw body, and whose AngleSharpHtmlDocument gives you a parsed DOM tree:
var httpStatus = e.CrawledPage.HttpResponseMessage.StatusCode;
var rawPageText = e.CrawledPage.Content.Text;
var angleSharpHtmlDocument = crawledPage.AngleSharpHtmlDocument; //AngleSharp parserHaving AngleSharp's document attached at no extra cost is the detail that makes this practical. You do not parse HTML yourself, and you do not add an HTML parser dependency, because parsing already happened on the way in.
The disallowed events matter just as much for a crawl that does not complete. PageCrawlDisallowed fires when a URL is rejected before fetching, and PageLinksCrawlDisallowed carries a DisallowedReason explaining why the links on a page were not followed:
void crawler_PageLinksCrawlDisallowed(object sender, PageLinksCrawlDisallowedArgs e)
{
CrawledPage crawledPage = e.CrawledPage;
Console.WriteLine($"Did not crawl the links on page {crawledPage.Uri.AbsoluteUri} due to {e.DisallowedReason}");
}Being told why a URL was skipped is what lets you debug a crawl that silently returns half the pages you expected.
CrawlConfiguration is where politeness and scale are decided
The README points at CrawlConfiguration.cs for the documentation of each option, and the class is described as having a lot of them. The ones that change behavior most:
var crawlConfig = new CrawlConfiguration();
crawlConfig.CrawlTimeoutSeconds = 100;
crawlConfig.MaxConcurrentThreads = 10;
crawlConfig.MaxPagesToCrawl = 1000;
crawlConfig.UserAgentString = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/74.0.3729.169 Safari/537.36";
crawlConfig.ConfigurationExtensions.Add("SomeCustomConfigValue1", "1111");
crawlConfig.ConfigurationExtensions.Add("SomeCustomConfigValue2", "2222");
etc...Four of those deserve comment. CrawlTimeoutSeconds bounds the whole crawl, which is your protection against a crawler that follows links forever. MaxConcurrentThreads sets how wide the crawl goes, and it is the number most likely to get a crawler blocked if you raise it carelessly. MaxPagesToCrawl is a hard budget. UserAgentString is set explicitly here, which is the polite thing to do and also makes your traffic identifiable as automated.
The ConfigurationExtensions dictionary is the escape hatch: string keys and string values for anything this class does not model, which you can read through your own pluggable components.
The class lives under the Abot2.Poco namespace, and it is worth knowing that the whole library carries a 2 suffix in its namespaces while the NuGet package is simply Abot. So `using Abot2.Core;` and `using Abot2.Crawler;` in the quick start are correct, not a typo.
JavaScript rendering is not here, and that is the main limitation
The README is direct about this. Every capability that most people assume a modern crawler has is delegated to a separate project, AbotX, which the README recommends for more capable work:
- Crawl multiple sites concurrently - Execute/Render Javascript - Avoid getting blocked by sites - Auto Tuning - Auto Throttling - Pause/Resume live crawls - Simplified pluggability/extensibility
That list is effectively the feature gap between this repository and a production crawler. Abot itself fetches HTTP and parses the response; if your target renders its content with JavaScript, you get the unrendered shell and nothing else. The same applies to auto-throttling and pause and resume, which are the features that make long crawls survivable.
The README's own summary of Abot explains the intended split. Abot takes care of the low level plumbing, listing multithreading, HTTP requests, scheduling and link parsing, and you register for events to process page data. AbotX is where the harder operational problems live.
Read that division as deliberate rather than incomplete. A library that ships a headless browser dependency forces every consumer to accept that dependency, and Abot's pitch is the opposite: very lightweight, not over engineered, and with no out of process dependencies, no databases and no installed services.
The trade is real, though. Building a JavaScript-capable crawler on this foundation means implementing the rendering yourself and plugging it in, which is possible precisely because the core interfaces are replaceable.
A mature codebase with an old target framework and a redirected homepage
The README says version 2.0 and above targets .NET Standard 2.0, while versions below 2.0 target .NET Framework 4.0. That statement is old enough to raise a question about current .NET versions, and it is paired with topic tags mentioning netcore2 and netcore3, both of which have been out of support for years. Nothing in the visible documentation updates the target framework story, so if you are on a current .NET release, verify that the package you resolve actually loads before planning around it.
The repository tree is well organized for what it is. Abot2.sln sits at the root with the library in Abot2/, and there are separate projects for a demo, unit tests and integration tests, plus a TestResponses.saz file. That saz file is Fiddler's session archive format, which tells you the integration tests replay recorded HTTP traffic rather than hitting the network, a good sign for test stability.
The build is on Azure Pipelines via azure-pipelines.yml, and there is a msbuild.rsp plus a .nuget directory, indicating a restore-on-build convention that predates current SDK defaults.
The popularity numbers explain why this has survived. There are 544 forks against only 8 open issues, a fork-to-issue ratio that suggests the repository is used far more widely than it is actively triaged. The last push recorded is 2026-07-17, so the project is not dormant, though the README's AppVeyor badge and Google Groups link for questions both point at infrastructure that has since been wound down, and the PayPal donation link suggests a maintainer who has been maintaining this for a long time.
Editorial conclusion
Abot is a good fit when you want a crawler that will not fight you, and a poor fit if you need JavaScript rendering out of the box. The event model is genuinely the simplest part of the design: register handlers, receive a CrawledPage with parsed HTML already attached, and never think about the worker pool. Everything below that surface, from the scheduler to the page requester to the content extractor, is an interface you can substitute, which is how you get at behavior the default implementation does not offer. The boundary worth respecting is that JavaScript execution, parallel multi-site crawling, and auto-throttling all live in a separate AbotX project rather than here, so the feature list you are comparing against is not this library's list. Install the NuGet package, wire up the four events, and read CrawlConfiguration.cs when you need to change how it treats politeness.
Frequently asked questions
What is Abot used for?
Abot is a C# web crawler framework. It handles multithreading, HTTP requests, scheduling and link parsing, and exposes the results through events, so you register a handler for PageCrawlCompleted and receive a CrawledPage with the raw text and an AngleSharp document already parsed. Core interfaces are replaceable, so you can substitute your own scheduler, requester or content extractor.
Can Abot render JavaScript on pages it crawls?
Not in this repository. JavaScript execution and rendering are listed among the capabilities provided by AbotX, a separate companion project. Abot itself fetches and parses HTTP responses, so a page whose content is rendered client-side will come back as the unrendered shell unless you implement rendering yourself and plug it in.
How do I install Abot in a C# project?
Through NuGet, with the command Install-Package Abot in the Package Manager Console. Note that the namespaces carry a 2 suffix, so source files use using directives such as Abot2.Core, Abot2.Crawler and Abot2.Poco even though the package name is Abot.
How do I control how aggressive an Abot crawl is?
Through CrawlConfiguration. MinCrawlDelayPerDomainMilliSeconds sets the wait between requests to one domain, MaxConcurrentThreads sets crawl width, MaxPagesToCrawl sets a hard budget, and CrawlTimeoutSeconds bounds the run. The README notes that Auto Tuning and Auto Throttling live in AbotX rather than here.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/sjdirect-abot)