Model or dataset
paulpierre/markdown-crawler avatar
paulpierre/markdown-crawler

markdown-crawler: turn a site into a folder of Markdown files for RAG

A multithreaded 🕸️ web crawler that recursively crawls a website and creates a 🔽 markdown file for each page, designed for LLM RAG

474 stars54 forksPythonMIT

At a glance

What is it?
A multithreaded Python crawler that writes one Markdown file per page, aimed at LLM document pipelines. It installs from PyPI in one command, but the README leaves rate limiting, robots.txt and resume semantics undocumented.
Who is it for?
Adopt markdown-crawler when you need a small, MIT-licensed Python library that turns a bounded set of pages into Markdown files you can chunk by heading, and when you are willing to read the source because the README does not document rate limiting, robots.txt handling or what --base-dir reuse actually skips. Do not adopt it as a general-purpose site mirroring or scraping framework; it has no parser plugin system, no JavaScript execution and no documented politeness controls.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 96 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What markdown-crawler is for, and who reaches for it

The repository describes itself as a multithreaded web crawler that recursively crawls a website and creates a Markdown file for each page. The stated primary purpose is LLM document parsing: the README says it was created so that large documents can be chunked and processed for RAG use cases, on the argument that Markdown is human readable, keeps document structure and has a small footprint.

The audience is narrow and identifiable. You are building a retrieval corpus and you want headings and paragraphs preserved rather than a flat text dump. You have a bounded target: a documentation site, a wiki, a knowledge base, a game or film wiki whose corpus you want to reconstruct. The README lists RAG, fine-tuning corpora, agent knowledge for tools like autogen, and online RAG learning as the intended uses. If you want a general scraper that extracts structured fields from product pages, this is not that tool. It converts pages to Markdown and follows links. That is the whole job.

The crawl loop: requests, BeautifulSoup, markdownify, one file per URL

The dependency list in pyproject.toml names the three moving parts: requests for fetching, beautifulsoup4 for parsing, markdownify for conversion. The flow implied by the README and the CLI surface is: take a base URL, fetch it with a browser-like User-Agent, parse the HTML with BeautifulSoup, optionally narrow the parsed region with a CSS selector supplied through target_content, convert the surviving DOM to Markdown with markdownify, and write it to a file under base_dir whose path mirrors the URL structure.

Depth control is the boundary of the recursion. The README's example sets max_depth to 3, which it describes as the base URL plus three levels of children. Link eligibility is filtered by two switches that read backwards from their names: is_domain_match and is_base_path_match. The README states that setting is_domain_match to False means only pages in the same domain as the base URL are crawled, and is_base_path_match to False means all URLs in the same domain are included even when they do not begin with the base URL. Read that carefully before you trust a filter. valid_paths is the allowlist of relative paths, and exclude_paths, added in v0.0.9 as the --exclude-paths or -x flag, is the denylist. Threading is the throughput mechanism: num_threads sets how many workers run in parallel.

The output layout matters for downstream chunking. Because each page becomes its own file with its heading hierarchy intact, a chunker can split on headings without a separate segmentation pass. That is the real design decision here, and it is a good one for retrieval: a per-page file gives you a natural document boundary that a single concatenated dump does not.

Installing markdown-crawler and running a first crawl

The README gives two installation routes. The published package comes from PyPI, and the project page listed in the repository metadata is the PyPI entry for markdown-crawler. The CLI entry point is declared in pyproject.toml as markdown_crawler.cli:main, exposed under the command name markdown-crawler.

Install from PyPI:

bash
pip install markdown-crawler

Then run the CLI against a small page. The README's own example uses a Wikipedia article with five threads, depth three, and an output directory named markdown:

bash
markdown-crawler -t 5 -d 3 -b ./markdown https://en.wikipedia.org/wiki/Morty_Smith

After the run, the README shows the markdown directory containing one file per crawled page and the converted contents inside. Start with depth one and a single thread on any host you do not own, then raise both once you have seen the file names and confirmed the selector you want.

If you prefer the library, the README's snippet imports md_crawl and passes the same parameters as keyword arguments:

python
from markdown_crawler import md_crawl
url = 'https://en.wikipedia.org/wiki/Morty_Smith'
md_crawl(url, max_depth=3, num_threads=5, base_path='markdown')

Note the naming inconsistency: the CLI flag is --base-dir with the short form -b, while the library parameter in the README example is base_path. The usage block lists --base-dir, and the example.py description in the README says base_dir. Treat the README snippet as the authority for the library call and the usage block as the authority for the CLI, and check the signature before wiring either into a script.

Where it breaks: selectors, politeness and the missing controls

The most likely silent failure is target_content. It accepts CSS selectors, and the README says you can supply several and their results are concatenated. If the selector matches nothing on a given page, you get a file that is empty or nearly empty, and nothing in the documented CLI output tells you the extraction was wrong rather than the page being thin. On a site with heterogeneous templates, that failure is per-page and easy to miss across hundreds of files. Verify the selector on two or three page types before a full run.

The second gap is politeness. The README documents num_threads as the parallelism knob and nothing else. There is no documented delay between requests, no retry policy, no rate limit, and no mention of robots.txt anywhere in the README. A multithreaded crawler pointed at a host without an agreed limit is a load problem, and the documentation gives you no built-in way to be gentle other than lowering the thread count to one. If your target publishes a crawl-delay or forbids crawling, this tool does not enforce that for you.

The third gap is resume semantics. The feature list says you can continue scraping where you left off, and the README shows files persisting in the output directory. What is not documented is the exact rule: whether an existing file causes a skip, whether the directory listing is read at startup, or whether a partially written file is treated as complete. The README does not document rollback either. If you need a crawl you can interrupt and trust, read the source before relying on that feature.

Finally, there is no JavaScript execution. The v0.0.9 release added a browser-like User-Agent specifically to get past JavaScript checks, which is a header-level workaround, not a rendering engine. Pages that build their content client-side will convert to Markdown as whatever the server returned. That is a hard boundary, not a bug.

How it differs from Crawl4AI and from writing your own requests loop

Crawl4AI appears in the related searches for this project, and the comparison is worth making because the two tools sit at different points on the same spectrum. Crawl4AI is a broader crawling and extraction stack aimed at LLM data pipelines, with its own crawler, browser handling and extraction layer. markdown-crawler is deliberately smaller: three dependencies, one conversion path, one output format. The difference in approach is scope. If your pages render client-side, or you need structured extraction beyond Markdown, the larger framework is the direction to look. If your pages are server-rendered and you want files on disk with no service to run, the smaller tool is easier to reason about.

The other alternative is a hand-written loop: requests plus BeautifulSoup plus markdownify is roughly twenty lines, and those are exactly the three dependencies this project lists. What you get by using markdown-crawler instead is the parts that are tedious to get right: depth tracking, URL validation, domain and base path filtering, an allowlist and a denylist, threading, and a CLI. The v0.0.9 release notes list fixes for a UnicodeEncodeError on non-ASCII content, an UnboundLocalError in get_target_content, and null href links being skipped, which is a fair summary of the bugs you would otherwise hit yourself. The trade-off is that you inherit its opinions about output layout and its gaps around politeness, and you cannot swap the parser without forking.

Maintenance, licence and the upgrade cost you should price in

The last push to the default branch was on 2026-06-26, which is also the date of the v0.0.9 release. The repository is not archived. The release history is uneven: 0.0.4 and 0.0.8 landed in October 2023, and 0.0.9 arrived in June 2026 with six issue fixes, a configurable heading style, the exclude-paths flag, and a test suite the README describes as 60 tests at 95 percent coverage. That is a long dormancy followed by a concentrated release, so plan for the project moving in bursts rather than continuously.

Upgrading is low cost in dependency terms. The runtime requirements are beautifulsoup4, markdownify and requests, all widely used, and pyproject.toml sets requires-python to >=3.8. The risk is behavioural rather than structural: v0.0.9 changed how files are written (UTF-8 on all writes), added a User-Agent header to every request, and changed how target_content handles null links. If you have a pipeline that depends on byte-identical output, re-run a sample after upgrading and diff the files.

The licence is MIT, stated in the README and in the classifiers. That permits commercial use, modification and redistribution provided the copyright notice and permission notice are included. It ships with no warranty. That is the licence text, not legal advice; if you are redistributing the tool inside a product, have your own counsel read the notice.

Editorial conclusion

Adopt markdown-crawler when you need a small, MIT-licensed Python library that turns a bounded set of pages into Markdown files you can chunk by heading, and when you are willing to read the source because the README does not document rate limiting, robots.txt handling or what --base-dir reuse actually skips. Do not adopt it as a general-purpose site mirroring or scraping framework; it has no parser plugin system, no JavaScript execution and no documented politeness controls. Before pointing it at a production host, verify three things yourself: whether the target site's terms permit crawling, whether the default thread count is acceptable to that host, and whether the target_content CSS selector you pass actually matches the container you want, because a wrong selector silently produces empty or partial files.

Frequently asked questions

What is the difference between crawling and scraping in markdown-crawler?

The project does both in one pass: it crawls by following links recursively up to max_depth, and it scrapes by parsing each fetched page with BeautifulSoup and converting it to Markdown. The README frames the output as one Markdown file per page, which is the scraping half, while the depth and path filters control the crawling half.

What are the disadvantages of Markdown as the output format for a crawl?

Markdown keeps headings and paragraphs but discards layout and most presentation detail, so anything encoded only in CSS or in a widget will not survive the conversion. The README's argument for the format is that it stays human readable, holds document structure and keeps a small footprint, which is a trade of fidelity for chunkability.

Can markdown-crawler be used to generate Markdown for a chatbot or LLM pipeline?

The README states the project was primarily created for large language model document parsing, so that large documents can be chunked and processed for RAG. It lists RAG, fine-tuning corpora and agent knowledge as intended uses, and the output is one Markdown file per page.

Official sources

  1. License: MIT
  2. paulpierre/markdown-crawler on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/paulpierre-markdown-crawler.svg)](https://hysenlabs.com/projects/paulpierre-markdown-crawler)