# recipe-scrapers: turning cooking sites into structured recipe data

> An MIT-licensed Python library with per-site scrapers on top of schema.org parsing, which hands you ingredients and instructions and leaves fetching and bot protection to you.

**hhursev/recipe-scrapers** — Python package for scraping recipes data

- Repository: https://github.com/hhursev/recipe-scrapers
- Website: https://docs.recipe-scrapers.com
- Stars: 2,240 · Forks: 677
- Language: Python
- License: MIT
- Published: 2026-10-06 · Updated: 2026-10-06 · Language: en
- Canonical page: https://hysenlabs.com/projects/hhursev-recipe-scrapers

## A parser first and a scraper library second

The description is short, and it undersells the design: a Python package for extracting recipe data from cooking websites. What matters is where the responsibility boundary sits. The library is exclusively an HTML parser. Fetching, session handling, request scheduling and retry policy all stay with you, and the README says so directly, recommending that you provide both the HTML content and its source domain for the best results.

Given HTML and a URL, it runs through three layers of markup. Standard HTML structure is the fallback, schema.org markup is the main path covering JSON-LD, Microdata and RDFa, and OpenGraph metadata fills in what the rest misses. That ordering is the right one, because schema.org is what Google reads for recipe rich results, and a site that wants those results in search has already emitted machine-readable recipe data.

The output is a consistent set of accessors rather than a dict you have to know the shape of. `title()`, `instructions()`, `to_json()` and `image()` are the visible ones, and the README points you at `help(scraper)` for the complete list. Every method is a call, so the object is explorable in a REPL rather than something you have to read docs to discover.

The result is a library that fits neatly in the middle of a pipeline: your HTTP client, then recipe-scrapers, then whatever you store the data in.

## Installing is one line, and requests is an extra

The install is the shortest possible statement.

```bash
pip install recipe-scrapers
```

The dependency list is deliberately small, and `pyproject.toml` shows exactly three: `beautifulsoup4 >= 4.12.3`, `extruct >= 0.17.0` and `isodate >= 0.6.1`. BeautifulSoup handles the HTML tree, `extruct` extracts the structured metadata across JSON-LD, Microdata and RDFa in one pass, and `isodate` parses the ISO 8601 duration strings that schema.org uses for cooking and prep times. None of these reaches for the network, which is the consistency check on the parsing-only claim.

Anything that does fetch is an optional extra called `online`, which pulls in `requests>=2.31.0`. That split is the clearest signal of the boundary. Install the base package and you get parsing with no HTTP client at all, which is what you want in a service that already has a fetcher with its own retry and rate-limit logic.

The version floor is `>= 3.10`, and the classifiers run through Python 3.14, including 3.10, 3.11, 3.12, 3.13 and 3.14. That floor is recent rather than theoretical: the 15.11.0 release dropped Python 3.9 because it had reached end of life.

Packaging is setuptools with `version` and `readme` marked dynamic, and the author list is three names. The presence of a second and third maintainer matters more than it looks for a library with this many site-specific rules.

## Two entry points, and the difference matters for production

The quick version takes a URL and is the example most people copy.

```python
from recipe_scrapers import scrape_me
```

The example then calls the accessors on the result.

```python
scraper.title()
scraper.instructions()
scraper.to_json()
```

That example works, and the caveat in the README is the part worth internalising: for anything beyond a script, implement your own fetching and pass the HTML in through `scrape_html`, supplying the URL alongside it. The domain argument matters because it is how the library decides which site-specific scraper to apply, and how it resolves relative image URLs.

The offline path has its own documented behaviour. Higher-quality image detection is on by default, meaning that when a page offers several images the library picks the best one rather than the first. You can turn that off per call with `best_image=False`, or globally through the settings module.

```python
from recipe_scrapers.settings import settings
```

A global `BEST_IMAGE_SELECTION` setting can then be set to `False` to keep the site's first image. Global mutable settings are the one design choice here that scales badly, because they are process-wide and make the behaviour of two libraries in one interpreter depend on import order. Prefer the per-call argument if you are writing a service.

The other programmatic entry point is the registry itself, exposed as `SCRAPERS`, and calling `.keys()` on it gives you the full list of supported sites without leaving Python. That is the first call to make when you are evaluating the library.

## What a site-specific scraper is actually for

The supported sites list is long and gets longer, and understanding why it exists is the key to using the library well.

Most sites emit schema.org Recipe markup correctly, and the generic path reads it without help. A site-specific scraper exists when the markup is present but wrong in some way the standard path cannot fix. The instruction list might be one undifferentiated blob instead of separate steps, or split per paragraph with no way to know where a step begins. Times might be written as prose that no ISO parser will touch. Yield might say two loaves in a field that expects a number. Nutrition might be missing, or present as an image rather than data.

When that happens, a scraper subclass overrides the relevant method and returns something usable. The cost is obvious once you state it: a scraper is a promise about one website's markup, and websites redesign. The library handles that churn by sheer contributor volume, and the release notes show the mechanism. 15.10.0 on 2025-11-12 added more than fifty scrapers alongside improvements to normalisation, grouping, equipment handling and image fetching. 15.11.0 on 2025-12-10 added another batch including `howtocook` and `iamafoodblog`. 15.12.0 on 2026-08-08 added cookingclassy, 365daysofbakingandmore, cloudykitchen and more.

Most of that work traces to a single contributor credited by name in all three notes, who was added as an author and a maintainer of the package in 15.12.0. That is the clearest signal in the repository about what happens when a site redesigns and nobody volunteers to fix it: the corresponding issue sits open until it does. 137 open issues is the number to watch for a library whose quality is a function of how much unreviewed HTML exists in the wild.

## Repository plumbing that shows how the site list is maintained

The tree tells you how the site list is produced and kept in sync. At the root sit `generate.py`, `templates/` and `scripts/`, which together describe a generator pipeline rather than a hand-maintained list. `mkdocs.yaml` and `docs/` build the documentation site at docs.recipe-scrapers.com, and `tests/` covers the per-site behaviour.

That combination matters more than it looks. A library with hundreds of site-specific rules is only trustworthy if every site has a test fixture, because that is the only thing standing between a silent markup change and wrong data in your database. The presence of a coverage service in the README badges, alongside a unit test workflow, points the same way.

The testing burden is also the reason contributors cluster their work. A new scraper is a small file plus a test case, and the release notes show exactly that pattern repeating: one pull request per site, batched into a minor release. Nothing about adding a site is architecturally difficult, which is precisely why the project can absorb fifty at a time without a redesign.

For a user, the practical consequence is that the test suite is public. If a scraper you rely on has a fixture, its expected behaviour is written down in the repository, and when it changes you will find out in the diff rather than in your data.

## Where it fits next to a plain extruct call

The fair comparison is with using `extruct` directly. If you only need the raw schema.org blob from a handful of well-behaved sites, calling `extruct` yourself gives you the JSON-LD, Microdata and RDFa in a few lines and no extra abstraction. There is a real class of project for which that is the better answer, and recipe-scrapers is not it.

What you buy is the normalisation. Instructions come back as steps rather than a paragraph blob, durations come back as `isodate` objects rather than strings, and equipment has its own handling after the 15.11.0 cleanup that tidied the mixin usage. You also buy the per-site overrides described above, and the uniform accessor surface that means a new site does not change your calling code.

What you pay is a dependency and an abstraction you cannot fully see. When a result looks wrong you have to decide whether the generic parser, the site scraper or your own fetching is at fault, and the site scraper is code you would otherwise not have at all. The registry means you can always check which scraper was chosen for a URL, which is the first thing to do when output surprises you.

One more limit belongs in the decision rather than the footnote, and the README puts it in bold. The package does not circumvent or bypass bot protection on any site. If you need to collect at scale from sites that actively resist it, this is the wrong layer and no configuration changes that.

The project is MIT-licensed, last pushed on 2026-09-28, has 2240 stars and 677 forks, and lists three authors and three maintainers. The fork count is unusually high relative to stars for a library, which reads as adoption rather than curiosity: people are forking it to add the one site they needed.

## Conclusion

recipe-scrapers fits a specific and common need: turning a URL you already fetched into a title, a list of ingredients and ordered instructions. The schema.org path alone covers most well-marked-up sites, and the per-site scrapers exist to catch the ones where the markup is technically valid but practically unhelpful. Three things to weigh before adopting it. The library deliberately does no fetching, so you supply HTML and a domain and you own the HTTP layer, retries and rate limits, and the README states plainly that it does not circumvent bot protection. Custom scrapers are a per-site maintenance commitment, which is why the 137 open issues and the site-specific contributions in the release history are the number to watch rather than stars. And the version floor moved recently, since 15.11.0 dropped Python 3.9 support on 2025-12-10 and the package now requires 3.10 or later. Start by calling `SCRAPERS.keys()` and checking whether the sites you care about are already covered before writing anything.

## FAQ

### How can I scrape recipes from websites?

Use `recipe-scrapers` with its `scrape_html` entry point. Pass the page HTML you fetched yourself along with the source URL, then call accessors such as `title()`, `instructions()`, `image()` and `to_json()` on the result. The library parses schema.org markup in JSON-LD, Microdata and RDFa form, plus OpenGraph metadata, and falls back to standard HTML structure. It does no fetching itself and does not bypass bot protection, so you own the HTTP layer, retries and rate limits.

### How many cooking websites does recipe-scrapers support?

The supported sites list is long and changes often, with more than fifty scrapers added in the 15.10.0 release alone. Rather than counting, call `SCRAPERS.keys()` at runtime to get the full registry of supported domains. The site list is the fastest-moving part of the project, so treat the registry as the source of truth rather than any documentation snapshot.

### What Python version does recipe-scrapers require?

Python 3.10 or later. `pyproject.toml` sets `requires-python = ">= 3.10"` and the classifiers cover 3.10 through 3.14. The 15.11.0 release dropped Python 3.9 because that version had reached end of life, so older environments need an older release of the package.

### Which third-party libraries does recipe-scrapers depend on?

Three, all MIT-compatible and all parsing-only: `beautifulsoup4 >= 4.12.3` for the HTML tree, `extruct >= 0.17.0` for structured metadata across JSON-LD, Microdata and RDFa, and `isodate >= 0.6.1` for the duration strings used in cooking times. Network fetching is not a base dependency; `requests >= 2.31.0` only appears in the optional `online` extra.

## Sources

- [hhursev/recipe-scrapers on GitHub](https://github.com/hhursev/recipe-scrapers)
- [License: MIT](https://github.com/hhursev/recipe-scrapers/blob/main/LICENSE)
- [Project website](https://docs.recipe-scrapers.com)
- [README](https://github.com/hhursev/recipe-scrapers/blob/main/README.md)
- [Releases](https://github.com/hhursev/recipe-scrapers/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/hhursev-recipe-scrapers
