ScrapeGraphAI: LLM-Driven Scraping Pipelines for Python
Python scraper based on AI
At a glance
- What is it?
- ScrapeGraphAI turns a natural-language prompt and a URL into structured data by routing the page through a LangChain-based graph and an LLM of your choice. It is a good fit when the page structure is unknown or changes often, and a poor fit when you need deterministic, cheap, high-volume extraction.
- Who is it for?
- Adopt ScrapeGraphAI when the target pages are irregular, the schema is not fixed, and you can accept LLM latency and token cost per page. Do not adopt it for high-volume, stable pages where a CSS selector or an XPath already works, or when you cannot send page content to a third-party model.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 22 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 22, 2026, and from our analysis. They are not legal advice.
DEEP OPEN-SOURCE ANALYSIS
What ScrapeGraphAI solves, and for whom
Classic scrapers break when markup changes. A selector written against a product page fails the moment the site renames a class or reorders a table, and someone has to notice and fix it. ScrapeGraphAI takes the opposite approach: you describe the fields you want in a prompt, and the library asks a language model to find them in the page content. The README states the library "uses LLM and direct graph logic to create scraping pipelines for websites and local documents (XML, HTML, JSON, Markdown, etc.)".
The intended user is a Python developer who already knows what data they need but does not want to maintain a selector layer per site. It also suits one-off extraction jobs, where writing a bespoke parser costs more than a few model calls. It is less suited to teams that need byte-identical output across runs, since the model is doing the interpretation.
The graph pipelines and how a page becomes a dictionary
The core abstraction is a graph. Each pipeline is a named graph made of nodes, and the README lists several: SmartScraperGraph for a single page, SearchGraph for the top results of a search engine, SpeechGraph for audio output, and ScriptCreatorGraph for generating a Python script. The repository also ships example directories for CSV, JSON, XML, Markdown, document and omni scrapers, so the pipeline you pick decides the input and output shape rather than the other way round.
The data flow for SmartScraperGraph is: fetch the source with a browser or fetcher, reduce the HTML to something a model can read, send the prompt plus that content to the configured LLM, and parse the response back into a Python dictionary. The configuration object carries the model name, token budget and a JSON format flag, which is what pushes the model toward a parseable answer. The README's own example output is a nested dictionary with a description string, a list of founder objects and a social media links map, which shows the expected shape: keys come from your prompt, not from a schema file you write.
The dependency list in pyproject.toml is worth reading before you commit. It pins langchain 1.2.0 or newer, langchain-openai, langchain-mistralai, langchain-aws, langchain-ollama and langchain_community, plus playwright and undetected-playwright for fetching, html2text and beautifulsoup4 for conversion, and tiktoken for token counting. That is a large surface area for a scraping library, and it means your resolution of LangChain packages is coupled to this project's.
Installing ScrapeGraphAI and running SmartScraperGraph
The README gives a two-step install. The first command pulls the library from PyPI; the second installs the browser binaries that Playwright needs to fetch pages, and the README marks it as important for that reason.
pip install scrapegraphai
# IMPORTANT (for fetching websites content)
playwright installThe README recommends a virtual environment to avoid conflicts with other libraries. Note that pyproject.toml declares requires-python as ">=3.12,<4.0", while the repository's Dockerfile builds on python:3.11-slim and installs the published package. If you follow the Dockerfile rather than the metadata, you are relying on the released artifact's own constraints.
For a first run, the README uses a local Ollama model. The config names the model, a token budget and a JSON format flag, and the graph takes a prompt and a source URL.
from scrapegraphai.graphs import SmartScraperGraph
graph_config = {
"llm": {
"model": "ollama/llama3.2",
"model_tokens": 8192,
"format": "json",
},
"verbose": True,
"headless": False,
}
smart_scraper_graph = SmartScraperGraph(
prompt="Extract useful information from the webpage, including a description of what the company does, founders and social media links",
source="https://scrapegraphai.com/",
config=graph_config
)
result = smart_scraper_graph.run()The README prints the result with json.dumps(result, indent=4) and shows a dictionary containing a description, a founders list and social_media_links. Setting headless to False opens a visible browser window, which is useful while debugging a page that blocks silent requests. To switch providers, the README says you only need to change the llm config, and gives an OpenAI block with an api_key and the model "openai/gpt-4o-mini".
If you would rather run the model locally in a container, the repository includes a docker-compose.yml that starts an ollama service on port 11434 with a named volume for its models.
services:
ollama:
image: ollama/ollama
container_name: ollama
ports:
- "11434:11434"
volumes:
- ollama_volume:/root/.ollama
restart: unless-stoppedWhere the LLM approach costs you: determinism, tokens and anti-bot pages
The trade-off is explicit. A selector-based scraper returns the same value every time the markup is the same. A model-mediated scraper can return a different phrasing, a missing key, or a plausible but wrong value, and the README's own sample output shows an empty name field for one of the founders, which is exactly the kind of gap you have to handle downstream. The "format": "json" flag improves the odds of parseable output; it does not guarantee the content is correct.
Cost scales with page size. Every run sends the reduced page content plus your prompt to the model, and the config carries a model_tokens budget. Long pages, or a pipeline that visits many pages, multiply that. The repository depends on tiktoken, so token accounting exists, but the README does not document a cost estimator or a caching layer for repeated fetches of the same URL.
Fetching is the other weak point. The dependency list includes undetected-playwright, which signals that some targets resist ordinary browser automation. That is an arms race, not a solved problem, and a site that fingerprints automation will still fail. If your target is a heavily protected endpoint, an official API or a paid extraction service will be more reliable than prompting a model over a blocked page.
ScrapeGraphAI compared with Firecrawl and with plain BeautifulSoup
The repository's own topics list firecrawl-alternative, so the comparison is fair game. Firecrawl is a hosted extraction service with a documented API and SDKs; you send a URL and get markdown or structured data back, and the crawling, rendering and retry logic live on their infrastructure. ScrapeGraphAI is a library you run yourself, and the model is yours to choose, including a local Ollama model that keeps page content on your own machine. The difference is operational: Firecrawl moves the browser fleet and scaling problem to a vendor, while ScrapeGraphAI moves the model choice and the data boundary to you.
The other comparison is with BeautifulSoup, which pyproject.toml already depends on. BeautifulSoup parses markup you have already fetched; it has no opinion about which fields matter. ScrapeGraphAI uses it as one step in a larger graph. If your target pages are stable and you can write the selectors in an afternoon, BeautifulSoup plus requests is faster, free per page and fully deterministic. Reach for the graph pipelines when the pages are irregular or the schema is still moving.
Licence, maintenance and what an upgrade actually costs
The project is MIT licensed, which permits commercial use and modification provided the copyright notice and permission notice are retained. That is the standard permissive position; it is not legal advice, and if you redistribute the library inside a product you should read the LICENSE file in the repository rather than this summary.
The repository is not archived, and the last push was on 2026-09-07, with v2.2.4 released the same day and v2.2.3 shortly before. Releases are cut through semantic-release, as the .releaserc.yml and SEMANTIC_COMMITS.md files in the repository root indicate, so version numbers track commit types rather than a hand-written roadmap. The CHANGELOG.md at the root is the place to read before upgrading.
Upgrade cost is dominated by the LangChain dependency floor. Because pyproject.toml requires langchain 1.2.0 or newer alongside several provider packages, a major LangChain change can force a ScrapeGraphAI upgrade even when you did not want one. The Makefile shows the project's own check pipeline: uv sync for install, ruff, black and isort for lint, mypy for types, and pytest with coverage. Running the same commands against your fork is the cheapest way to see whether an upgrade breaks your graphs.
Editorial conclusion
Adopt ScrapeGraphAI when the target pages are irregular, the schema is not fixed, and you can accept LLM latency and token cost per page. Do not adopt it for high-volume, stable pages where a CSS selector or an XPath already works, or when you cannot send page content to a third-party model. Before writing it into a pipeline, verify two things in your own environment: that the model you configure returns valid JSON for your prompt, and that Playwright's browser binaries install cleanly on your build image.
Frequently asked questions
What is ScrapeGraphAI used for?
It is a Python library for building scraping pipelines that use an LLM and graph logic to extract information from websites and local documents such as XML, HTML, JSON and Markdown. You give it a prompt describing the fields you want and a source, and it returns a dictionary.
Is ScrapeGraphAI free?
The library is MIT licensed and installs from PyPI with pip install scrapegraphai, so the code itself is free to use. Running it still costs whatever your chosen model costs: a local Ollama model is free to run on your own hardware, while the OpenAI configuration in the README requires an API key and is billed by that provider.
How much does ScrapeGraphAI cost?
The repository does not publish pricing for the library, which is MIT licensed and free to install. Cost depends on the LLM you configure in graph_config: a local model such as ollama/llama3.2 runs on your own machine, while hosted models are billed by their provider. The README also points to a separate hosted service at scrapegraphai.com for scraping at scale, whose pricing is not described in the README.
How to use ScrapeGraphAI?
Install it with pip install scrapegraphai and then playwright install, build a graph_config dictionary naming your model, and pass a prompt and source URL to a pipeline such as SmartScraperGraph. Calling run() on the graph returns a dictionary, which the README prints with json.dumps.
Can ChatGPT scrape websites?
A chat model on its own has no fetching step, which is why ScrapeGraphAI pairs a browser-based fetcher with the model. The README shows the same pipeline pointed at an OpenAI model by changing the llm config to include an api_key and the model openai/gpt-4o-mini.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/scrapegraphai-scrapegraph-ai)
Community notes