llm-scraper: turning any webpage into typed data with a browser and an LLM
Turn any webpage into structured data using LLMs
At a glance
- What is it?
- llm-scraper is a TypeScript library that pairs Playwright with a model provider and a Zod schema to return structured objects instead of parsed HTML. It is a good fit when the page shape changes often; it is the wrong tool when you need deterministic, cheap, high volume extraction.
- Who is it for?
- Adopt llm-scraper when the target pages are few, their markup shifts, and you already pay for a model provider: the Zod schema plus Playwright page object is a smaller surface than a hand-written parser. Do not adopt it for high-volume crawling where per-page token cost and non-determinism matter, or where you cannot send page content to an external API.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 21 days ago.
- What is it written in?
- Mainly TypeScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What llm-scraper solves, and who it is actually for
Classic scraping breaks when a site changes a class name. llm-scraper takes a different route: it loads the page in a real browser, converts it into a representation a model can read, and asks the model to fill a schema you define. The README describes the library as a TypeScript tool that extracts structured data from any webpage using LLMs, and the feature list names GPT, Sonnet, Gemini, Llama and Qwen model series as supported. The audience is narrow and specific. You need Node or a bundler, you need TypeScript to get the type-safety the README advertises, and you need an API key or a local model server. If you are writing a one-off script to pull a table from a page you control, this is a large dependency for a small job. If you maintain a handful of extractors against pages that are redesigned every few months, the schema survives the redesign while a CSS-selector parser does not.
How the browser, the formatter and the schema fit together
The data flow has four moving parts. Playwright opens the page and gives you a page object; that object is what you hand to the scraper. The scraper converts the page into one of six formatting modes listed in the README: html for pre-processed HTML, raw_html for unprocessed HTML, markdown, text via Readability.js, image for a screenshot on multi-modal models, and custom for a function you supply. The converted content goes to the model together with your schema, which is defined with Zod or JSON Schema. The model returns an object, and the library hands it back to you typed. The package metadata shows the runtime dependencies are @ai-sdk/provider, ai and turndown, with playwright and zod sitting in devDependencies, so you install the browser and the schema library yourself rather than getting them transitively. That split is worth noticing: the README's install line asks for zod, playwright and llm-scraper together, which matches what the package.json implies. Version 2.0 moved to Vercel AI SDK 6 support, so the provider packages you import come from the @ai-sdk scope and the call shape follows that SDK.
Installing llm-scraper and running a first extraction
The README's getting-started sequence is three steps: install dependencies, initialize a model provider, then construct the scraper. Start with the packages the README names.
npm i zod playwright llm-scraper
npm i @ai-sdk/openaiThe second command pulls the OpenAI provider; the README gives equivalent installs for @ai-sdk/anthropic and @ai-sdk/google. Next, create the provider and the scraper instance. The README's OpenAI example is the shortest.
import { openai } from '@ai-sdk/openai'
import LLMScraper from 'llm-scraper'
const llm = openai('gpt-4o')
const scraper = new LLMScraper(llm)Now a real extraction. The README's Hacker News example launches Chromium, navigates, defines a Zod schema, and calls run with the Output.object helper and format: 'html'. The schema below asks for exactly five stories, each with a title, points, author and comments URL.
import { chromium } from 'playwright'
import { z } from 'zod'
import { Output } from 'ai'
import LLMScraper from 'llm-scraper'
const browser = await chromium.launch()
const page = await browser.newPage()
await page.goto('https://news.ycombinator.com')
const schema = z.object({
top: z.array(z.object({
title: z.string(),
points: z.number(),
by: z.string(),
commentsURL: z.string(),
})).length(5).describe('Top 5 stories on Hacker News'),
})
const { data } = await scraper.run(page, Output.object({ schema }), { format: 'html' })
console.log(data.top)What you should see is an array of five objects, shaped like the README's sample output with fields such as title, points, by and commentsURL. Close the page and the browser when you are done; the example does both. If you want partial results as they arrive, the README says to replace run with stream and iterate the returned stream with for await.
Streaming and code generation: two features with different costs
The streaming mode is a thin wrapper. You call scraper.stream instead of scraper.run, and the README shows a for await loop over the stream printing data.top on each chunk. That is useful for a UI that fills in rows as the model works, and it changes nothing about the extraction itself. Code generation is the more interesting one. The generate function asks the model to produce a Playwright script that scrapes according to your schema, and the README then evaluates that script on the page, parses the result with schema.parse, and logs it. Read that sequence carefully, because it moves the work from the model to the browser: once you have the generated code, the per-page model call is gone, and the extraction becomes ordinary deterministic DOM traversal. The trade-off is that generated code is only as good as the page it was generated against. The README does not document what happens when the generated script runs against a page whose markup has since changed, and it does not describe a validation step before page.evaluate. Treat generate as a way to bootstrap a parser you then commit and review, not as a runtime path you leave unattended.
Where llm-scraper is the wrong tool
Three limits are visible from the documentation alone. First, cost and latency are per page and per token. Every run sends page content to a model, and the formatting mode you choose decides how much content that is: raw_html sends the page unprocessed, while text runs it through Readability.js first. For a crawl of thousands of pages, that arithmetic rarely works in your favour, and a selector-based scraper or an HTTP-level parser will be orders of magnitude cheaper. Second, results are not guaranteed to satisfy the schema in the way a type system suggests. The README's own example constrains the array with .length(5), but the library's contract is that a model fills the schema; nothing in the README describes a retry or repair loop when the model returns four items or a string where a number was requested. Validate before you persist. Third, the data leaves your machine unless you point the provider at a local server. The README shows an Ollama setup through ollama-ai-provider-v2 and a Groq setup through createOpenAI with a baseURL and an apiKey from the environment, which is the pattern to copy if you need a self-hosted or alternative endpoint. If your pages contain personal data or material you are not permitted to send to a third-party API, this library is the wrong layer and you should extract locally. On the legal side, the README says nothing about permitted use; whether scraping a given site is allowed depends on its terms, robots directives and the law where you operate, which is outside what this repository documents.
How it compares with Crawl4AI and with plain Playwright
Crawl4AI is the comparison people reach for, and the difference is architectural rather than cosmetic. Crawl4AI is a Python crawling framework: it handles fetching and traversal across many URLs and produces markdown aimed at feeding a model. llm-scraper does not crawl. It takes a Playwright page you already opened and returns an object matching a schema you already wrote, in TypeScript, with the model call inside the library. So the split is: if your problem is breadth (discover and fetch many pages, then hand text to a model), a crawler is the right shape. If your problem is depth on a known page (get typed fields out of this one page, repeatedly, in a Node service), llm-scraper fits without a Python sidecar. The second comparison is plain Playwright. A hand-written page.locator call is deterministic, free per run and fast, and it breaks when the markup changes. llm-scraper trades that determinism and cost for tolerance of markup drift. Neither is universally better; the deciding question is how often the target page changes versus how often you run the extraction.
Maintenance, licence and upgrade exposure
The repository is not archived, and the last push was on 2026-09-08, which is recent enough that the project is being worked on. The README marks a 2.0 release with Vercel AI SDK 6 support and updated examples, and package.json pins the package at 2.0.0 with ai at ^6.0.77 and @ai-sdk/provider at ^3.0.8. That is the main upgrade cost to plan for: the library sits on top of the AI SDK, so a major AI SDK release is a migration for you, not just for the maintainer. The provider packages carry their own version lines, and the README's install commands are the place to check which one matches your chosen model. The licence is MIT, declared in both package.json and LICENSE.md, which permits commercial use and modification provided the copyright notice and permission notice are retained; this is a description of the licence text, not legal advice, and if you redistribute the library inside a product you should read LICENSE.md yourself. There is no published release list in the repository, so pinning to 2.0.0 in your own lockfile is the only version statement you can make with confidence.
Editorial conclusion
Adopt llm-scraper when the target pages are few, their markup shifts, and you already pay for a model provider: the Zod schema plus Playwright page object is a smaller surface than a hand-written parser. Do not adopt it for high-volume crawling where per-page token cost and non-determinism matter, or where you cannot send page content to an external API. Before committing, verify the formatting mode that fits your pages (html, text, markdown, image or custom), confirm your provider package matches the AI SDK 6 line the README targets, and check that the schema constraints you rely on, such as array length, actually hold in the returned object.
Frequently asked questions
What is llm-scraper?
It is a TypeScript library that extracts structured data from a webpage using an LLM. You open the page with Playwright, define a schema with Zod or JSON Schema, and the scraper returns an object matching that schema.
Which LLM is best for web scraping with llm-scraper?
The README lists GPT, Sonnet, Gemini, Llama and Qwen model series as supported and gives setup snippets for OpenAI, Anthropic, Google, Groq and Ollama. It does not rank them or report accuracy per model, so the choice comes down to which provider you can call and whether the mode you need is multi-modal.
Is AI data scraping legal?
The README does not address legality or permitted use. Whether scraping a given site is allowed depends on that site's terms, its robots directives and the law where you operate, and the repository documents none of that.
Is web scraping illegal?
The repository takes no position on this. It documents how to extract data with a browser and a model, and leaves the question of whether a particular site permits scraping to its terms and to the law where you operate.
Can ChatGPT do web scraping?
Not by itself, which is the gap llm-scraper fills: Playwright fetches and renders the page, and a model such as gpt-4o fills the schema you define. The README's example uses openai('gpt-4o') as the provider.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/mishushakov-llm-scraper)