Open-source project
IonicaBizau/scrape-it avatar
IonicaBizau/scrape-it

scrape-it: a declarative Node.js scraper with a selector object instead of a parser

đź”® A Node.js scraper for humans.

4,071 stars218 forksJavaScriptMIT

At a glance

What is it?
scrape-it turns a plain JavaScript object of CSS selectors into a structured scrape result. It is a good fit for server-rendered HTML and a poor fit for anything that needs a browser, a crawl queue or retry logic.
Who is it for?
Adopt scrape-it when your target pages are server-rendered and you want the output shape to be the same object you wrote as input. Do not adopt it if the page needs JavaScript to render, if you need a crawl queue, or if you expect the library to retry failed requests; the README says there is only a simple request module and no fancy crawling.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 11 days ago.
What is it written in?
Mainly JavaScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What scrape-it actually solves, and for whom

Most scraping scripts fail not at the download step but at the extraction step. You fetch HTML, then write a chain of querySelector calls, then flatten the results into an object, then repeat that for every field. scrape-it removes the second and third steps by letting you describe the output object directly, with CSS selectors as the values. The README example passes a single object where keys are field names and values are selectors, and the resolved result has the same shape.

The audience is a Node.js developer who needs a handful of fields from a page and does not want to add a parsing framework. The package is published on npm as scrape-it, the licence is MIT, and the repository ships a TypeScript declaration file at lib/index.d.ts alongside the compiled lib/index.js, so typed consumers get the option names without a separate types package. It is not aimed at teams building a general-purpose crawler. The README is explicit that there is no fancy way to crawl pages with it.

The selector object is the whole architecture

There is no schema language and no separate parser configuration. The object you pass to scrapeIt is both the query and the result template. A string value is a selector whose text content is extracted. An object value with a selector key and an attr key pulls an attribute instead of text. Adding how: "html" returns the inner HTML rather than the text. Adding convert runs a function over the extracted value, which the README example uses to turn a date string into a Date object.

Lists are handled by listItem. A field whose value is an object containing listItem queries every matching element and returns an array. If that object also has a data key, each element is scraped recursively with the nested object, which is how the README builds an articles array where every entry has its own createdAt, title, tags, content and classes. A nested tags field with only listItem and no data returns a flat array of strings. Omitting the selector inside a list item reads an attribute from the list element itself, which the example uses to collect class names.

The request layer is deliberately thin. The README states that scrape-it has only a simple request module for making requests. That single sentence explains most of the project's boundaries: no headless browser, no cookie jar configuration described in the README, no retry policy, no concurrency control. The response callback also receives status, so the HTTP status code is available next to data, but the README does not document any built-in handling for non-200 responses.

Installing scrape-it and scraping one page

Installation is a normal npm or yarn install, and the package exposes scrapeIt as its main export. The README gives both package managers:

bash
# Using npm
npm install --save scrape-it

# Using yarn
yarn add scrape-it

A first scrape is a single call with a URL and a selector object. The README's Promise example fetches a page and pulls three fields, one of them an attribute:

js
const scrapeIt = require("scrape-it")

scrapeIt("https://ionicabizau.net", {
    title: ".header h1"
  , desc: ".header h2"
  , avatar: {
        selector: ".header img"
      , attr: "src"
    }
}).then(({ data, status }) => {
    console.log(`Status Code: ${status}`)
    console.log(data)
});

What you should see is the status code printed first, then a data object with title, desc and avatar keys. The avatar value is the src attribute string, not the element text. The same call works with await, and the README's async example destructures only data from the result.

For a list, add a listItem field. This is the shape the README uses for a set of articles, trimmed to the parts that matter for a first run:

js
const { data } = await scrapeIt("https://ionicabizau.net", {
    articles: {
        listItem: ".article"
      , data: {
            title: "a.article-title"
          , createdAt: {
                selector: ".date"
              , convert: x => new Date(x)
            }
          , content: {
                selector: ".article-content"
              , how: "html"
            }
        }
    }
})

The result is data.articles, an array with one entry per matched .article element. If the array comes back empty, the selector did not match the markup the server actually returned, which is the most common first-run failure.

Where scrape-it stops: JavaScript-rendered pages

The README's own FAQ is the clearest statement of the limit. It says you cannot directly parse ajax pages with scrape-it, and then lists three ways out. If the ajax response is JSON, you do not need a scraping library at all; parse the JSON. If the ajax endpoint returns HTML, point scrape-it at the ajax URL instead of the page URL and scrape that response. If the request is too complicated to reverse-engineer, load the page in a headless browser such as Google Chrome, Electron or PhantomJS and then call .scrapeHTML on the HTML you already have.

That third option is the honest answer for single-page applications. scrape-it will not run your JavaScript for you, so a page whose content is injected after load will yield an empty result with no error. The failure is silent, which makes it worse than a crash: your script exits zero and writes empty arrays.

The same thinness shows up in crawling. The README says you can parse a list of URLs from the initial page and then use Promises to fetch each one, or hand the download step to a different crawler and scrape local files. Both are reasonable, and both mean you are writing the queue, the rate limiting and the error handling yourself. There is no robots.txt handling, no delay between requests and no deduplication in the library.

Local files are the one case that is explicitly supported. The FAQ says to read HTML with fs.readFile and pass it to .scrapeHTML. That makes scrape-it usable as a pure parser over HTML you obtained some other way, which is a cleaner separation than most scraping libraries offer.

scrape-it compared with Cheerio and a headless browser

The nearest alternative is Cheerio. Both parse HTML with CSS selectors, and the difference is where the structure lives. Cheerio gives you a jQuery-like API: you load the HTML, then call $(".header h1").text() for each field and assemble the object yourself. scrape-it inverts that. You declare the object, and the library walks it. For a page with three fields the difference is small. For a nested list of articles, each with a date, tags and HTML content, the declarative version stays readable because the output shape is visible in the source, while the Cheerio version accumulates assignment statements.

The cost of that inversion is flexibility. Cheerio lets you branch mid-parse, filter by position, or compute a field from several elements. scrape-it's convert function is the escape hatch, but it operates on one extracted value at a time, so cross-field logic has to happen after the scrape returns.

The other alternative is a headless browser such as Puppeteer or Playwright. That is not really a competitor, it is the tool the README tells you to reach for when the page needs JavaScript. A browser runs the page, waits for the network to settle and hands you the rendered DOM, at the cost of a much heavier runtime and slower per-page throughput. scrape-it is the cheaper option when the server already returns the markup, and the wrong option when it does not.

Maintenance, releases and the MIT licence in practice

The repository is not archived, and the last push was on 2026-09-20. The most recent release is 6.1.15, published the same day, preceded by 6.1.14 on 2026-07-07 and 6.1.13 on 2026-06-29. The version numbers tell you what to expect from an upgrade: these are patch releases on the 6.x line, so the public API described in the README has been stable across them. There is no 7.x migration to plan for in the information available.

The practical upgrade cost is therefore low, but the maintenance cost of your own scraper is not. Selectors are tied to the target site's markup, and no version bump of scrape-it will fix a selector that stopped matching. The package.json test script is simply node test, and the repository has a test directory plus a .travis.yml at the top level, so the test suite runs on plain Node without a test framework wrapper. That is worth knowing if you plan to fork.

The licence is MIT. In practice that means you can use the package commercially, modify it and redistribute it, provided the copyright notice and permission notice are kept. It does not grant any rights over the sites you scrape, and it offers no warranty. Nothing in the repository addresses the legality or the terms of service of the pages you point it at, which is a separate question from the software licence.

Editorial conclusion

Adopt scrape-it when your target pages are server-rendered and you want the output shape to be the same object you wrote as input. Do not adopt it if the page needs JavaScript to render, if you need a crawl queue, or if you expect the library to retry failed requests; the README says there is only a simple request module and no fancy crawling. Before committing, check that the site does not block plain HTTP clients, confirm the selectors against the real markup rather than a browser's rendered DOM, and read the lib/index.d.ts file for the exact option names.

Frequently asked questions

How do I install scrape-it?

The README gives two options: npm install --save scrape-it or yarn add scrape-it. A separate CLI is available as a different package, scrape-it-cli, installed globally.

Can scrape-it parse pages that load content with ajax?

Not directly. The README states that scrape-it has only a simple request module, so you cannot parse ajax pages with it. It suggests requesting the ajax endpoint itself if that endpoint returns HTML, or loading the page in a headless browser and then using .scrapeHTML.

Does scrape-it crawl a whole website?

No. The README says there is no fancy way to crawl pages with scrape-it, and suggests parsing the list of URLs from the initial page and fetching each one with Promises, or using a separate crawler and scraping the downloaded files.

Can scrape-it read HTML from a local file instead of a URL?

Yes. The README's FAQ says to use .scrapeHTML to parse HTML read from local files with fs.readFile.

What is the licence for scrape-it?

The package.json declares the license as MIT, and the repository contains a LICENSE file at the top level.

Official sources

  1. IonicaBizau/scrape-it on GitHub
  2. License: MIT
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/ionicabizau-scrape-it.svg)](https://hysenlabs.com/projects/ionicabizau-scrape-it)