x-ray: a composable scraper for pulling structured records out of HTML
The next web scraper. See through the <html> noise.
At a glance
- What is it?
- x-ray is a Node.js scraping library built around a jQuery-like selector string with an @attribute suffix. It is small, MIT licensed, and its last push was on 2026-08-31, but the released version 2.3.4 dates from 2019-07-07.
- Who is it for?
- Reach for x-ray when you already have a Node service and want a page's list items turned into JSON with a two-line schema, especially if you need pagination with delay, throttle and limit controls. Skip it when the data only appears after client-side JavaScript runs, because the default HTTP driver returns what the server sends and the PhantomJS driver is a separate package.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 30 days ago.
- What is it written in?
- Mainly JavaScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem x-ray solves, and who ends up using it
Most scraping code starts as a loop over a parsed DOM. You fetch a page, load it into cheerio, walk a list of nodes, pull a title out of one child and an href out of another, and push a plain object into an array. That loop is where the interesting bugs live: the selector for the third field is subtly wrong, the attribute lives on a sibling, the page has thirty results and you only wanted five, and the next page link is a relative URL. x-ray takes the loop and turns it into a schema you declare once.
The audience is Node developers who need a bounded set of records from a page they can already see in a browser. The README's opening example points at a blog index, selects .post, and asks for an object with a title from h1 a and a link from .article-title@href, then follows .nav-previous a@href for three pages and writes results.json. That is the whole shape of the tool: a URL, a scope selector, a schema, and a chain of modifiers. If you are comfortable writing that, you are the target reader.
It is not a general crawling framework. There is no queue, no robots.txt handling, no retry policy, and no persistence beyond streaming to a file. The README describes the crawl as a breadth-first walk from one page to the next, which is a description of pagination, not of site-wide discovery.
How the selector string and the schema fit together
The mechanism is a small selector language layered on cheerio. A selector is a jQuery-like string, and if you append @attribute the library reads that attribute instead of the element's innerText. So title returns text, img.logo@src returns the src value, and body@html returns inner HTML. The README lists exactly those four cases as the examples, and the default when no attribute is given is innerText.
The second piece is the schema. You can pass a string, an array, an array of objects, or nested objects, and the README is explicit that the schema is not tied to the structure of the page. That matters because it means you can flatten a nested card layout into a flat record, or group fields the way your downstream consumer wants them. A two-argument call xray(url, scope, selector) applies the scope first, which the README compares to $(scope).find(selector) in jQuery.
Collections have their own rule, and it is the kind of detail that trips people up. x('ul', 'li') selects only the first list item. x('ul', ['li']) selects all of them. Collections of collections work the same way: x(['ul'], ['li']) selects every list item in every list. If your output has one record where you expected twenty, the missing brackets are the first thing to check.
The same call signature accepts raw HTML in place of a URL, which is how the README's Pear example works. That is useful for testing a schema against a saved fixture instead of hitting the network on every run.
Installing x-ray and running a first scrape
The README gives one installation line. It is a normal npm package, so the dependency lands in your project's node_modules and you require it by name.
npm install x-rayThe package.json declares engines of node >= 6.0.0, so the runtime floor is old, but the dependency list is not: cheerio is pinned to ~0.22.0 and bluebird to ^3.4.7. Expect the cheerio version to be the constraint that bites first if you also parse HTML elsewhere in the same process.
A first real use is the README's own example. It scrapes a blog index, takes the title and link from each post, follows the previous-page link, stops after three pages, and writes the accumulated records to results.json.
var Xray = require('x-ray')
var x = Xray()
x('https://blog.ycombinator.com/', '.post', [
{
title: 'h1 a',
link: '.article-title@href'
}
])
.paginate('.nav-previous a@href')
.limit(3)
.write('results.json')The array around the object is what makes this return a list rather than a single record. Drop the brackets and you get one post. What you should see after the process exits is results.json containing the records from up to three pages, written incrementally as each page is scraped, which is why the README notes that an error on a later page does not lose what came before.
If you would rather handle the data in memory, .then() is promisified and .stream() returns a readable stream. The README's Express example pipes x('http://google.com', 'title').stream() straight into the response. The README warns that then() must be the last call in the chain, because the other methods are not promisified.
Pacing, pagination and where the chain can go wrong
The modifiers are where x-ray earns its place over a hand-written loop. .paginate(selector) picks a URL out of the current page and visits it. .limit(n) caps how many pagination requests happen. .delay(from, [to]) waits between requests, accepting either a millisecond number or a string like '1s' or '10s'. .concurrency(n) sets parallel requests and defaults to Infinity, which is worth reading twice: out of the box, nothing stops x-ray from opening as many requests as the chain produces. .throttle(n, ms) caps requests to n per ms window, and .timeout(ms) bounds each request.
The failure mode that the API invites is an unbounded chain. Because concurrency defaults to Infinity and limit is opt-in, a paginate selector that matches a link on every page will keep going until the site stops offering one or the process dies. The README presents these controls as a way to scrape responsibly, which reads as an acknowledgement that the defaults are not conservative. Set concurrency and limit explicitly on any run against a host you do not own.
.abort(validator) is the other escape hatch. The validator receives the result object for the current page and the next URL, and pagination stops when it returns true. That is the right place to stop on a repeated page, a login wall, or a result set that has stopped changing. The README does not document rollback or resumption after an abort, so a long crawl that aborts leaves you with whatever was already streamed to disk and no built-in way to continue from there.
Drivers are the last piece. xray.driver(driver) swaps the request layer. The README names two: request-x-ray, for setting headers, cookies or HTTP methods, and x-ray-phantom, for rendering pages or interacting with elements created by JavaScript. Both are separate packages. The README also says the author would like to see a Tor driver in the future, which is a statement of intent, not a feature.
When x-ray is the wrong tool
The default driver fetches HTML over HTTP. If the records you want are rendered by client-side JavaScript, the response x-ray parses will not contain them, and no selector will fix that. The README's answer is the PhantomJS driver, but that is a separate install and a heavier runtime, and it is worth checking whether that package has kept pace with the x-ray version you are pinning.
The second limitation is the selector language itself. It is jQuery-like, not CSS, and it does not cover every extraction shape. If you need XPath, or you need to join two nodes by a shared attribute value, or you need to compute a value from several elements, you are back to writing code around the library. The README's examples are all single-field extractions from a scoped node.
Third, the release history is thin. The most recent release is 2.3.4, dated 2019-07-07, and before that 2.1.0 from 2016-03-25. The repository's last push was on 2026-08-31, so there is activity on master, but a user adopting this today should treat the published API as settled rather than evolving. That is fine for a scraping script you write once. It is less fine if you were hoping for a stream of fixes to selector edge cases.
Finally, x-ray has no opinion about whether you are allowed to scrape a given site. It will fetch whatever URL you give it. That judgement is yours, and so is the rate at which you do it.
How x-ray differs from a hand-rolled cheerio script
The honest alternative is not another scraping library. It is the twenty lines of cheerio you would write yourself. Load the page, select the list, map over the nodes, build objects, follow the next link in a loop. You get total control, no dependency on x-ray-parse, and no surprise about what a selector means.
The difference in approach is where the declarative boundary sits. A cheerio script puts the traversal in imperative code and the data shape emerges from it. x-ray puts the data shape in the schema and hides the traversal. That trade pays off when the same extraction runs against several similar pages and you want to change the output structure without rewriting the walk. It costs you when the page needs logic that does not fit a selector string, because then you are mixing a declarative chain with imperative escape hatches and the result is harder to read than either alone.
There is a second difference in the pagination chain. Doing paginate, limit, delay, throttle and abort by hand is genuinely tedious, and getting the streaming-to-file behaviour right so a mid-crawl error does not discard earlier results takes real care. If pagination is central to your task, x-ray's chain is the part that justifies the dependency. If you are scraping a single page once, it is not.
Licence and the cost of keeping it in your tree
x-ray is MIT licensed, and the LICENSE file sits at the repository root. MIT is permissive: you can use it in closed-source software, and the main obligation is preserving the copyright notice and licence text in what you distribute. This is a description of the licence, not legal advice; if the scraping output feeds a product with its own compliance constraints, that is a question for your counsel, not for the package metadata.
The upgrade cost is unusual. With the latest release at 2.3.4 from 2019-07-07, there is no migration treadmill to keep up with, but there is also no expectation that a reported selector bug gets a patch release. The practical risk is dependency drift rather than API churn: cheerio ~0.22.0 and bluebird ^3.4.7 are the pinned pieces, and if another part of your stack needs a newer cheerio, you may end up with two copies in the tree. Check that before you add x-ray to a project that already parses HTML.
On maintenance: the repository is not archived, and the last push was on 2026-08-31. That is recent activity on the default branch. It is not the same thing as a release cadence, and the two should not be confused when you are deciding how much to depend on the published package.
Editorial conclusion
Reach for x-ray when you already have a Node service and want a page's list items turned into JSON with a two-line schema, especially if you need pagination with delay, throttle and limit controls. Skip it when the data only appears after client-side JavaScript runs, because the default HTTP driver returns what the server sends and the PhantomJS driver is a separate package. Before you commit, check three things: that the site's terms permit scraping, that x-ray-parse still handles the selector syntax you need, and what your own paginate and abort chain does when a page returns a redirect instead of the next link.
Frequently asked questions
How do I install x-ray?
The README gives a single command, npm install x-ray, which adds the package to your project's dependencies. The package.json declares engines of node >= 6.0.0, so check your Node version before installing.
How do I make x-ray return all list items instead of just the first one?
Wrap the selector in an array. The README states that x('ul', 'li') selects only the first list item, while x('ul', ['li']) selects all of them, and x(['ul'], ['li']) selects every list item in every list.
Does x-ray scrape pages that build their content with JavaScript?
Not with the default driver, which fetches HTML over HTTP. The README lists a PhantomJS driver, x-ray-phantom, for rendering pages or interacting with elements created dynamically by JavaScript, and it is a separate package you install and pass to xray.driver().
What is the default concurrency in x-ray?
The README states that xray.concurrency(n) defaults to Infinity, so a chain that paginates will keep opening requests unless you set a limit. Use concurrency, throttle, delay and limit together to bound a run.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/matthewmueller-x-ray)