Mozilla Readability: extracting article text from web pages in Node.js and the browser
A standalone version of the readability lib
At a glance
- What is it?
- The library behind Firefox Reader View is published on npm as @mozilla/readability. It scores DOM nodes to isolate the main article, returns structured metadata, and leaves sanitizing and fetching to you.
- Who is it for?
- Adopt Readability if you already have a DOM and need the article body plus metadata, and you are prepared to sanitize the output yourself; its Firefox Reader View lineage is the strongest signal of what it handles well. Do not adopt it if you need fetching, JavaScript rendering, or sanitization in the same package, because the README explicitly delegates those elsewhere.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 57 days ago.
- What is it written in?
- Mainly JavaScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What @mozilla/readability extracts, and what it refuses to do
Readability.js is the standalone version of the readability library used for Firefox Reader View. That sentence from the README defines the scope precisely: the package does one job, turning a DOM document into an article object. It does not fetch pages, it does not render JavaScript, and it does not sanitize. The security section states that sanitizing unsafe content out of the input is explicitly not something the project aims to do, and points at DOMPurify and CSP instead, noting that the Firefox integration uses both.
That division of labour is the main thing to understand before adopting it. A typical pipeline is: fetch HTML, build a DOM with jsdom, hand the document to Readability, then sanitize the returned content string before it reaches a browser or a database. If you expected a one-call URL-to-article function, this is not it, and the missing pieces are not oversights but stated policy. The audience is developers building reader modes, archive pipelines, feed builders, or text analysis jobs who already have a DOM and want the article body separated from navigation, ads and comment threads.
How the parse() call turns a DOM into an article object
The API is two constructors and one method. new Readability(document, options) wraps a DOM document, and parse() returns an object with title, content, textContent, length, excerpt, byline, dir, siteName, lang and publishedTime. The options object is entirely optional and carries the tuning knobs: charThreshold defaults to 500 characters, below which no result is returned; nbTopCandidates defaults to 5 and controls how many top candidates are compared; maxElemsToParse defaults to 0, meaning no limit; keepClasses defaults to false, so classes are stripped except those listed in classesToPreserve.
Two options reveal how the extraction works internally. linkDensityModifier adds a number to the base link density threshold during what the README calls the shadiness checks, which is the heuristic that penalizes nodes full of links, the signature of navigation and related-post blocks. disableJSONLD controls metadata extraction: by default Readability gives precedence to Schema.org fields in JSON-LD, and setting it to true skips that parsing. The serializer option controls how content is produced from the root element, defaulting to el => el.innerHTML; passing the identity function el => el returns a DOM element instead of a string, which matters if you plan to process the tree further rather than serialize it.
One behavioural detail deserves emphasis because it causes real bugs. The README states that parse() works by modifying the DOM and removes some elements from the page. If you hold a reference to the original document, clone it first with document.cloneNode(true) and pass the clone to the constructor.
Installing @mozilla/readability and parsing your first page in Node.js
The package is on npm under the scoped name. Node.js does not ship a DOM, so the README points at external libraries and uses jsdom in its example. Install both, then build a document and parse it.
npm install @mozilla/readability jsdomThe README's Node.js example constructs a JSDOM instance with a url option, then passes doc.window.document to the constructor. The url option is not cosmetic: it lets Readability convert relative URLs for images and hyperlinks into absolute ones.
var { Readability } = require('@mozilla/readability');
var { JSDOM } = require('jsdom');
var doc = new JSDOM("<body>Look at this cat: <img src='./cat.jpg'></body>", {
url: "https://www.example.com/the-page-i-got-the-source-from"
});
let reader = new Readability(doc.window.document);
let article = reader.parse();After this runs, article is either an object with the properties listed above or null when the content falls under charThreshold. A real HTML page is the next step, and the same pattern applies: fetch the markup, hand it to JSDOM with the page's URL, parse. The README warns that jsdom can execute scripts and fetch remote resources, that these are disabled by default, and that you should keep them that way for untrusted input.
In a browser the setup is shorter, because a document already exists. For web-based projects the README says to load the Readability.js script from your webpage and use a document reference you already have, such as one fetched via XMLHttpRequest or from a same-origin iframe you can access.
isProbablyReaderable as a cheap gate before the expensive parse
Before parsing, you can ask whether parsing is likely to pay off. isProbablyReaderable(document, options) returns a boolean, and the README describes it as a quick-and-dirty check that is likely to produce both false positives and false negatives. It exists to avoid bogging down a time-sensitive process, such as loading and showing a webpage, with the complex logic in the core of Readability.
Its options are minContentLength, default 140, the minimum node content length used to decide; minScore, default 20, the minimum cumulated score; and visibilityChecker, defaulting to isNodeVisible. The documented usage is to instantiate Readability only when the check passes.
/*
Only instantiate Readability if we suspect
the `parse()` method will produce a meaningful result.
*/
if (isProbablyReaderable(document)) {
let article = new Readability(document).parse();
}The honest framing in the README is worth taking at face value. This function is a heuristic gate, not a guarantee, and the project invites improvements to its logic provided performance does not deteriorate. If your pipeline can afford the full parse, the gate buys you little; if you are deciding whether to show a reader-mode button in a page that is already loading, it is the intended tool.
Where Readability is the wrong tool: untrusted input and non-article pages
The first limitation is security, and the project states it plainly. If you use Readability with untrusted input, in HTML or DOM form, the README strongly recommends a sanitizer library like DOMPurify to avoid script injection when you use the output, plus CSP for defense in depth. Readability will happily return content containing whatever the page contained. Treat the content field as untrusted until you have sanitized it.
The second limitation is scope. Pages that are not articles, such as product listings, forums, documentation indexes or single-page applications that render content with JavaScript, are outside what a DOM-based extractor can see if the markup never contained the text. Readability operates on a document you supply; whatever is missing from that document is missing from the result.
The third is the threshold behaviour. charThreshold defaults to 500 characters, so short pages return nothing. That is a sensible default for articles and a poor one for short notes, and the option exists precisely because the right value depends on your corpus. There is also no documented rollback or undo for the DOM mutation parse() performs, which is why the clone advice in the README is the practical mitigation rather than a recovery step.
Readability.js against a full readability service or a readability formula
The name collides with two unrelated things that dominate search results: readability scores such as Flesch-Kincaid or Fry graphs, which measure how hard a text is to read, and hosted readability services that fetch and clean pages for you. Mozilla Readability does neither. It does not compute a reading level, and it does not fetch anything.
A hosted extraction service is the closest functional alternative, and the difference is architectural rather than cosmetic. A service owns the fetching, the headless browser, the proxy rotation and the sanitization, and hands you clean text over HTTP. Readability gives you a library that runs inside your own process against a DOM you built. That means no network dependency and no per-page cost beyond your own compute, but it also means you own jsdom, you own the sanitizer, and you own the failure cases. If you are already running a crawler, the library fits into it; if you have no crawler, a service removes work that Readability deliberately leaves to you. The same distinction applies to readability formulas: those tools answer how difficult a passage is, while this package answers which part of the page is the passage.
Maintenance, Node.js support and the Apache-2.0 licence
The repository is not archived, and the last push was on 2026-08-04, which places it within the last two months. The package.json declares version 0.6.0, an engines field of node >=14.0.0, and a test script running mocha over test/test-*.js. There is a CHANGELOG.md, a SECURITY.md and a CONTRIBUTING.md at the top level, plus index.d.ts for TypeScript consumers and index.js as the main entry point. The codebase is small: Readability.js, Readability-readerable.js and JSDOMParser.js at the root, with tests in test/.
Upgrade cost is low by construction. The public surface is one constructor, one parse method, one readerable check and a documented options object. There is no plugin system and no configuration file to migrate, so version bumps are unlikely to require changes beyond the options you set. The 0.x version number is the honest signal here: the API has not been declared stable at 1.0, and the README's invitation to improve isProbablyReaderable suggests its behaviour can shift.
The licence is Apache-2.0, with a NOTICE file and the copyright line reading Copyright (c) 2010 Arc90 Inc. Apache-2.0 permits commercial use and modification and requires that you retain the licence and notice. This is a description of the licence text, not legal advice; if you redistribute the library or a modified version, read LICENSE.md and NOTICE and confirm the attribution your product needs.
Editorial conclusion
Adopt Readability if you already have a DOM and need the article body plus metadata, and you are prepared to sanitize the output yourself; its Firefox Reader View lineage is the strongest signal of what it handles well. Do not adopt it if you need fetching, JavaScript rendering, or sanitization in the same package, because the README explicitly delegates those elsewhere. Before committing, verify two things: that your Node.js version satisfies the engines field (node >=14.0.0), and that isProbablyReaderable returns true for the pages you care about, since it is documented to produce both false positives and false negatives.
Frequently asked questions
What is @mozilla/readability?
It is a standalone version of the readability library used for Firefox Reader View, published on npm as @mozilla/readability. You create a Readability object from a DOM document and call parse() to get the article title, content, textContent, byline, siteName and other metadata.
How do I install and use Mozilla Readability in Node.js?
Install @mozilla/readability and a DOM library such as jsdom, build a JSDOM instance passing the page URL as the url option, then pass doc.window.document to the Readability constructor and call parse(). The url option lets Readability convert relative image and hyperlink URLs to absolute ones.
Does Mozilla Readability sanitize the HTML it returns?
No. The README states that sanitizing unsafe content out of the input is explicitly not something the project aims to do, and strongly recommends a sanitizer library like DOMPurify plus CSP when you use the output with untrusted input.
Why does parse() return null or nothing for some pages?
The charThreshold option defaults to 500, the number of characters an article must have in order to return a result. Pages shorter than that threshold produce no result unless you lower the option.
Does parse() change the document I pass in?
Yes. The README states that parse() works by modifying the DOM and removes some elements from the web page, and suggests passing a clone made with document.cloneNode(true) to the constructor to avoid that.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/mozilla-readability)