Defuddle: extract a page's main content as Markdown, from the browser or the CLI
Get the main content of any page as Markdown.
At a glance
- What is it?
- Defuddle is a TypeScript library and CLI that strips comments, sidebars, headers and footers from a web page and returns clean HTML or Markdown. It is honest about being a work in progress, and its main trade-off is that it removes fewer uncertain elements than Mozilla Readability does.
- Who is it for?
- Adopt Defuddle if you need cleaned article content plus metadata in a browser extension or a Node pipeline, and you would rather keep an uncertain element than lose it. Do not adopt it if you need a stable, finished extraction API: the README itself warns that Defuddle is very much a work in progress, and there is no documented rollback or migration path between releases.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 3 days ago.
- What is it written in?
- Mainly TypeScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What Defuddle removes, and who needs that
A web page is mostly not the article. Navigation, comment threads, newsletter prompts, related-post rails and footers all sit in the same document as the text you wanted, and anything downstream of a naive fetch inherits all of it. Defuddle takes a URL or an HTML string, finds the primary content, and returns cleaned HTML or Markdown. The README defines the name plainly: to remove unnecessary elements from a web page, and make it easily readable.
The project was built for Obsidian Web Clipper, the browser extension that saves pages into an Obsidian vault. That origin explains the output choices: a clipper needs Markdown, footnotes that survive the trip, code blocks that keep their fences, and metadata like title, author and published date to write into frontmatter. The README states Defuddle is designed to run in any environment, so the same logic is available to a Node script or a terminal one-liner. If you are building a read-later tool, an archive, a research pipeline, or a note-taking workflow, you are the intended user. If you need to scrape structured product data or tables, this is the wrong shape of tool: it is built to find prose.
How the extraction works: mobile styles, metadata, and three bundles
The mechanism the README describes is a scoring and removal pass over the DOM. Defuddle looks at a page's mobile styles to guess which elements are unnecessary, which is an unusual signal: a responsive stylesheet often hides exactly the furniture a reader does not want, so the CSS becomes evidence about what is chrome and what is content. It also extracts more metadata than a typical extractor, including schema.org data, and returns that raw schemaOrgData alongside parsed fields.
The output object carries author, content, description, domain, favicon, image, language, metaTags, parseTime, published, site, title, wordCount and, when debug is on, a debug object with the content selector and the removals. That debug field is the most useful part for anyone tuning extraction, because it tells you which selector won and what got dropped rather than leaving you to guess.
There are three bundles, and picking the wrong one is the most common integration mistake. The core bundle (defuddle) is for the browser and has no dependencies. The full bundle (defuddle/full) adds math equation parsing and Markdown conversion, using mathml-to-latex and temml for MathML-to-LaTeX fallbacks. The Node bundle (defuddle/node) accepts any DOM Document from linkedom, JSDOM or happy-dom and includes the full math and Markdown capabilities. The README recommends the core bundle for most use cases and notes it still handles math content, just without the conversion fallbacks.
Installing Defuddle and parsing your first page
The library installs from npm. For Node usage you also need a DOM implementation; the README shows linkedom and JSDOM as the two options.
npm install defuddle
npm install linkedomTo use the defuddle command globally, install with the -g flag, or skip installation and run it through npx.
npm install -g defuddle
npx defuddle parse https://example.com/articleThe CLI accepts a file path, a URL, or HTML piped over stdin. Running it against a local file prints cleaned HTML by default; adding --markdown switches the output format, and --json returns metadata and content together.
npx defuddle parse page.html --markdown
npx defuddle parse page.html --json
curl -L https://stephango.com/saw | npx defuddle parse --markdownIn Node, import the node bundle rather than the default one. The README is explicit that for defuddle/node to import properly, the module format in your package.json has to be set to type module.
import { parseHTML } from 'linkedom';
import { Defuddle } from 'defuddle/node';
const { document } = parseHTML(html);
const result = await Defuddle(document, 'https://example.com/article', {
markdown: true
});
console.log(result.content);
console.log(result.title);What you should see is a result object whose content holds the cleaned article and whose title, author and published fields are filled in when the page exposes them. If a field comes back empty, the page did not publish it in a form Defuddle recognizes, which is a page problem rather than a parse failure.
Where Defuddle gives up, and the 403 problem
The README opens with a warning: Defuddle is very much a work in progress. Treat that as a real constraint rather than boilerplate. Extraction heuristics are tuned against real pages, and real pages change constantly, so an extractor that is still moving will occasionally return a page's comment section as the main content or drop a paragraph that mattered. The debug option exists precisely because this happens.
There is a second, more mundane failure mode. Some sites return 403 to the default User-Agent, and the CLI exposes --user-agent for that case. The README's example passes a full Safari UA string. This is a workaround, not a fix: if a site blocks automated requests by policy, spoofing the header is a decision you make, not something Defuddle resolves for you.
The third boundary is the one the project states about itself. Defuddle is more forgiving than Mozilla Readability and removes fewer uncertain elements. Forgiving means it errs toward keeping things. If your pipeline needs a tight, predictable body of text and you would rather lose a stray caption than carry a sidebar, that default works against you. There is no documented rollback or version-pinning guidance in the README either, which matters for a library at 0.19.x where minor releases can change extraction behavior.
Defuddle versus Mozilla Readability
The README positions Defuddle as a possible replacement for Mozilla Readability and then lists the differences, which is more useful than a feature table. The first is tolerance: Defuddle removes fewer uncertain elements, so it keeps more of the page. The second is output consistency, specifically for footnotes, math and code blocks, which Readability does not normalize in the same way. The third is the mobile-styles signal, which Readability does not use. The fourth is metadata breadth, including schema.org data.
Those differences point in one direction. Readability optimizes for a clean article body and has been the default in reader modes and Firefox for years; Defuddle optimizes for fidelity, because its first consumer is a clipper writing into someone's permanent notes. If you are rendering a distraction-free reading view, Readability's aggressiveness is a feature. If you are archiving a page you may never revisit, keeping the footnote and the code fence intact matters more than a tidy DOM. Choose on that axis rather than on which one is newer.
Licence, releases, and what maintenance costs you
Defuddle is MIT licensed. That permits commercial and closed-source use, modification and redistribution provided the copyright notice and permission notice are preserved; the LICENSE file in the repository root is the authoritative text, and this is a description of the licence rather than legal advice. The practical implication is that embedding Defuddle in a proprietary product carries no copyleft obligation.
The release cadence visible in the repository is steady: 0.19.4 on 2026-09-17, 0.19.3 on 2026-08-22, and 0.19.2 on 2026-07-22. The last push to the default branch was on 2026-09-20. The version numbers stay in the 0.x range, which conventionally signals that the API is not yet frozen, and the README's own warning agrees. For an upgrade budget, that means pinning a version and reading the CHANGELOG before bumping, because extraction output is the kind of thing that changes silently and shows up as missing paragraphs in your archive rather than as a build error.
There is also a repository script named size, run through scripts/check-bundle-size.mjs, which suggests bundle size is tracked deliberately. That is relevant if you ship the core bundle into a browser extension, since the README advertises it as dependency-free.
A one-line clipper from the terminal
The most direct use of Defuddle outside a browser is piping fetched HTML into the CLI. The README's own example fetches a page with curl and pipes it to npx defuddle parse with --markdown, which produces Markdown on stdout. Adding --frontmatter prepends YAML frontmatter with fields like title, author and source, which is the shape Obsidian and most static site generators expect. Adding --output writes to a file instead of stdout, and --property pulls a single field such as title or domain when you only need one value.
npx defuddle parse https://example.com/article --markdown --frontmatter --output note.md
npx defuddle parse page.html --property titleThe second command is the useful one for scripting: it prints just the title, so you can use it in a shell pipeline without parsing JSON. Note that --lang takes a BCP 47 code such as en, fr or ja, and is described as a preferred language rather than a filter, so it guides extraction rather than restricting it.
Editorial conclusion
Adopt Defuddle if you need cleaned article content plus metadata in a browser extension or a Node pipeline, and you would rather keep an uncertain element than lose it. Do not adopt it if you need a stable, finished extraction API: the README itself warns that Defuddle is very much a work in progress, and there is no documented rollback or migration path between releases. Verify first that your DOM implementation works under defuddle/node, that your package.json is set to type module, and that the sites you target do not return 403 to the default User-Agent, since the CLI exposes a --user-agent flag specifically for that case.
Frequently asked questions
How do I install Defuddle?
Install the library with npm install defuddle, and for Node usage also install a DOM implementation such as linkedom or jsdom. For the command line, use npm install -g defuddle or run it without installing via npx defuddle parse.
How do I use Defuddle on a web page?
Run npx defuddle parse followed by a URL, a file path, or HTML piped over stdin. Add --markdown for Markdown output, --json for metadata plus content, and --frontmatter to prepend YAML fields such as title, author and source.
How does Defuddle compare with Mozilla Readability?
The README describes Defuddle as more forgiving and as removing fewer uncertain elements, with consistent output for footnotes, math and code blocks, plus a mobile-styles signal and more metadata including schema.org data. Readability strips more aggressively, which suits a reading view rather than an archive.
What alternatives to Defuddle exist?
Mozilla Readability is the alternative the README names, and it is the one Defuddle is designed to be able to replace. The difference is in how much is removed: Defuddle keeps more uncertain elements and normalizes footnotes, math and code blocks.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/kepano-defuddle)