metascraper: unified metadata extraction from any URL
Scrape metadata from any URL using Open Graph, JSON-LD, HTML meta tags, and smart fallbacks.
At a glance
- What is it?
- metascraper is an MIT-licensed Node.js library that pulls author, date, image, logo and other fields out of a page using Open Graph, JSON-LD, Twitter Cards and HTML fallbacks. It is a parsing layer, not a fetcher, and that distinction shapes everything about how you deploy it.
- Who is it for?
- Adopt metascraper if you already control HTML retrieval and want deterministic, rule-based field extraction without a hosted dependency. Skip it if you need proxy rotation, paywall handling or bot-detection evasion, since the README points those users at the managed Microlink API instead.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 13 days ago.
- What is it written in?
- Mainly HTML, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 24, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The gap metascraper fills between a raw page and usable fields
Fetching a URL is easy. Deciding which of the many competing metadata sources on that page should win is not. A single article can carry an og:title, a twitter:title, a JSON-LD headline, a schema.org Microdata property, an RDFa attribute and a plain <title> element, and they frequently disagree. metascraper exists to resolve that disagreement into one object with a stable shape.
The README states the design principles directly: high accuracy for online articles by default, simple addition or overriding of rules, and no restriction of rules to CSS selectors or text accessors. That last point matters more than it reads. A rule in metascraper is not a selector string; it is a function that receives the parsed DOM and the URL and returns a value. That is what lets the library consult JSON-LD and RDFa, which do not map cleanly onto selector syntax.
The intended audience is a developer building link previews, unfurl cards, archive tooling or content pipelines in Node.js. The README notes that metascraper is a collection of tiny packages, so you install only the extractors you need rather than a monolithic parser. If you are working in another language, this is not the tool for you; the library ships as npm packages.
Two inputs, one output: the actual data flow
metascraper requires exactly two inputs: the target URL and the HTML markup behind that URL. It does not fetch anything. That separation is deliberate, and the README is explicit that the markup needs to be as accurate as possible, which is why the maintainers built html-get, a companion package that uses a headless browser to retrieve rendered HTML.
The flow is: retrieve HTML with your own client, hand it plus the URL to a metascraper instance, receive a flat object of fields. The instance is constructed from an array of rule bundles, and each bundle contributes one or more fields. In the README example the array contains metascraper-author, metascraper-date, metascraper-description, metascraper-image, metascraper-logo, metascraper-publisher, metascraper-title and metascraper-url, and the resulting object carries exactly those keys.
That composition model is the strongest part of the design. You are not configuring an extractor through options; you are choosing which extractors exist. Unknown fields simply do not appear in the output, so the shape of the result is a direct function of the array you passed in. The cost is that adding a field later means adding a dependency, and the README points to rules bundles for custom detection when the official ones do not cover what you need.
Installing metascraper and extracting your first page
The README does not give a bare npm install line, so the packages come from the npm registry under the names shown in its example. Install the core library plus the individual extractors you want, then wire them together.
The example below is the README's own pattern, trimmed to the retrieval and extraction steps. html-get returns the rendered markup, and the metascraper instance is built from the rule bundles. Each require() call corresponds to one output field.
const getHTML = require('html-get')
const browserless = require('browserless')()
const getContent = async url => {
const browserContext = browserless.createContext()
const promise = getHTML(url, { getBrowserless: () => browserContext })
promise.then(() => browserContext).then(browser => browser.destroyContext())
return promise
}With the HTML in hand, construct the extractor. The array order does not determine precedence; each bundle owns its own field.
const metascraper = require('metascraper')([
require('metascraper-author')(),
require('metascraper-date')(),
require('metascraper-description')(),
require('metascraper-image')(),
require('metascraper-logo')(),
require('metascraper-publisher')(),
require('metascraper-title')(),
require('metascraper-url')()
])
getContent('https://microlink.io')
.then(metascraper)
.then(metadata => console.log(metadata))
.then(browserless.close)
.then(process.exit)According to the README, the output looks like this: an object with author, date, description, image, logo, publisher, title and url keys, where date is an ISO 8601 string and url is the canonical article URL. If a field is missing from the page, expect it to be absent or empty rather than present with a placeholder.
The README also documents an API surface beyond the simplest call: metascraper(options) accepts html, htmlDom, omitPropNames, pickPropNames, rules, url and validateUrl. pickPropNames is the one worth knowing about early, since it lets you limit the returned object without changing which bundles you installed. There is also an environment variable, METASCRAPER_RE2, documented in the README.
Where metascraper stops being the right tool
The clearest limitation is stated by the project itself. The README says that running this at scale means operating headless browsers, proxies and antibot workarounds, and then offers the managed Microlink API as the alternative for teams that do not want to manage that infrastructure. That is an unusually candid admission, and it should be read as a scope boundary rather than marketing.
So the failure mode is not usually in the parsing. It is upstream. If the HTML you hand metascraper is a bot-check page, a consent wall, a login redirect or an empty JavaScript shell, the extractor will faithfully return metadata for that page instead of the article you wanted. metascraper cannot tell the difference, because it never sees the request. A page that renders its Open Graph tags client-side will look empty to a plain HTTP client and populated to a headless browser.
Second, accuracy is bounded by what publishers emit. The README frames the default behaviour as high accuracy for online articles, which is a narrower claim than high accuracy for arbitrary web pages. Product listings, forums, documentation sites and single-page applications do not follow article conventions, and the fallback chain has less to work with. If your corpus is not articles, budget time for custom rules.
Third, the dependency model cuts both ways. Because metascraper is a collection of tiny packages, a project that needs ten fields carries ten dependencies with their own release cadence. The repository is a monorepo using lerna and pnpm-workspace.yaml, and the packages under packages/ version independently, so upgrades are not always a single bump.
metascraper versus hand-rolled Open Graph parsing
The obvious alternative is to skip the library and read og: tags yourself with a parser such as cheerio. For a narrow case, that is defensible. If you only ever need og:title, og:image and og:description, a few selector lookups will do it, and you avoid the dependency tree entirely.
The difference in approach is what happens when those tags are missing. A hand-rolled reader returns undefined and you write the fallback logic yourself: check twitter:title, then JSON-LD, then the <title> element, then strip the site name suffix. metascraper's rule bundles already encode that chain, and the README's principle of not restricting rules to CSS selectors is precisely what allows JSON-LD and RDFa to participate in the same chain. Reproducing that across a dozen fields is the work you are buying.
The trade-off is transparency. With your own selectors, every extraction decision is visible in your code. With metascraper, the precedence lives inside each bundle, and when a field comes back wrong you debug inside a dependency. The README points to a rules bundles section listing official and community bundles, which is where you would look to understand or replace a given field's logic.
There is also a hosted option on the other side of the spectrum. The README describes the Microlink API as handling proxy rotation, paywalls, bot detection and restricted platforms such as major social networks, with pay-as-you-go pricing that starts for free. That is a different product with a different cost structure: you trade infrastructure work for a per-request dependency on a third party. metascraper sits between the two, and choosing it means accepting that you own the fetching layer.
Maintenance, licensing and what an upgrade actually costs
The repository is not archived, and its last push was on 2026-09-18. The release history shows v5.58.3 on 2026-09-18, v5.58.2 on 2026-09-17 and v5.58.1 on 2026-09-17, which indicates a rapid patch cadence rather than a slow, batched one. For a library whose accuracy depends on the metadata publishers emit, frequent small releases are the expected shape.
That cadence has a practical consequence for upgrades. Because the project is a monorepo of independently versioned packages, the version numbers you pin are per-extractor, and a patch to metascraper-image does not imply a patch to metascraper-title. If you pin exact versions, expect to bump several entries rather than one. The README's benchmark section is the place to look if you want to understand performance characteristics before upgrading, though it is a section heading in the README rather than something this article can report numbers for.
The licence is MIT, which is permissive and imposes no copyleft obligation on your application. That is a factual statement about the licence identifier, not legal advice; if your organisation has specific compliance requirements around dependency licences, run them through your own process. Note that the README promotes a commercial hosted service from the same maintainers, so the library and the API are separate offerings with separate terms.
Editorial conclusion
Adopt metascraper if you already control HTML retrieval and want deterministic, rule-based field extraction without a hosted dependency. Skip it if you need proxy rotation, paywall handling or bot-detection evasion, since the README points those users at the managed Microlink API instead. Before committing, verify two things yourself: how your chosen fetcher renders JavaScript-heavy pages, and which of the packages/ rule bundles match the fields you actually need.
Frequently asked questions
What is metascraper used for?
It extracts unified metadata from a web page, resolving Open Graph, Microdata, RDFa, Twitter Cards, JSON-LD and regular HTML into a single object. The README describes it as a library for scraping metadata from an article, with a default emphasis on accuracy for online articles.
Does metascraper fetch the URL for me?
No. It requires two inputs, the target URL and the HTML markup behind that URL, and the README states the markup needs to be as accurate as possible. The maintainers built html-get, which uses a headless browser, for that retrieval step.
What does the metascraper output object look like?
The README's example returns keys such as author, date, description, image, logo, publisher, title and url. The date is an ISO 8601 representation and the url is the article URL.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/microlinkhq-metascraper)