Trafilatura: extracting article text and metadata from raw HTML
Python & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML
At a glance
- What is it?
- Trafilatura is an Apache-2.0 Python package and command-line tool for crawling, downloading and extracting main text and metadata from web pages, with output as TXT, Markdown, CSV, JSON, HTML, XML or XML-TEI. Its own rule-based extractor leads the pipeline, with jusText and readability-lxml as fallbacks.
- Who is it for?
- Adopt Trafilatura if you need article bodies and metadata as structured files without running a database, and if a rule-based extractor with jusText and readability-lxml fallbacks fits your tolerance for imperfect recall. Do not adopt it as a general HTML parser: for scraping prices, tables or form-driven flows, BeautifulSoup or a browser automation stack is the right layer.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem Trafilatura solves: HTML noise versus article text
A news page arrives as navigation bars, cookie notices, related-article rails, comment threads and a footer, wrapped around perhaps 600 words of prose. If you are building a corpus, a retrieval index or a fine-tuning set, that wrapper is the problem. Trafilatura's stated goal is to go from raw HTML to the essential parts, focusing on actual content and avoiding noise caused by recurring elements. The README frames the extractor as striking a balance between precision (limiting noise) and recall (including all valid parts).
The audience follows from that. The project's own history places it at the crossroads of linguistics and NLP, started to create text databases for research at the Berlin-Brandenburg Academy of Sciences. The pyproject classifiers list Scientific/Research, Education and Information Technology as intended audiences, and the topic list includes corpus-builder, text-mining, news-aggregator, rag and llm. In practice this is a tool for people who need clean article bodies at volume, not for people who need a specific DOM node from one page.
How the extraction pipeline is put together
The README describes a full chain rather than a single function: discovery, downloads, scraping, and extraction of main text, metadata and comments. Discovery covers sitemaps in TXT and XML form plus feeds in ATOM, JSON and RSS, with URL filtering and deduplication. Inputs can be live URLs, previously downloaded HTML files, or already-parsed HTML trees, and the README states that online and offline input can be processed in parallel, with what it calls polite processing of download queues.
Extraction is where the design choice shows. Trafilatura ships its own rule-based extractor and uses jusText and readability-lxml as fallbacks when that fails. It can return main text, metadata (title, author, date, site name, categories, tags), and structural elements: paragraphs, headings, lists, quotes, code, line breaks and inline formatting. Comments, links, images and tables are optional. Language detection and speed optimizations are listed as optional add-ons. Output formats are TXT, Markdown, CSV, JSON, HTML, XML and XML-TEI.
Two consequences are worth stating plainly. First, because the primary extractor is rule-based rather than a trained model, behaviour is inspectable and deterministic for a given input, but it is also tuned against the kinds of pages the maintainers and contributors have seen; sites with unusual layouts are where fallbacks and misses appear. Second, the README does not document rollback or a per-site override mechanism, so the practical unit of control is the extraction call and its options, not a persistent site profile.
Installing Trafilatura and running a first extraction
The package is on PyPI and requires Python 3.10 or later according to pyproject.toml. A standard install is enough for the library and the CLI:
pip install trafilaturaAfter that, the command-line entry point is available. The README's own Python example uses fetch_url and extract, so the equivalent first check is to fetch one page and print the extracted text:
from trafilatura import fetch_url, extract
downloaded = fetch_url("https://github.blog/2019-03-29-leader-spotlight-erin-spiceland/")
extract(downloaded)The README shows that this returns the article body as a string beginning with the opening sentence of the post. To get metadata as well, the documented call passes output_format and with_metadata:
from trafilatura import fetch_url, extract
downloaded = fetch_url("https://github.blog/2019-03-29-leader-spotlight-erin-spiceland/")
extract(downloaded, output_format="json", with_metadata=True)According to the README, that returns a JSON string carrying title, author and text fields. The documentation site linked from the README covers command-line usage, Python usage, R usage and the core functions; the repository also ships an interactive notebook, docs/Trafilatura_Overview.ipynb. If you prefer not to fetch at extraction time, the README lists previously downloaded HTML files and parsed HTML trees as accepted inputs, which is the route to take when you want reproducible runs over a fixed corpus.
Where Trafilatura is the wrong tool
Trafilatura is built around article-shaped documents. If your target is a product listing, a paginated table, a search result page or anything behind a login or heavy JavaScript, the extraction model does not match the task, and the fallback chain will not rescue you. The README does not claim browser rendering, session handling or form interaction, and none of the listed features cover them.
There is a second, subtler limit. The precision and recall balance is a single global trade-off: the extractor is tuned once, not per site. When a page's main content is interleaved with something the rules read as boilerplate, or when a site's markup is unusual enough that the rule-based pass fails, you get either noise or a short extraction, and the README does not document a supported way to pin a site to a hand-written rule. For a corpus spanning many domains this is normally acceptable, because errors average out. For a pipeline that must be correct on one high-value site, it is a real risk, and the honest answer is to measure that site specifically before you commit.
Metadata deserves the same scepticism. Title, author, date, site name, categories and tags are listed as extractable, but they come from what the page declares. Pages with missing or malformed metadata will yield empty fields, and a JSON record with a null author is a data problem you inherit, not one Trafilatura introduces.
Trafilatura compared with BeautifulSoup and readability-lxml
The most common comparison is with BeautifulSoup, and the difference is one of layer. BeautifulSoup is an HTML parsing library: you locate elements with selectors and decide what text matters. Trafilatura decides for you, using a rule-based extractor with jusText and readability-lxml as fallbacks. If you already know the structure of your pages and want exact control, BeautifulSoup is the more direct instrument, and Trafilatura's heuristics are an extra layer you would be fighting.
Against readability-lxml the relationship is closer, because readability-lxml is inside Trafilatura as a fallback rather than beside it. The README positions Trafilatura's own extractor as the primary path, with readability-lxml used when that does not succeed. So the practical difference is not which algorithm is smarter in the abstract; it is that Trafilatura adds discovery (sitemaps, feeds), download queue management, metadata extraction and seven output formats around the extraction step, while readability-lxml gives you the extraction step. If you only need one page's body in Python, readability-lxml is the smaller dependency. If you need a corpus, the surrounding machinery is the reason to pick Trafilatura.
On accuracy the README cites external evaluations rather than internal numbers: ScrapingHub's article extraction benchmark, Lejeune and Barbaresi (2020), and Bevendorff et al. (2023), where the README states Trafilatura is the best single tool by ROUGE-LSum Mean F1 page scores. Those are the project's cited results, not something to take on faith for your own domains; the repository contains an evaluation readme under tests/ if you want to reproduce the setup.
Licence, maintenance and what an upgrade costs
Trafilatura is distributed under Apache-2.0. The README adds an important boundary: versions prior to v1.8.0 were under GPLv3+. If you are auditing a dependency tree that pinned an old release, that distinction matters, and it is the kind of thing to check before assuming the whole history is permissively licensed. Nothing here is legal advice; read the LICENSE file in the repository for the operative text.
The release cadence visible in the repository is uneven. v2.0.0 landed on 2024-12-03, v2.1.0 on 2026-06-07 and v2.2.0 on 2026-07-31, with the last push to the default branch on 2026-08-28. The gap between the 2.0 and 2.1 releases is the practical warning: this is a project where a major version can sit unchanged for well over a year, then move twice in two months. Budget for reading HISTORY.md before an upgrade rather than assuming a patch-level bump is inert.
The README states that the project's future depends on community support and points to GitHub Sponsors and ko-fi. That is a maintenance-cost signal as much as a funding one: for a tool you build a corpus pipeline on, the version you pin and the extraction behaviour you validated are things you own, because the project's continuity is explicitly tied to contributions.
Editorial conclusion
Adopt Trafilatura if you need article bodies and metadata as structured files without running a database, and if a rule-based extractor with jusText and readability-lxml fallbacks fits your tolerance for imperfect recall. Do not adopt it as a general HTML parser: for scraping prices, tables or form-driven flows, BeautifulSoup or a browser automation stack is the right layer. Before committing, verify three things on your own corpus: how your target sites fare under the default extractor and the --format options you intend to ship, whether the metadata fields you need (author, date, site name) actually come back populated for those sites, and whether the requires-python >=3.10 floor and the Apache-2.0 licence terms match your deployment.
Frequently asked questions
What are the advantages of Trafilatura compared to BeautifulSoup?
BeautifulSoup is an HTML parsing library where you choose the elements; Trafilatura decides what the main content is using its own rule-based extractor, with jusText and readability-lxml as fallbacks. Trafilatura also bundles discovery through sitemaps and feeds, download queue management, metadata extraction and output as TXT, Markdown, CSV, JSON, HTML, XML or XML-TEI.
How to install trafilatura?
Install it from PyPI with pip install trafilatura. The package requires Python 3.10 or later according to pyproject.toml, and the README links to an installation page in the documentation for further detail.
What is the Trafilatura library?
It is a Python package and command-line tool that gathers text and metadata from the web, covering crawling, downloading, scraping and extraction. The README describes it as modular, requiring no database, with output convertible to commonly used formats.
How to use trafilatura?
The README's example imports fetch_url and extract, downloads a page, and calls extract on the result; passing output_format="json" with with_metadata=True returns a JSON string with title, author and text. The documentation links from the README cover command-line, Python and R usage.
Is trafilatura open source?
Yes. It is distributed under the Apache 2.0 license, and the README notes that versions prior to v1.8.0 were under GPLv3+.
Is trafilatura free?
The package is distributed under the Apache 2.0 license, so there is no licence fee to use it. The README does ask users who depend on it to consider sponsoring the project on GitHub or ko-fi, and states that its future depends on community support.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/adbar-trafilatura)