Library / SDK
WikiExtractor/wikiextractor avatar
WikiExtractor/wikiextractor

WikiExtractor: turning a Wikipedia XML dump into plain text

A tool for extracting plain text from Wikipedia dumps. Warning**: problems have been reported on Windows due to poor support for StringIO in the Python implementation on Windows.

4,006 stars1,000 forksPythonAGPL-3.0

At a glance

What is it?
WikiExtractor is a Python script that parses a Wikipedia database backup dump and writes cleaned plain text documents to disk. It needs no third-party libraries, but template expansion and Windows support are the two places where expectations tend to break.
Who is it for?
Adopt WikiExtractor if you need plain text corpora from a Wikipedia dump and are comfortable running a Python 3 script over a multi-gigabyte bz2 file on Linux or macOS. Do not adopt it if you need structured wikitext, infobox fields or template parameters preserved, or if your pipeline runs on Windows, where the README reports StringIO problems.
Can I use it commercially?
Yes, with strict conditions. AGPL-3.0 is a network copyleft licence: if people use a modified version over a network, for example as a hosted service, you must offer them its source code under the same licence.
Is it still maintained?
Yes. The repository last received commits 4 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 26, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What WikiExtractor extracts, and what it refuses to give you

A Wikipedia database backup dump is an XML file containing every page, its wikitext source, and the metadata around it. That source is not readable prose. It is markup: links, tables, refs, and template invocations such as {{Infobox person}} whose visible text only exists after the template is expanded. WikiExtractor's job is to walk that dump and emit cleaned plain text, one document per article, in a directory of files of similar size.

The intended audience is narrow and practical: people building text corpora for search, language modelling, or linguistic analysis who want sentences rather than wiki syntax. The README describes the tool as a script that "extracts and cleans text from a Wikipedia database backup dump", and the output format is a doc element carrying id, url and title attributes. If you need the wikitext itself, or a parsed abstract with typed fields, this is the wrong layer of the stack.

Template expansion, multiprocessing and the template cache

The mechanism the README describes is a two-pass approach. WikiExtractor preprocesses the whole dump to extract template definitions, then uses those definitions to expand templates in the articles. That is why the first run over a full English dump is expensive: the tool is effectively reading the dump twice, and the template definitions are the reason.

Two optimisations are documented. Articles are handled in parallel through multiprocessing, and a cache of parsed templates is kept. The cache is only useful for repeated extractions, which tells you the intended workflow is iterative: extract once, then re-extract with different output settings without paying the full template cost again. The --templates option writes those definitions to a local file so a later run can reload them. The README is explicit that reloading is only safe if template definitions have not changed, which in practice means the cache belongs to one dump revision, not to the tool.

There is a documented escape hatch. --no-templates significantly speeds up the extractor by avoiding template expansion altogether. That is a real trade-off, not a free win: skipping expansion means template-generated text never appears in the output. The README does not spell out what is lost per article, so the only reliable way to judge is to extract a small slice both ways and diff the results.

Installing WikiExtractor and running a first extraction

The README gives three installation routes. You can invoke the module directly from a checkout with no installation at all, install from PyPI, or install locally with setup.py. The package declares no install_requires, so nothing else is pulled in, and python_requires is >=3.6.

The PyPI route is the shortest:

bash
pip install wikiextractor

The installer also places two console scripts on your path, wikiextractor and extractPage. The first is equivalent to running the module; the second pulls a single page out of a dump, which is useful for inspecting one article without processing the whole file.

With the package installed, point the tool at a dump file. The README uses the English pages-articles dump as its example:

bash
python -m wikiextractor.WikiExtractor <Wikipedia dump file>

Output lands in the directory given to -o, split into files of similar size (the -b option controls bytes per file and defaults to 1M). Without --json each file holds doc elements with id, url and title; with the flag, each file is JSON instead. Expect a long-running job on a full dump, and expect the directory to fill with many small files rather than one large one.

If you plan to re-extract, save the templates on the first pass and reload them afterwards:

bash
python -m wikiextractor.WikiExtractor <Wikipedia dump file> [--templates <extracted template file>]

There is a separate entry point, cirrus-extractor.py, for Cirrus dumps, which the README says already contain expanded templates. If you have a Cirrus dump, use that script: it accepts a Cirrus JSON dump as input and writes doc elements that also carry language and revision attributes.

The Windows warning and other places it breaks

The README carries a blunt warning: problems have been reported on Windows due to poor support for StringIO in the Python implementation on Windows. That is not a footnote about an edge case. It is a statement that a supported platform is unreliable, and the release history backs the point: v3.1.0 is labelled "Windows compatibility", and the previous release in the series, v3.0.7, dates from 2023-01-24. The gap between v3.0.7 and the 2026 releases suggests the project went quiet for a stretch and then returned with a Windows-focused fix. The last push to the repository was on 2026-08-10.

The second failure mode is silent rather than loud. --no-templates does not warn you that template-derived text is missing; it just runs faster. A corpus built with that flag will differ from one built without it, and the difference is concentrated exactly where infoboxes, taxoboxes and citation templates would have contributed text. If your downstream task depends on that content, the speed gain is a trap.

The third constraint is the cache. The README says saving templates speeds up the next extraction assuming template definitions have not changed. Nothing in the documentation describes how the tool detects a stale cache. Treat the template file as bound to one dump revision and regenerate it when you change dumps.

WikiExtractor compared with using the MediaWiki API or a dump parser

The obvious alternative for many people is not another extractor but the MediaWiki API: fetch rendered page text over HTTP for the articles you care about. The difference in approach is fundamental. The API gives you current, rendered HTML for a bounded set of pages, with rate limits and network dependence. WikiExtractor gives you an offline batch conversion of an entire dump, with no network access at extraction time and no per-page request budget. If you need a few thousand articles, the API is simpler. If you need millions, or you need a reproducible corpus pinned to a dump revision, the API is the wrong shape.

A second alternative is to parse the dump yourself with a general XML or wikitext library. That gives you control over which templates you expand and which structures you keep, at the cost of writing and maintaining the expansion logic that WikiExtractor already ships. The trade-off is honest: WikiExtractor decides for you what cleaned text means, and its decisions are not configurable beyond the flags listed in the usage output.

Licence, maintenance and what an upgrade costs

The code is released under the GNU Affero General Public License v3.0, and setup.py classifies it as AGPLv3 or later. The AGPL's network clause is the part that matters for anyone embedding this in a service. If you modify WikiExtractor and let users interact with it over a network, the licence's terms reach further than the GPL's would. Whether that affects your deployment is a question for your own counsel, not for this article. Note also that the output you generate from a Wikipedia dump carries its own licensing obligations from Wikipedia, which are separate from the tool's licence.

Upgrade cost is low in one sense and non-zero in another. There are no runtime dependencies to reconcile, so installing a new version cannot break your environment through a transitive package. But the output format and the template handling are the contract, and a change there invalidates any corpus you built with an earlier version. Because the template cache is tied to a dump revision and the README does not describe cache invalidation, a version bump is a reasonable moment to regenerate templates and re-extract rather than reuse an old cache.

Editorial conclusion

Adopt WikiExtractor if you need plain text corpora from a Wikipedia dump and are comfortable running a Python 3 script over a multi-gigabyte bz2 file on Linux or macOS. Do not adopt it if you need structured wikitext, infobox fields or template parameters preserved, or if your pipeline runs on Windows, where the README reports StringIO problems. Before committing, verify that the version you install matches the dump you downloaded, and check whether --no-templates is acceptable for your corpus, because that flag changes what the output contains.

Frequently asked questions

How do I use WikiExtractor?

Install it with pip install wikiextractor, then pass a Wikipedia dump file to the wikiextractor command with an output directory. The README's example input is the English pages-articles dump, and the tool writes cleaned text into files of similar size in that directory.

Can I legally download Wikipedia?

The README points to the official Wikimedia dump server for backup dumps, and WikiExtractor is built to process those files. The tool's own licence is AGPL-3.0, which is separate from the licensing of the Wikipedia content you extract.

How do I download a file from Wikipedia?

The README links to the Wikimedia dumps site and gives the English pages-articles dump as its example input, for instance enwiki-latest-pages-articles.xml.bz2. WikiExtractor does not download anything itself; you fetch the dump and pass the local file to the script.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
  4. Release notes
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/wikiextractor-wikiextractor.svg)](https://hysenlabs.com/projects/wikiextractor-wikiextractor)