pdfminer.six: text extraction from PDF source code, and what it refuses to do
Community maintained fork of pdfminer - we fathom PDF
At a glance
- What is it?
- pdfminer.six is a community fork of PDFMiner that reads text, positions, fonts and colors straight from the PDF's own content streams. It is a parser first and a convenience library second, and that shapes both what it does well and where it breaks.
- Who is it for?
- Adopt pdfminer.six when you need text, coordinates, fonts or colors pulled from the PDF's own content streams, and when a pure-Python dependency set matters. Do not adopt it if you need a rendered page image, a table object, or a maintained response time on bug reports; the README itself says the best way to get an issue resolved is to submit a pull request.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Activity is slowing. The repository last received commits 6 months ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem pdfminer.six solves, and the developer it is aimed at
Most tools that claim to read PDFs give you a string and hide everything else. pdfminer.six does the opposite. It parses the document and exposes the intermediate structure: the interpreter walks the content stream, and the rendering devices decide what to do with each operator. The README states the project "focuses on getting and analyzing text data" and extracts text "directly from the sourcecode of the PDF." It can also report the exact location, font or color of the text.
That framing tells you who the library is for. It is for a developer who has to answer questions like: which page did this sentence come from, what font size was it set in, is this column left of that one, and does the color change mid-paragraph. It is also for someone who wants to replace one stage of the pipeline. The README says the project is "built in a modular way such that each component of pdfminer.six can be replaced easily," and that you can implement your own interpreter or rendering device.
The scope is deliberately narrower than "read PDFs." There is no rendering to an image in the feature list. There is no table model. If your task is to turn a scanned invoice into rows in a database, this is not the layer you want to start at.
How the parser works: content streams, devices and layout analysis
The architecture visible in the repository is a classic interpreter plus device split. The `pdfminer/` package holds the parser, and the command-line tools live in `tools/` as `pdf2txt.py` and `dumppdf.py`; pyproject.toml registers both as script files. The `cmaprsrc/` directory holds the CID-to-code source tables, and the Makefile converts them into `pdfminer/cmap/to-unicode-Adobe-CNS1.json.gz` and the equivalent files for Adobe-GB1, Adobe-Japan1 and Adobe-Korea1 using `tools/conv_cmap.py`. Those four files are what make CJK text resolvable to Unicode, and the Makefile shows the encodings each one is built from (cp950, cp936, cp932, euc-kr and others).
Above the parser sit the devices. A layout analysis stage groups characters into lines and blocks, which is why the output has a reading order at all rather than a stream of glyphs in drawing order. The README lists "Automatic layout analysis" as a feature, and the pyproject keywords include "layout analysis" as a distinct concept from "pdf parser." That separation is the useful part: the same parsed document can be rendered as text, HTML, hOCR, or inspected for form fields, table of contents entries and tagged content.
The dependency list is short: `charset-normalizer >= 2.0.0` and `cryptography >= 36.0.0`. Cryptography is there because the README claims support for RC4 and AES encryption. Image extraction is optional and pulls in Pillow.
Installing pdfminer.six and running a first extraction
The README requires Python 3.10 or newer; pyproject.toml sets `requires-python = ">=3.10"` and the classifiers list 3.10 through 3.14. Install the library from PyPI:
pip install pdfminer.sixIf you plan to extract embedded images, the README documents an extra that installs Pillow:
pip install 'pdfminer.six[image]'The package ships a command-line entry point. Point it at a PDF and it writes the extracted text to standard output:
pdf2txt.py example.pdfIn Python, the high-level API is a single function. The README gives exactly this example:
from pdfminer.high_level import extract_text
text = extract_text("example.pdf")
print(text)What you should see is the text of the document in reading order, not a rendered page. If you need the coordinates, fonts or colors that the README mentions, `extract_text` is the wrong entry point: you have to drop to the layout and interpreter classes that `extract_text` wraps. The repository includes `samples/simple1.pdf` through `samples/simple5.pdf` and `samples/jo.pdf` if you want a file to try the command on before you point it at production documents.
Where pdfminer.six fails: reading order, tables and the maintainer bottleneck
The first limitation is that layout analysis is a heuristic, and the README does not promise otherwise. Multi-column pages, footnotes, sidebars and text that is positioned rather than flowed will be grouped by proximity and baseline, and the result can interleave columns or drop a caption into the middle of a paragraph. Nothing in the README describes a configuration that guarantees a correct reading order for a given document class. You have to test it against your own files.
The second is that tables are not modelled. A table in a PDF is a set of positioned glyphs plus ruling lines; nothing in the feature list turns that into rows and columns. If your pipeline needs structured cells, you are building that layer yourself on top of the character positions pdfminer.six exposes.
The third is maintenance. The repository is not archived, and the last push was on 2026-03-13. The README is unusually direct about capacity: it describes the project as "a community-maintained project with limited maintainer availability" and says "the best way to get an issue resolved is to submit a pull request yourself." That is a statement about response time, not about code quality, but it should shape how you plan around a bug in a document you must parse. The project also keeps a `fuzzing/` directory and an `atheris` dependency group, which tells you malformed or hostile PDFs are treated as an input class worth testing against. Fuzzing finds crashes; it does not make a parser safe against every crafted file.
pdfminer.six compared with pdfplumber and PyMuPDF
pdfplumber is the closest comparison because the relationship is structural rather than competitive. pdfplumber is built on top of pdfminer.six and adds a table-extraction layer and a higher-level object model for words, lines and rectangles. The difference in approach is where the abstraction sits: pdfminer.six gives you the parsed document and the devices, while pdfplumber gives you convenience objects and table finding on top of the same parsing core. If you need tables, starting at pdfminer.six means writing the layer that pdfplumber already ships.
PyMuPDF takes a different route entirely. It binds to a C rendering engine rather than implementing the parser in Python, so it can rasterize pages and tends to be the choice when rendering or speed of page handling matters. The trade-off is the dependency: pdfminer.six installs as pure Python with two dependencies, which matters in environments where a compiled extension is awkward to build or audit. The README's feature list is about parsing and extraction, not rendering, so if your problem is "show me the page as an image," pdfminer.six is the wrong tool and PyMuPDF is the right shape of tool.
Against pypdf, the split is similar: pypdf targets document manipulation (merging, splitting, page-level operations) alongside text extraction, while pdfminer.six stays on analysis. If you need to rewrite a PDF rather than read it, look elsewhere.
Licence, upgrade cost and what the release cadence implies
The licence is MIT, declared both in the README badge area and as `license = "MIT"` in pyproject.toml. That is permissive and imposes no copyleft obligation on your own code, but the repository also states that it "includes code from pyHanko" and that the original license for that code is kept at `docs/licenses/LICENSE.pyHanko`. If you vendor the package or redistribute it, read that file rather than assuming a single uniform licence across every file. This is a description of what the repository says, not legal advice.
The recent releases are dated 2025-12-29, 2025-12-30 and 2026-01-07, with the last push on 2026-03-13. Releases are versioned by date rather than semver, which means an upgrade decision cannot be made from the version number alone; you have to read CHANGELOG.md to see whether a parsing behaviour changed. For a library whose output is text, a change in layout analysis is a change in your data, so pinning the version and diffing output across an upgrade on a fixed sample set is the practical approach. The `samples/` directory gives you files to diff against, though several of them are deliberately small and simple.
The dependency floor is low enough that upgrades are unlikely to be blocked by transitive churn: charset-normalizer and cryptography, plus Pillow only if you opt into the image extra. The Python floor moves with the project, currently 3.10, so an environment pinned to 3.9 cannot use current releases.
Editorial conclusion
Adopt pdfminer.six when you need text, coordinates, fonts or colors pulled from the PDF's own content streams, and when a pure-Python dependency set matters. Do not adopt it if you need a rendered page image, a table object, or a maintained response time on bug reports; the README itself says the best way to get an issue resolved is to submit a pull request. Before committing, verify on your own corpus that the layout analysis produces the reading order you expect, and check whether the CJK cmap data shipped in pdfminer/cmap covers the documents you will parse.
Frequently asked questions
How do I install pdfminer.six?
Install Python 3.10 or newer, then run pip install pdfminer.six. If you also need embedded image extraction, the README documents pip install 'pdfminer.six[image]', which adds Pillow.
How do I use pdfminer.six to extract text from a PDF?
The README gives two routes: the command line, pdf2txt.py example.pdf, or the Python high-level function extract_text from pdfminer.high_level. Both return the document text in reading order rather than a rendered page.
What is pdfminer.six?
It is a community maintained fork of the original PDFMiner, written entirely in Python, that parses and analyzes PDF documents and extracts text directly from the PDF source code. The README also lists support for extracting images, HTML and hOCR.
Is pdfminer.six safe?
The repository keeps a fuzzing/ directory and an atheris dependency group, which indicates malformed PDFs are treated as a test input class. The README does not make a security guarantee, so a fuzzing setup should be read as evidence of testing, not as a safety claim.
How does pdfminer.six differ from pdfplumber?
pdfplumber is built on top of pdfminer.six and adds a higher-level object model and table extraction. Choosing pdfminer.six means working with the parsed document and devices directly, and writing any table layer yourself.
How does pdfminer.six differ from PyMuPDF?
pdfminer.six implements the parser in pure Python and focuses on text analysis, while PyMuPDF binds to a C rendering engine and can rasterize pages. If your task needs a page image, pdfminer.six is the wrong tool.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/pdfminer-pdfminer-six)