CLI tool
jsvine/pdfplumber avatar
jsvine/pdfplumber

pdfplumber: character-level PDF plumbing in Python

Plumb a PDF for detailed information about each char, rectangle, line, et cetera — and easily extract text and tables.

10,778 stars925 forksPythonMIT

At a glance

What is it?
pdfplumber exposes every character, line and rectangle a machine-generated PDF contains, then builds text and table extraction on top of it. It is a good fit when you need coordinates and layout, and the wrong tool when your PDFs are scans.
Who is it for?
Adopt pdfplumber when your input is machine-generated PDFs and you need coordinates, not just strings: the .chars, .lines and .rects lists are the reason to pick it over a plain text extractor. Do not adopt it for scanned documents, since the README states it works best on machine-generated rather than scanned PDFs, and there is no OCR step in the dependency list (pdfminer.six, Pillow, pypdfium2).
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 56 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What pdfplumber exposes that a text extractor hides

Most PDF libraries answer one question: what does this page say? pdfplumber answers a different one: what is on this page, where, and at what size? Each page carries lists named .chars, .lines, .rects, .curves and .images, and each entry in those lists is a dictionary describing one object. That is the whole design. Text extraction and table extraction are built on top of those dictionaries rather than being the primitive.

The audience follows from that. If you are parsing invoices where the vendor name sits in a fixed band at the top, or statements where a column boundary is a drawn line, you need the coordinates. If you only want the words, a simpler extractor will be less code. The README is explicit that the library works best on machine-generated rather than scanned PDFs, so anyone whose corpus came off a flatbed scanner is in the wrong place.

How the parsing pipeline is put together

pdfplumber is built on pdfminer.six, and the README names that dependency directly. pdfminer.six does the low-level work of walking the PDF's content streams and producing layout objects; pdfplumber wraps those objects in its own Page and PDF classes and adds geometry helpers, table finding and a visual debugging layer.

The dependency list in requirements.txt pins the pieces: pdfminer.six, Pillow and pypdfium2. Pillow and pypdfium2 are what make the visual debugging path possible, since rendering a page to an image requires an imaging library and a PDF renderer. The Python package declares python_requires=">=3.8" in setup.py, while the README says the project is currently tested on Python 3.10 through 3.14. Those two statements are not the same claim, and the gap is worth noticing: the floor is 3.8, the tested ceiling starts at 3.10.

Layout analysis parameters pass through to pdfminer.six's layout engine via the laparams keyword, either in pdfplumber.open(...) or through the CLI's --laparams flag. That is the main tuning surface when default object grouping does not match the document in front of you.

Installing pdfplumber and reading your first page

The README gives a single install command. It pulls the pinned pdfminer.six, Pillow and pypdfium2 dependencies along with the package itself.

bash
pip install pdfplumber

The Python entry point is pdfplumber.open, which accepts a path, a file object loaded as bytes, or a file-like object loaded as bytes. The README's basic example opens a file, takes the first page and prints the first character dictionary. Running it against a machine-generated PDF should print a dictionary of attributes for one character rather than a string.

python
import pdfplumber

with pdfplumber.open("path/to/file.pdf") as pdf:
    first_page = pdf.pages[0]
    print(first_page.chars[0])

Two keyword arguments are documented for open and are easy to miss. Pass password="..." for a password-protected file, and pass unicode_norm with one of "NFC", "NFD", "NFKC" or "NFKD" to pre-normalize extracted text. Invalid metadata is treated as a warning by default; strict_metadata=True turns that into an exception instead.

There is also a command line entry point, installed as a console script named pdfplumber. The README's example downloads a sample PDF and converts it to CSV, where each row describes a character, line or rectangle.

bash
curl "https://raw.githubusercontent.com/jsvine/pdfplumber/stable/examples/pdfs/background-checks.pdf" > background-checks.pdf
pdfplumber background-checks.pdf > background-checks.csv

The CLI takes --format with csv, json or text, a 1-indexed --pages list such as 1,11-15, a --types filter over char, rect, line, curve, image and annot, plus --laparams and --precision. The text format is not the same shape as the other two: the README states it returns a plain-text representation produced by Page.extract_text(layout=True), so it is a rendering of the page rather than a per-object dump.

Cropping, filtering and the strict bounding box default

Page.crop(bounding_box, relative=False, strict=True) returns a page restricted to a 4-tuple of (x0, top, x1, bottom). Objects that fall partly inside are sliced to fit; objects that fall entirely outside are dropped. within_bbox and outside_bbox are the stricter and looser siblings, keeping only objects entirely inside or entirely outside the box.

The default strict=True is a real design decision. The README says that with strict enabled, the crop's bounding box must fall entirely within the page's bounding box. If your coordinates come from an external tool and drift a few points past the page edge, the call raises rather than quietly returning something slightly wrong. That is the right default for data pipelines, and it is also the first thing to check when a crop that worked on one document fails on the next.

Page.filter(test_function) is the other escape hatch: it returns a page containing only the objects for which your function returns true. Between crop and filter you can build a page view that is exactly the region you care about, then run extraction on that view rather than on the whole page.

Where pdfplumber breaks down

Scanned PDFs are the clearest failure case, and the README states it plainly rather than burying it: the library works best on machine-generated PDFs. There is no OCR engine in the dependency list, so a scanned page yields image objects and little else. If your corpus is scans, you need an OCR step before pdfplumber sees the file, and at that point you are choosing between running OCR yourself and picking a tool that bundles it.

Table extraction is the second soft spot, though for a different reason. The README documents extract_tables but the available documentation does not describe the strategy it uses to decide where a table begins and ends. Tables drawn with ruling lines are the easy case; tables implied only by whitespace alignment are where you should expect to write your own tuning. Budget time for that rather than assuming the default call works on your documents.

Version pinning is a smaller but real constraint. requirements.txt pins pdfminer.six to an exact version, so an upstream pdfminer.six release does not reach you until pdfplumber updates that line. That is deliberate, and it also means a bug fixed in pdfminer.six stays fixed only as fast as pdfplumber moves.

pdfplumber vs PyMuPDF and vs docling

The comparison people search for most is pdfplumber against PyMuPDF. The difference is architectural. PyMuPDF binds to a C library and is a general PDF toolkit: rendering, editing, redaction and extraction all live in one fast package. pdfplumber is pure Python on top of pdfminer.six and is narrower by design. It gives you dictionaries for each character, line and rectangle, and it lets you crop and filter pages before extracting. If you need to modify a PDF or render pages at speed, pdfplumber is the wrong shape. If you need to reason about where a character sits relative to a drawn line, the per-object dictionaries are the point.

docling comes at the problem from the document-conversion side, aiming at structured output from whole documents with models in the loop. pdfplumber does not attempt that. It hands you geometry and leaves interpretation to your code. Choosing between them is choosing between a converter that decides for you and a toolkit that does not.

Maintenance, licence and what an upgrade costs

The repository is not archived, and the last push was on 2026-08-06. Releases in the recent series are v0.11.8 on 2025-11-08, v0.11.9 on 2026-01-05 and v0.11.10 on 2026-06-15, so the cadence over that window is roughly a few months between tagged releases rather than a continuous stream. The project is MIT licensed, which permits commercial and closed-source use; the LICENSE.txt file is at the repository root, and anyone with compliance requirements should read it rather than take a summary as authoritative.

Upgrade cost is dominated by the pinned pdfminer.six version rather than by pdfplumber's own API. The library exposes a small surface (PDF, Page, open, crop, filter, extract_text, extract_tables, plus the CLI), so a minor version bump is unlikely to force rewrites. The Makefile shows the project's own gate before a release: pytest, then black, isort, flake8 and mypy --strict. That strict mypy run is worth knowing about if you plan to contribute, and it also tells you the maintainer treats type correctness as part of the contract.

Editorial conclusion

Adopt pdfplumber when your input is machine-generated PDFs and you need coordinates, not just strings: the .chars, .lines and .rects lists are the reason to pick it over a plain text extractor. Do not adopt it for scanned documents, since the README states it works best on machine-generated rather than scanned PDFs, and there is no OCR step in the dependency list (pdfminer.six, Pillow, pypdfium2). Before committing, open one representative page from your own corpus and print first_page.chars[0] and first_page.extract_tables() to see whether the geometry you get matches what your downstream code expects.

Frequently asked questions

How do I install pdfplumber?

The README gives one command, pip install pdfplumber, which also installs the pinned pdfminer.six, Pillow and pypdfium2 dependencies. Installing the package also puts a pdfplumber console script on your path.

Is pdfplumber free?

Yes. The repository is MIT licensed, which permits commercial and closed-source use. Read LICENSE.txt at the repository root for the actual terms rather than relying on a summary.

What is pdfplumber used for?

It exposes detailed information about each text character, rectangle and line in a PDF, and builds text and table extraction on top of that. It works best on machine-generated rather than scanned PDFs.

Which is better, pdfplumber or PyMuPDF?

They are built differently. PyMuPDF binds to a C library and covers rendering, editing and extraction; pdfplumber is pure Python on top of pdfminer.six and gives you per-object dictionaries plus page cropping and filtering. Pick pdfplumber when you need geometry, and PyMuPDF when you need to modify or render documents.

How do I use pdfplumber in Python?

Call pdfplumber.open on a path or a bytes file object, then read pdf.pages. Each page exposes .chars, .lines, .rects, .curves and .images as lists of dictionaries, and the README's basic example prints first_page.chars[0].

Official sources

  1. Issues
  2. jsvine/pdfplumber on GitHub
  3. License: MIT
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/jsvine-pdfplumber.svg)](https://hysenlabs.com/projects/jsvine-pdfplumber)