CLI tool
py-pdf/pypdf avatar
py-pdf/pypdf

pypdf: a pure-Python PDF library for splitting, merging and extracting text

A pure-python PDF library capable of splitting, merging, cropping, and transforming the pages of PDF files

10,225 stars1,635 forksPythonNOASSERTION

At a glance

What is it?
pypdf is a pure-Python library for reading, splitting, merging and transforming PDF files. It is the maintained successor to PyPDF2, and its limits matter as much as its features.
Who is it for?
Adopt pypdf when you need scriptable page-level surgery on PDFs inside a Python process and cannot ship a native binary. Do not adopt it as an OCR engine, and do not expect text extraction to reconstruct complex layouts.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 4 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 26, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What pypdf does that a PDF viewer cannot

A desktop PDF viewer is built for one human reading one document. pypdf is built for a program handling many documents without a human. The README describes it as a pure-python PDF library capable of splitting, merging, cropping, and transforming the pages of PDF files, and it also adds custom data, viewing options and passwords, plus retrieval of text and metadata.

The audience is narrow but real: backend engineers who need to slice a 400-page scan into per-chapter files, merge invoice PDFs into one archive, stamp a password on an export, or pull plain text out of a report so it can be indexed. Because the library is pure Python, it installs from PyPI with no system package, no compiled extension and no external binary. That single property decides most adoption questions: it works in a Lambda zip, a slim Docker image, or a locked-down CI runner where you cannot apt-get anything.

What it is not is equally clear from the README. pypdf reads the text layer that already exists in a PDF. It is not an OCR engine, so a scanned page with no embedded text yields nothing useful.

The reader, the writer and the page objects

The mental model is small. A PdfReader wraps an existing file and exposes a pages sequence, indexed from zero. Each entry is a page object that carries its own content stream, and the writer side holds a PdfWriter that you append pages to and then write out to a new file. Nearly every pypdf task is some combination of pulling page objects out of a reader and pushing them into a writer.

The README gives the canonical entry point: construct a PdfReader over a path, take len(reader.pages) for the page count, index into reader.pages, and call extract_text() on the page. That last call is where the abstraction leaks. A PDF page is a sequence of drawing operators, not a text grid, so extract_text() reconstructs reading order from positioning hints. Columns, tables, footnotes and rotated text can come back interleaved or in the wrong order. Treat the output as a good approximation for prose and a poor one for tabular data.

The package metadata shows the dependency footprint is deliberately thin: the only runtime dependency is typing_extensions, and only for Python below 3.11. Everything heavier is optional. AES encryption and decryption need the crypto extra, which pulls in cryptography. Font work needs fonttools, image work needs Pillow, and right-to-left text needs arabic-reshaper and python-bidi. You pay for those only if you use them.

Installing pypdf and extracting text from a real file

Installation is one pip command. The README gives it directly, and there is no build step or system dependency to satisfy first.

bash
pip install pypdf

If you plan to open files protected with AES encryption, install the crypto extra instead. Without it, encrypted files raise an error rather than silently returning garbage, which is the safer failure mode.

bash
pip install pypdf[crypto]

A first real use is text extraction. The README shows the shape of it, and the only change worth making immediately is to print the version, because the project asks for exactly that in bug reports.

python
import pypdf
from pypdf import PdfReader

print(pypdf.__version__)
reader = PdfReader("example.pdf")
print(len(reader.pages))
page = reader.pages[0]
print(page.extract_text())

You should see the installed version, then the page count, then the text of page one on stdout. If the page count is right and the text is empty, the file most likely has no text layer and needs OCR before pypdf can help. Merging is the other common first task: create a PdfWriter, call append on it once per input path, then write the result to a new file. The documentation covers merging, cropping and transforming under separate user-guide pages, and there is a separate CLI project, pdfly, for people who want these operations from a shell rather than from Python.

Where pypdf quietly fails

Text extraction quality is the first honest limitation. Because the library infers order from glyph positions, multi-column layouts, tables and documents with heavy floating elements will produce text that reads correctly in fragments and incorrectly as a whole. If your downstream consumer is a search index that only needs keywords, that is fine. If it is a parser expecting rows and columns, you will spend more time repairing output than the extraction saved.

Scanned documents are the second. pypdf has no OCR path in the README or the package extras. A page that is a single embedded image returns empty or near-empty text, and no amount of configuration changes that.

Encryption is the third, and it is a packaging trap rather than a bug. AES support lives behind the crypto extra, so a library that works on your laptop can fail in a container built from the same requirements file if the extra was dropped. The failure appears at runtime, on the first encrypted file, which is usually in production.

Finally, pypdf is the wrong tool for rendering. It transforms pages and writes PDFs; it does not rasterize them for display. If your goal is a thumbnail or a pixel-perfect preview, you need a renderer, and pypdf will not get you there.

pypdf versus PyMuPDF, pdfplumber and PyPDF2

The comparison that matters most is with PyMuPDF. PyMuPDF wraps a compiled rendering engine, so it can rasterize pages and its text extraction is generally stronger on difficult layouts. The cost is a native dependency that must be built or installed for each platform. pypdf trades that extraction quality for portability: one pip install, no binary, works anywhere CPython runs. If you need rendering or you are fighting messy layouts, PyMuPDF is the better fit and you should accept the deployment weight.

pdfplumber is the other targeted alternative. It is built around extracting tables and words with coordinates, which is exactly the case where pypdf's plain extract_text() is weakest. If your task is pulling structured tables out of reports, pdfplumber's model fits the problem better than pypdf's page-level text dump.

The PyPDF2 comparison is a naming question more than a technical one. pypdf is the continuation of that codebase under a new name, and the README points to a migration guide for the 1-to-2 transition, noting that pypdf 3.1.0 and above include significant improvements over previous versions. New code should import from pypdf. Old code importing PyPDF2 is running an abandoned name, and the search traffic still going to PyPDF2 is a good sign that a lot of that code has not been touched.

One more distinction worth stating plainly: pypdf and PyMuPDF are different projects, not two spellings of the same thing, and questions that treat them as interchangeable tend to come from people who have not yet hit the rendering wall.

Release cadence, licence and the cost of staying current

The repository is not archived, and the most recent push was on 2026-09-21. Releases are frequent: 6.19.0 on 2026-09-16, 6.18.1 on 2026-09-11 and 6.18.0 on 2026-09-07. That cadence is a real maintenance cost on the consumer side. A library that ships minor versions every few days will occasionally change behaviour in ways that show up as different extracted text or different writer output, so pinning a version and reading the changelog before bumping is cheaper than discovering the change in production.

The upside of the cadence is that PDF edge cases get fixed rather than accumulating. The project also states its contribution expectations clearly: bug reports need a minimal complete verifiable example with the PDF attached and the output of print(pypdf.__version__). That is a high bar, and it is the reason issues get resolved instead of stalling on unverifiable reports.

The licence situation needs care. The package metadata declares BSD-3-Clause and ships a LICENSE file, but the repository's licence field as reported by the hosting platform is NOASSERTION, meaning the platform could not classify it automatically. For most users the pyproject.toml declaration is the operative one, but if your organisation runs automated licence scanning, expect that field to raise a flag that a human has to clear. This is a description of what the metadata says, not legal advice; check the LICENSE file yourself before shipping.

Installing is free and there is no paid tier, no telemetry and no account. The upgrade cost is entirely in testing your own corpus after each bump.

Editorial conclusion

Adopt pypdf when you need scriptable page-level surgery on PDFs inside a Python process and cannot ship a native binary. Do not adopt it as an OCR engine, and do not expect text extraction to reconstruct complex layouts. Before committing, run it against your own worst-case file and print pypdf.__version__ so any bug report carries the version you actually tested.

Frequently asked questions

What is pypdf in Python?

It is a pure-Python library for reading and writing PDF files. The README lists splitting, merging, cropping, transforming pages, adding custom data, viewing options and passwords, and retrieving text and metadata.

Which is better, pypdf or PyPDF2?

pypdf is the continuation of that codebase under a new name, and the README points to a migration guide for the 1-to-2 transition. It also notes that pypdf 3.1.0 and above include significant improvements compared to previous versions, so new code should import from pypdf.

What is the difference between pypdf and PyMuPDF?

pypdf is pure Python and installs with a single pip command and no external binary. PyMuPDF wraps a compiled rendering engine, which gives it rendering and generally stronger extraction on difficult layouts at the cost of a native dependency. If you need to rasterize pages, pypdf does not do that.

Is pypdf an OCR tool?

No. pypdf retrieves the text layer already present in a PDF, and neither the README nor the package extras list an OCR path. A scanned page with no embedded text returns empty or near-empty output.

How do I install pypdf?

Install it with pip install pypdf. If you need AES encryption or decryption, install pip install pypdf[crypto] instead, which pulls in the cryptography package.

How do I use pypdf to extract text from a PDF?

Construct a PdfReader over the file path, index into reader.pages, and call extract_text() on the page object. The README shows exactly this sequence, and the result is the text layer reconstructed from the page content, so complex layouts may come back in the wrong order.

Official sources

  1. Issues
  2. Project website
  3. py-pdf/pypdf on GitHub
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/py-pdf-pypdf.svg)](https://hysenlabs.com/projects/py-pdf-pypdf)