pikepdf: PDF repair and object-level editing in Python, backed by qpdf
A Python library for reading and writing PDF, powered by QPDF
At a glance
- What is it?
- pikepdf wraps the qpdf C++ library to give Python code a dictionary-style view of PDF objects, automatic repair on open, and lossless image extraction. It is a tool for operating on existing PDFs, not for generating them from HTML or rendering them to images.
- Who is it for?
- Adopt pikepdf when your job is to repair, merge, rotate, encrypt or inspect PDFs that already exist, and when you want qpdf's structural handling without shelling out to it. Do not adopt it to generate PDFs from HTML or templates (weasyprint or reportlab), to render pages to images (PyMuPDF or pypdfium2), or to pull text and tables out of documents (pdfminer.six or pdfplumber); the README names those as the wrong tools for this one.
- Can I use it commercially?
- Yes, with conditions. MPL-2.0 is a weak copyleft licence: you can use it inside commercial and closed-source software, but if you distribute changes to its own files, you must publish those changes under the same licence.
- Is it still maintained?
- Yes. The repository last received commits 23 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 24, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem pikepdf solves: existing PDFs that are damaged, encrypted or need surgery
Most PDF work in Python splits into two jobs that look similar and are not. One is creating a document from content you control. The other is taking a PDF someone else produced, which may be malformed, encrypted, or structurally strange, and changing it without breaking it. pikepdf is built for the second job. Its README describes it as a library to "Read, write, repair, and transform PDFs in Python", and the repair part is not decoration: opening a file through pikepdf runs it through qpdf, which the README says "automatically repairs structural damage" on open. That means a file with a broken cross-reference table or a damaged object stream may open and save cleanly without you writing recovery code.
The audience is developers building tools and libraries that operate on existing PDFs. The README lists the fit directly: repairing, sanitizing or normalizing malformed PDFs; merging, splitting, rotating, cropping or rearranging pages; editing XMP and DocumentInfo metadata; working with encrypted files; low-level object and stream manipulation; and linearizing for web delivery. It also lists what it is not for, which is unusually honest for a README: generating PDFs from HTML or templates, rendering PDFs to images, and extracting text or tables. Those three exclusions matter more than the feature list, because they are where people reach for the wrong library.
How pikepdf works: a Python view over qpdf's object model
pikepdf is a binding, not an independent PDF implementation. The README states it is "powered by qpdf", a C++ library, and the build configuration confirms the arrangement: pyproject.toml declares nanobind as a build requirement and scikit-build-core as the build backend, with a comment noting that wheel-building CI downloads qpdf's source into a ./qpdf directory and compiles libqpdf before building. So the parsing, repair and writing logic lives in C++, and pikepdf exposes it to Python.
The API shape follows the PDF specification rather than hiding it. Dictionaries, arrays, streams and names map to Python types, and the README's example shows direct attribute access on a page object: page.MediaBox returns the media box array, page.Resources.XObject exposes image and form XObjects, and page.Rotate = 90 sets rotation. Pages behave like a list, so del pdf.pages[-1] removes the last page and merged.pages.extend(src.pages) appends pages from another document. That directness is the point: you are editing the object graph, not calling high-level verbs that may or may not preserve what the original file contained.
Two mechanisms are worth calling out because they are easy to miss. Metadata editing goes through pdf.open_metadata(), a context manager that reads and writes XMP with automatic synchronization to DocumentInfo, so the two do not drift apart. And pikepdf exposes qpdf's Job API, which lets you run qpdf's command-line operations programmatically, either as an argument list such as Job(['pikepdf', '--check', 'document.pdf']) or as a JSON job dictionary. That gives you qpdf's full CLI surface from inside a Python process.
Installing pikepdf and a first real edit
The README gives one installation command. Binary wheels are available for Linux, macOS and Windows on x86-64 and ARM64/Apple Silicon, including free-threaded CPython 3.14, and the README states no compiler is required.
pip install pikepdfBuilding from source is a separate path the README points to in its documentation, and it requires compiling libqpdf. If a wheel exists for your platform, you should not need that.
A first use that exercises repair, page deletion and saving: open a PDF, check the page count, drop the last page, write a new file. The README's opening example does exactly this, and the comment notes that repair happens automatically on open.
import pikepdf
with pikepdf.Pdf.open('input.pdf') as pdf:
num_pages = len(pdf.pages)
del pdf.pages[-1]
pdf.save('output.pdf')After this runs, output.pdf exists with one fewer page than input.pdf. If input.pdf had structural damage, the save is where the repaired structure lands.
Merging is the next thing most people need, and it is a loop over sources rather than a special function. Note that the source documents are opened inside the loop and the merged document is created with Pdf.new().
from pikepdf import Pdf
with Pdf.new() as merged:
for filename in ['first.pdf', 'second.pdf', 'third.pdf']:
src = Pdf.open(filename)
merged.pages.extend(src.pages)
merged.save('merged.pdf')Encryption is handled at save time rather than through a separate tool. The README shows saving with an Encryption object that takes user and owner passwords, and shows passing encryption=False to strip encryption when the user password is not set.
import pikepdf
with pikepdf.open('input.pdf') as pdf:
pdf.save('encrypted.pdf', encryption=pikepdf.Encryption(
user='readpassword', owner='adminpassword'
))For a quick structural sanity check without writing Python, the Job API runs qpdf's check operation in-process. The README shows the call as Job(['pikepdf', '--check', 'document.pdf']).run().
Where pikepdf is the wrong tool, and what its repair behaviour costs you
The README is explicit about three exclusions, and they are the most useful part of the document. pikepdf is probably not what you want if you need to generate PDFs from HTML or templates, render PDFs to images, or extract text and tables. For generation it points to weasyprint and reportlab; for rendering, PyMuPDF and pypdfium2; for text and table extraction, pdfminer.six and pdfplumber. Choosing pikepdf for any of those means fighting the library's model, because its model is the PDF object graph, not layout or glyphs.
A second limitation is stated plainly in the README: digital signature-based encryption is not currently supported. If your workflow depends on certificate-based encryption rather than password-based AES-256, AES-128 or RC4, pikepdf will not cover it.
There is also a subtler trade-off in the repair behaviour itself. Automatic repair on open is presented as a feature, and for damaged files it is. But it means the document you save is not necessarily byte-identical to the document you opened, even if you changed nothing. For workflows that care about preserving a file's exact structure, such as some archival or signing pipelines, that is a property to verify on your own files rather than assume. The README does not document a mode that disables repair on open, so if you need strict pass-through fidelity, that is a gap worth confirming against the documentation before committing.
pikepdf compared with pypdf, PyMuPDF and pypdfium2
The README's own comparison section names the alternatives and their different approaches. pypdf is pure Python and, in the README's words, well-suited for straightforward PDF tasks without compiled dependencies. That is the real dividing line: pypdf installs anywhere Python runs, with no wheel to match and no C++ underneath, while pikepdf ships compiled wheels and depends on qpdf being built into them. If your deployment target is an unusual platform or you cannot accept a compiled extension, pypdf's approach is the safer one. If you need qpdf's repair and structural handling, pure Python is exactly what you are giving up.
PyMuPDF takes a different angle again. The README notes it offers comprehensive functionality, and it is the library the README itself recommends for rendering PDFs to images. Rendering is the capability pikepdf deliberately does not have. If your pipeline needs to turn pages into pixels, you are in PyMuPDF or pypdfium2 territory, not here. The README mentions pypdfium2 specifically for permissively licensed PDF rendering, which is a licensing distinction rather than a functional one.
The choice therefore comes down to what the job is. Text and tables: pdfminer.six or pdfplumber. Pixels: PyMuPDF or pypdfium2. HTML to PDF: weasyprint or reportlab. Structural editing, repair, encryption and object-level work on files that already exist: pikepdf, with qpdf doing the heavy lifting.
Licence, maintenance and the cost of upgrading
pikepdf is licensed under MPL-2.0, which the README calls a "liberal license" compatible with most open and closed source projects. The repository keeps third-party-licenses/ with a component-to-license mapping, and pyproject.toml lists license-files as LICENSE.txt plus third-party-licenses/*, with a comment explaining that this covers the compiled libraries vendored into binary wheels by auditwheel, delocate and delvewheel. That matters if you distribute the wheel: the file-level copyleft in MPL-2.0 applies to pikepdf's own source, but the bundled qpdf and other compiled components carry their own terms, which the third-party licence directory enumerates. That is a description of what the repository ships, not legal advice; check the mapping against your own distribution model.
The repository is not archived, and the last push was on 2026-09-06. Recent releases in the same window include v10.13.0.post1 and v10.12.0, both dated 2026-09-06, and v10.4.0. The version in pyproject.toml is 10.13.0.post1, and requires-python is >=3.10, so Python 3.9 and earlier are out.
The upgrade cost of a binding like this is mostly about the underlying C++ library rather than the Python surface. qpdf is compiled into the wheels, so a qpdf fix arrives when a new pikepdf wheel is published, not when you update a system package. The flip side is that building from source means compiling libqpdf yourself, and the Makefile carries a warning about the editable install path: it instructs using uv pip rather than bare pip, because the venv's old pip "does not reliably recompile the scikit-build-core editable extension" and repeated pip install -e . can silently reuse a stale build and drop C++ changes. If you are working on the C++ side, that is a real trap.
Editorial conclusion
Adopt pikepdf when your job is to repair, merge, rotate, encrypt or inspect PDFs that already exist, and when you want qpdf's structural handling without shelling out to it. Do not adopt it to generate PDFs from HTML or templates (weasyprint or reportlab), to render pages to images (PyMuPDF or pypdfium2), or to pull text and tables out of documents (pdfminer.six or pdfplumber); the README names those as the wrong tools for this one. Verify first that your target platform has a binary wheel, since the wheels cover Linux, macOS and Windows on x86-64 and ARM64 including free-threaded CPython 3.14, and that your PDF does not rely on digital signature-based encryption, which the README says is not currently supported.
Frequently asked questions
What is pikepdf?
pikepdf is a Python library for reading, writing, repairing and transforming PDFs, powered by the qpdf C++ library. It exposes PDF dictionaries, arrays, streams and names as Python types, and repairs structural damage automatically when a file is opened.
How do I install pikepdf?
The README gives a single command, pip install pikepdf. Binary wheels cover Linux, macOS and Windows on x86-64 and ARM64/Apple Silicon, including free-threaded CPython 3.14, and no compiler is required. Building from source is documented separately and compiles libqpdf.
How do I use pikepdf to merge or rotate pages?
Create a new document with Pdf.new(), open each source with Pdf.open() and extend merged.pages with src.pages, then save. To rotate, iterate over pdf.pages and call page.rotate(180, relative=True) before saving. Both patterns come from the README.
Is pikepdf free?
pikepdf is licensed under MPL-2.0, which the README describes as a liberal license compatible with most open and closed source projects. Binary wheels also bundle compiled third-party components whose licenses are enumerated in the repository's third-party-licenses directory.
What is the difference between pikepdf and qpdf?
qpdf is the C++ library that does the parsing, repair and writing; pikepdf is the Python binding over it. pikepdf also exposes qpdf's Job API, so qpdf command-line operations can be run programmatically from Python instead of as a separate process.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/pikepdf-pikepdf)