qpdf: Command-Line PDF Transformation Tool and C++ Library
qpdf: A content-preserving PDF document transformer
At a glance
- What is it?
- qpdf is a content-preserving PDF transformation tool that manipulates PDF structure through a command-line interface and a C++ library. It handles linearization, encryption, splitting, merging, and structural inspection but does not render PDFs or extract text.
- Who is it for?
- qpdf is the right tool for engineers who need to manipulate PDF structure programmatically: removing encryption, linearizing for web delivery, splitting pages, or merging files. It is not the right tool for text extraction, PDF rendering, or creating documents with rich content from scratch without supplying all content manually.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 4 days ago.
- What is it written in?
- Mainly C++, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What qpdf Solves and Who Uses It
PDF files have a complex internal structure: cross-reference tables, object streams, encryption dictionaries, and content streams can all require modification without changing the visible content of the document. qpdf is designed for exactly that class of operation: transforming PDF structure while preserving content.
The README describes the target user as "anyone who wants to do programmatic or command-line-based manipulation of PDF files." Typical use cases include removing password protection from a PDF the user has rights to modify, preparing a PDF for web delivery through linearization (which allows incremental loading in browsers), splitting a large PDF into individual pages, merging multiple PDFs into one, or inspecting the internal structure of a PDF for debugging.
qpdf is a low-level tool. It works directly with the PDF object model and does not interpret content streams to understand page layout, fonts, or images. The README is explicit: qpdf does not render PDFs or perform text extraction. Engineers who need to extract text from a PDF or convert it to another format need a different tool.
How qpdf Transforms PDF Structure
qpdf's core operation is reading a PDF file's object structure, applying specified transformations, and writing a new PDF file that contains the same content in the transformed structure. The README calls this "content-preserving": the visible output of the transformed PDF matches the input, but the internal representation changes.
Linearization restructures the PDF so that the first page is near the beginning of the file and necessary cross-reference information is placed to allow a browser to begin rendering before the full file is downloaded. Non-linearized PDFs require downloading the entire file before display. qpdf can both linearize a PDF and remove linearization.
Encryption support covers both reading encrypted PDFs (given the correct password or owner key) and writing PDFs with specified encryption settings. qpdf supports multiple PDF encryption standards corresponding to different Acrobat versions and key lengths.
Object stream compression affects file size. qpdf can add or remove cross-reference streams and object streams, which affect both the size and compatibility of the resulting file with older viewers.
The library exposes an API for C++ programs that need to perform these operations programmatically rather than through command-line invocations. This is the path for applications that generate or process PDFs as part of a workflow.
Installing qpdf and Verifying Releases
qpdf is available as pre-built binaries for common platforms, through system package managers, and buildable from source. The repository includes release signing using cosign from the sigstore project. Each release includes a sha256 checksum file; the checksum file itself is signed and can be verified:
cosign verify-blob qpdf-x.y.z.sha256 --bundle qpdf-x.y.z.sha256.sigstore \
[email protected] \
--certificate-oidc-issuer=https://github.com/login/oauthThe README specifies two valid signing identities: Jay Berkenbilt at [email protected] and Manfred Holger at [email protected]. The README-what-to-download.md file in the repository guides users on which artifact to select for their platform.
Starting with version 12.4.2, macOS binaries are available in the releases section. These binaries are not signed with an Apple Developer ID or notarized. macOS adds a quarantine attribute to downloaded files, which prevents execution until removed. The README provides the command to strip the quarantine attribute from the unzipped directory:
xattr -r -d com.apple.quarantine bin libVersion 12.4.2 was released on 2026-09-27. The project's last push to the repository was on 2026-09-26.
Using qpdf as a C++ Library
For C++ applications that need to incorporate PDF transformation, qpdf exposes its functionality as a library. The integration uses either pkg-config with the package name `libqpdf` or cmake with the package name `qpdf`.
A minimal CMakeLists.txt for a program that links against qpdf follows the standard cmake find_package pattern:
cmake_minimum_required(VERSION 3.16)
project(some-application LANGUAGES CXX)
find_package(qpdf)
add_executable(some-application some-application.cc)
target_link_libraries(some-application qpdf::libqpdf)The compiler requirement for building qpdf itself is C++20. Linking against the library in a consuming application requires only a C++17-compatible compiler, which is a more modest requirement.
The examples/ directory in the repository contains working C and C++ examples covering a wide range of API usage: creating PDF files (pdf-create.cc), splitting pages (pdf-split-pages.cc), overlaying pages (pdf-overlay-page.cc), setting form field values (pdf-set-form-values.cc), modifying PDF info metadata (pdf-mod-info.cc), counting pages (pdf-npages.cc), and several others. A C API example (pdf-c-objects.c and extend-c-api.c) demonstrates using qpdf from C code. A qpdf-job interface (qpdf-job.cc) shows driving transformations from a JSON job description rather than the C++ API directly.
The library dependencies are zlib and jpeg, both standard on Linux distributions. Optional crypto providers (gnutls, openssl) are conditionally required depending on the build configuration.
Zopfli Integration for Archival PDF Compression
qpdf supports optional integration with the zopfli compression library for generating flate-compressed streams. The zopfli algorithm produces smaller output than the default zlib compression, at the cost of significantly higher computation time: the README quotes approximately 100x slower than zlib.
The tradeoff makes zopfli appropriate for cases where output size matters and generation time does not, such as archival PDFs that will be stored and retrieved but not regenerated frequently.
Zopfli compression in qpdf is controlled through the `QPDF_ZOPFLI` environment variable: - If unset or set to `disabled`, zlib is used. - If set to `force`, zopfli is required and the process fails if zopfli was not compiled in. - If set to `silent`, zopfli is used when available and zlib is used silently otherwise. - Any other value uses zopfli when available and warns if not.
Building with zopfli support requires the zopfli library and its header to be installed at compile time.
Crypto Providers and Build Configuration
qpdf supports multiple crypto implementations for PDF encryption and decryption. The available providers are: - `gnutls`: uses the GnuTLS library; links libqpdf against GnuTLS. - `openssl`: uses OpenSSL. - `native`: the built-in implementation present in all versions before 9.1.0.
The native provider is no longer built by default if any external provider is found at build time, but it remains available. The README-hardening.md file in the repository covers compiler hardening flags and security build options.
Crypto providers can be selected both at compile time (which are available) and at runtime (which one is used). This gives operators the ability to swap providers without recompiling, within the set of providers that were enabled at build time.
The README notes that the exact provider built into a distribution's package of qpdf varies by distribution. Teams that need a specific provider for compliance or policy reasons should verify what their distribution ships or build from source with explicit configuration.
What qpdf Cannot Do and When to Look Elsewhere
The README is direct about qpdf's scope. It does not render PDFs, so it cannot produce images or screenshots of pages. It does not extract text, so it cannot be used for indexing document content or OCR workflows. It does not provide higher-level interfaces for working with page content, meaning it cannot reflow text, reformat tables, or interpret semantic structure.
For PDF creation with rich content, qpdf can create PDF files but requires the caller to supply all content in raw PDF form. The pdf-create.cc example in the examples/ directory demonstrates the API, but it requires constructing PDF objects directly. Teams who need a higher-level document creation API should use a library that abstracts the PDF object model.
For text extraction, tools such as pdftotext (from the poppler library) or Apache PDFBox provide that capability. qpdf occupies the structural manipulation layer below them.
qpdf is also not a PDF validator. While it can inspect PDF structure, it does not attempt to enforce full PDF specification compliance on its output. Documents it produces will generally be valid, but intentionally malformed input can sometimes produce output that is technically outside spec.
Poppler as the Nearest Alternative for Content Operations
The nearest alternative in the PDF tooling ecosystem is poppler, a PDF rendering and text-extraction library derived from Xpdf. The difference in approach is fundamental: poppler interprets PDF content to render pages and extract text; qpdf manipulates PDF structure without interpreting content.
For structural operations (linearization, encryption, splitting, merging), qpdf is the more direct tool. For content operations (rendering, text extraction, annotation display), poppler provides the capabilities qpdf deliberately excludes. The two are complementary rather than competing: a pipeline might use qpdf to decrypt and split a PDF, then poppler to extract text from the resulting pages.
qpdf's Apache-2.0 license is more permissive than poppler's GPL license, which matters for proprietary software that statically links against PDF libraries. Teams building closed-source applications that manipulate PDF structure should note this distinction.
Editorial conclusion
qpdf is the right tool for engineers who need to manipulate PDF structure programmatically: removing encryption, linearizing for web delivery, splitting pages, or merging files. It is not the right tool for text extraction, PDF rendering, or creating documents with rich content from scratch without supplying all content manually. Teams integrating qpdf into a C++ project should use the cmake package qpdf or pkg-config with libqpdf. Before deploying, verify which crypto provider (gnutls, openssl, or native) is compiled in, since the default changes depending on what is available at build time.
Frequently asked questions
What is qpdf used for?
qpdf is used for content-preserving transformations on PDF files: linearization for web delivery, adding or removing encryption, splitting and merging files, and inspecting PDF internal structure. It does not render PDFs or extract text.
Is qpdf safe?
qpdf releases are signed using cosign and include sha256 checksums. Valid signing identities are listed in the README as Jay Berkenbilt ([email protected]) and Manfred Holger ([email protected]). The README-hardening.md file documents compiler hardening options for security-conscious builds.
How do I install qpdf?
qpdf is available through system package managers on Linux distributions. Pre-built macOS binaries are available in the releases section starting with version 12.4.2; they require removing the quarantine attribute with xattr before running. The README-what-to-download.md file in the repository guides artifact selection.
How can I compress a PDF file using qpdf?
qpdf can recompress PDF streams using zlib by default. For smaller archival PDFs at the cost of much slower processing, qpdf can use the zopfli algorithm when built with zopfli support and the QPDF_ZOPFLI environment variable is set to a value other than 'disabled'.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/qpdf-qpdf)