pdfcpu: a Go PDF library and CLI with no external dependencies
PDF tooling for Go and the command line.
At a glance
- What is it?
- pdfcpu is an Apache-2.0 PDF processor written in Go, shipped as both a command-line tool and an importable library. It covers validation, merge, split, encryption, stamps and extraction, at the cost of a smaller feature set than mature C-based tools.
- Who is it for?
- Adopt pdfcpu if you want PDF handling inside a Go program or a single static binary with minimal external dependencies, and your work is validation, merge, split, optimize, encryption or stamping. Do not adopt it if you need to render pages to images, convert to or from other office formats, or repair files that most parsers reject: the command list does not include those operations.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 2 days ago.
- What is it written in?
- Mainly Go, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What pdfcpu is for, and who ends up using it
The project describes itself as a PDF processing library and command-line tool written in Go. The two halves matter separately. On the command line you get a single binary that validates, optimizes, splits, trims and merges PDFs, extracts and manipulates images, fonts and metadata, encrypts and decrypts, resizes and rotates pages, adds and removes stamps and watermarks, checks signature integrity, and manages attachments and portfolios. As a library you get the same operations callable from Go code, which is the part that has no equivalent in most PDF tooling.
The people who reach for it are usually Go developers who need PDF work inside a service, a batch job or a build pipeline, and who do not want to shell out to a C toolchain or link a native library. The README states the motivation plainly: comprehensive PDF processing implemented in Go for both individual files and automated batch processing, with a focus on correctness and independence from external dependencies. That last point is the design bet. Looking at go.mod, the direct dependencies are a JSON schema package, a TIFF decoder, a rune-width package, Cobra for the CLI, a YAML library, and three golang.org/x modules. There is no PDF engine underneath. Everything from the parser to the content stream writer is in the repository, under cmd/, pkg/ and internal/.
How the CLI and the Go API split the work
The repository is laid out around that split. cmd/ holds the Cobra-based command-line entry point, pkg/ holds the public packages including pkg/api, and internal/ holds the implementation the public packages sit on. The README points Go users at pkg.go.dev for the package documentation and at two directories, pkg/api/test and pkg/samples, for worked examples. That is the intended path into the library: read the API package docs, then read the tests, because the tests are the only place the README says to look for usage patterns.
The CLI itself is a thin layer. Every command listed in the README maps to a documented page on pdfcpu.io, and `pdfcpu [command] --help` is given as the way to get help without leaving the terminal. Commands take an output file first and inputs after it, which is visible in the merge example: `pdfcpu merge merged.pdf in1.pdf in2.pdf`. That ordering is worth noticing before you script anything, because it is the opposite of what some other PDF command-line tools do.
Config lives in a `config` command, and there is a separate command for changing the owner password and one for changing the user password, which tells you the encryption model distinguishes the two roles rather than treating a password as a single flag. The README does not document rollback or dry-run behaviour for mutating commands, so plan on keeping the originals.
Installing pdfcpu and running a first real job
The README does not inline install steps. It links to two pages on pdfcpu.io, one for the CLI and one for the Go API, and the Go module path is `github.com/pdfcpu/pdfcpu`. If you have a Go toolchain, the module path is enough to install the CLI:
go install github.com/pdfcpu/pdfcpu/cmd/pdfcpu@latestThat places a `pdfcpu` binary in your Go bin directory. The same command appears in the repository's Dockerfile, which builds the binary in a `golang:latest` stage and copies it into an `alpine:latest` image with the entrypoint set to `pdfcpu`.
docker build -t pdfcpu .
docker run -v $(pwd):/data -it --rm pdfcpu pdfcpu val test.pdfThe Dockerfile comments show this exact invocation, including `val` as the short form of validate and `/data` as the bound directory. Once the binary is on your path, validation is the first thing to run, because it tells you whether the rest of the toolkit will accept the file at all:
pdfcpu validate input.pdf
pdfcpu merge merged.pdf in1.pdf in2.pdfBoth lines are copied from the README's usage section. The first checks the file and reports problems; the second writes a new file called `merged.pdf` containing the two inputs in order. For the library side, the README directs you to the API documentation at pkg.go.dev and to the example directories rather than showing a code sample, so budget time to read the test files before writing your first call.
What pdfcpu will not do for you
The feature list is long, but the command table is the honest boundary. There is no render command, so turning a page into a PNG or JPEG is not something the CLI advertises; the `images` command extracts and manipulates images already embedded in a PDF, which is a different job. There is no conversion to or from DOCX, HTML or Markdown. There is no OCR. If your pipeline needs any of those, pdfcpu is the wrong tool and no amount of library glue will change that.
The second limitation is stated by the project itself. The README says pdfcpu supports PDF versions through PDF 2.0 (ISO 32000-2), and then adds that PDF 2.0 validation support is basic and continuously improving. Read that as a warning for anyone validating modern files: a clean `pdfcpu validate` on a PDF 2.0 document is weaker evidence than the same result on an older file.
The third is the release cadence. The most recent release in the repository is v0.16.0-rc.1, a release candidate, preceded by v0.15.0 and v0.14.0. A release candidate as the newest tag means the stable series you should pin to is the one before it. The README does not document a rollback path for commands that rewrite a file in place, so treat the output path as something you own.
pdfcpu against qpdf and the rest of the field
qpdf is the obvious comparison, and people search for it directly. Both are command-line PDF processors with validation, encryption and page assembly. The difference is the language and the embedding story. qpdf is C++ and is normally consumed as a binary or through bindings; pdfcpu is Go and is consumed either as a binary or as `github.com/pdfcpu/pdfcpu/pkg/api` inside a Go program. If your service is already Go and you want to avoid a subprocess and a native dependency in your container image, that is the whole argument for pdfcpu. If you are not writing Go, the argument mostly disappears, and qpdf's longer history in production is the counterweight.
Compared with Ghostscript, the split is different again. Ghostscript's centre of gravity is rendering and conversion, which pdfcpu does not offer; pdfcpu's centre of gravity is document structure, validation and assembly, which is a narrower but more predictable surface. The README's own framing supports that reading: it lists validation first and describes the project as focusing on correctness and independence from external dependencies, not on output fidelity for rendering.
Within Go specifically, the practical alternative is writing your own parser or pulling a thin wrapper around a C library. Neither is appealing. The cost of pdfcpu is that you inherit its bugs; the cost of the alternatives is that you inherit a build chain.
Licence, maintenance and what an upgrade costs
pdfcpu is Apache-2.0, with the licence text in LICENSE.txt. For most server-side and internal use that is a permissive licence with a patent grant and a requirement to preserve notices; if you redistribute a modified binary, read the file rather than this paragraph. The README also notes that the maintainer is a member of the PDF Association, which is context about the project's relationship to the specification, not a licence term.
On maintenance, the last push to the default branch was on 2026-09-20, and the newest tag is v0.16.0-rc.1 from the same day, so the repository is being worked on right now. That cuts both ways for adopters: you get active development, and you also get a moving target, with a release candidate sitting at the top of the tag list.
Upgrade cost is dominated by the CLI surface. There are roughly forty commands in the README table, each with its own documentation page, and the project publishes a changelog at pdfcpu.io/changelog. Before bumping a pinned version, read that changelog and re-run `pdfcpu validate` against a fixture set, because validation behaviour is the thing most likely to shift between minor versions. The library path adds a second axis: since pkg/api is the documented entry point, check the pkg.go.dev page for the version you are moving to rather than assuming the signature you compiled against last year still exists.
Editorial conclusion
Adopt pdfcpu if you want PDF handling inside a Go program or a single static binary with minimal external dependencies, and your work is validation, merge, split, optimize, encryption or stamping. Do not adopt it if you need to render pages to images, convert to or from other office formats, or repair files that most parsers reject: the command list does not include those operations. Before committing, run `pdfcpu validate` on a representative sample of your own files and check the changelog, because v0.16.0-rc.1 is a release candidate and the README states that PDF 2.0 validation support is basic and still improving.
Frequently asked questions
How do I install pdfcpu?
The README links to installation instructions on pdfcpu.io for both the CLI and the Go API, and the module path is github.com/pdfcpu/pdfcpu. The repository's Dockerfile installs the binary with `go install github.com/pdfcpu/pdfcpu/cmd/pdfcpu@latest` and copies it into an Alpine image.
How do I use pdfcpu?
Use it as a command-line tool or as a Go library. On the command line, `pdfcpu validate input.pdf` checks a file and `pdfcpu merge merged.pdf in1.pdf in2.pdf` combines two files; `pdfcpu [command] --help` prints help for any command. From Go, the README points at the pkg/api package documentation and the pkg/api/test and pkg/samples directories.
How can I merge PDF files in Golang with pdfcpu?
The CLI form is `pdfcpu merge merged.pdf in1.pdf in2.pdf`, with the output file named first. For Go code, the README directs you to the API documentation at pkg.go.dev and to the example directories in the repository rather than showing an inline snippet.
What is the difference between pdfcpu and qpdf?
Both are command-line PDF processors covering validation, encryption and page assembly. pdfcpu is written in Go and can also be imported as a library at github.com/pdfcpu/pdfcpu/pkg/api, while qpdf is a C++ tool normally consumed as a binary or through bindings.
What are alternatives to pdfcpu?
qpdf covers a similar command-line surface in C++, and Ghostscript differs in emphasis by concentrating on rendering and conversion, which pdfcpu's command list does not include. Within Go, the alternative is a wrapper around a native library or a parser of your own.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/pdfcpu-pdfcpu)