Library / SDK
J-F-Liu/lopdf avatar
J-F-Liu/lopdf

lopdf: a Rust library for reading and writing PDF object graphs

A Rust library for PDF document manipulation.

2,263 stars300 forksRustMIT

At a glance

What is it?
A specification-shaped Rust crate for building, editing and encrypting PDF files, with no rendering engine and no page layout abstraction on top.
Who is it for?
lopdf suits the job of producing a correct PDF when you already know what the file should contain: report generators, forms, anything assembling a document from a template. It also suits the job of taking a PDF apart, because the objects are addressable rather than hidden behind a page API.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 5 days ago.
What is it written in?
Mainly Rust, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 7, 2026, and from our analysis. They are not legal advice.

Editorial analysis

A library shaped like the specification rather than like a document

lopdf gives a Rust program a typed handle on the objects inside a PDF file. The description in Cargo.toml is a Rust library for PDF document manipulation, the keywords are pdf, editing, manipulation and merge, and the three repository topics are pdf-document, rust and rust-library. The package ships no binary of its own. The tree holds src/, tests/, benches/, examples/ and a separate pdfutil/ directory, which suggests the command line tool is developed alongside the crate rather than published inside it.

What you get is the PDF object model: dictionaries, arrays, streams, and the indirect objects that the cross reference table points at. That is a lower level of interface than most PDF libraries present. There is no page object with a draw_text method. You construct a dictionary, register it to get an object id, attach a content stream, and wire up the page tree yourself. The README is candid about the reason for that shape, pointing at the PDF 1.7 reference document and, separately, the ISO 32000-2 specification for PDF 2.0, and describing the crate as a useful reference for understanding the PDF file format. Code that needs to reason about fonts, operators and object lifetimes will find that framing comfortable.

Building a one page document out of dictionaries and content operators

The README's creation example is the best documentation in the repository because every line is annotated with the part of the specification it comes from. A document starts with a version string, and object ids are handed out by the library rather than chosen by the caller:

rust
use lopdf::dictionary;
use lopdf::{Document, Object, Stream};
use lopdf::content::{Content, Operation};

let mut doc = Document::with_version("1.5");
let pages_id = doc.new_object_id();

Fonts are dictionaries with Type, Subtype and BaseFont keys straight out of the spec, and the `dictionary!` macro is what makes the nesting readable. Note that a font dictionary is not usable on its own: it has to be registered in a resource dictionary under a short name, and that resource dictionary has to reach the page. The example makes this explicit, calling font dictionaries triply nested.

The page content is a vector of operations, and each operation pairs a PDF operator name with its operands:

rust
let content = Content {
    operations: vec![
        Operation::new("BT", vec![]),
        Operation::new("Tf", vec!["F1".into(), 48.into()]),
        Operation::new("Td", vec![100.into(), 600.into()]),
        Operation::new("Tj", vec![Object::string_literal("Hello World!")]),
        Operation::new("ET", vec![]),
    ],
};

`BT` opens a text element, `Tf` selects the font named F1 at size 48, `Td` moves the text matrix, `Tj` prints the literal, and `ET` closes the element. One point in the comments is easy to miss and expensive to rediscover: PDF pages have Y equal to zero at the bottom, so the example moves to 600 to print near the top. Once the content stream is encoded, the page dictionary needs Type, Parent and Contents, and the root of the page tree needs an id that was reserved earlier so the child can point back at it.

The feature table is where your build decisions actually live

Two features are on by default. `chrono-clock` adds conversions to and from `DateTime<Local>`, which pulls in `iana-time-zone` and its per platform chain, and `rayon` parallelises object stream and cross reference parsing. Both are on because earlier versions behaved that way rather than because they are cheap.

The interesting row is the date backends. `chrono`, `jiff` and `time` are alternatives, not layers: each one supplies conversions for the same `DateTime` value, so enabling more than one only adds dependencies. The README states the reason, which is that a PDF date expresses a fixed offset from UT and never a named zone under ISO 32000-1 section 7.9.4. So `chrono` without its `clock` feature is enough to read one faithfully. Enabling none is supported too, since `Object::as_datetime` needs no backend and `DateTime::as_str` returns the raw string for a caller who would rather parse it.

The optional features that change what the library can do are `serde` for serialising the object model, `async` for Tokio based document loading, `embed_image` for raster images through the `image` crate, `font_embedding` for TrueType fonts through `skrifa`, and `wasm_js` to select getrandom's wasm backend, which encryption needs on that target. `jiff` is the one with a deployment consequence rather than a compile time one, because named zone resolution requires a timezone database, bundled into the binary on Windows and on any wasm target.

Encryption and compression coverage is visible in the dependency list

Cargo.toml is the clearest statement of what the crate actually implements, because a PDF library's capabilities tend to show up as its crypto and compression dependencies. The encryption story is `aes` at version 0.9 with `cbc` and `ecb` modes, plus `md-5` and `sha2` for the revision hashes a password based PDF needs and `stringprep` for string preparation. That is the standard security handler surface for PDF, and the presence of `rand` and `getrandom` fits a library that has to generate keys as well as derive them.

Compression is handled by `flate2` for the Deflate family, `weezl` for LZW, and `brotli-decompressor` for the streams that use it. Decompression is named explicitly rather than left implicit, and there is an example file called `examples/decompression_bomb.rs`, which tells you the crate's authors thought about a decompressed stream growing without bound.

Parsing is `nom` at version 8.0, which is a meaningful detail for anyone tracking the grammar. Structures are held in `indexmap`, page and range bookkeeping in `rangemap`, errors go through `thiserror`, and logging through `log`. With `rayon` available as an optional feature, object stream and cross reference parsing can be split across cores, which is the expensive part of opening a large file. The wasm target is a first class citizen here: `wasm-bindgen-test` sits in dev dependencies, which is why the README mentions the wasm_js backend rather than treating browser support as an afterthought.

Twenty-five example files, and the awkward ones are the useful ones

The examples directory holds twenty-five files, and reading the file names is a decent map of what real PDF work turns into. The obvious starting points are `create.rs` and `create_text_heavy_pdf.rs`, which extend the README's hello world towards documents with a lot of text in them. For people arriving with existing files, the interesting ones are `merge.rs`, `extract_text.rs`, `extract_toc.rs`, `replace_partial_text.rs` and `print_annotations.rs`. `compress_existing_pdf.rs` and `object_streams.rs` deal with the size and structure questions, and `analyze_pdf.rs`, `analyze_references.rs`, `analyze_page_contents.rs` and `analyze_object_streams.rs` are the diagnostic tools.

The rest is where the project is unusually open about its own rough edges. There are four separate debug files for compression and object streams (`debug_compression_detailed.rs`, `debug_compression_full.rs`, `debug_object_stream.rs`, `debug_save_with_objstreams.rs`), two named after the original failure reports that prompted them, `decompression_bomb.rs` for untrusted input, and one called `final_compression_test.rs`. A crate shipping that many debug examples is telling you that compression and object streams are where its bugs live, which matches what the release notes say.

Two examples also indicate scope beyond document assembly. `add_barcode.rs` and `replace_partial_text.rs` are the kind of edits people want from a PDF tool, and both are fiddly at the object level because they mean rewriting existing content streams and fonts. Whether they hold up on arbitrary real world files is a separate question from whether the code is there to read.

Release notes lean towards crash fixes, and the stated Rust version disagrees with Cargo.toml

The version at the top of Cargo.toml is 0.45.0, and a release with that tag was published on 2026-09-08. Its changelog is dominated by parsing and filter fixes rather than features: correcting the PNG average predictor calculation, bounding object graph recursion depth, fixing four crash bugs triggered by crafted PDFs, correcting an ASCII85 group value overflow, adding ASCIIHexDecode and RunLengthDecode decoding, and adding sub byte depth PNG and TIFF predictor support. There is also a fix for the `time` feature, which had a `From<Time>` implementation that never compiled, and CI changes enforcing rustfmt and broader clippy checks.

That list is a useful signal for anyone feeding lopdf untrusted input. A parser of a format designed in 1992 that gets handed files from the internet is the exact situation where crafted input bugs live, and a maintainer publishing four crash fixes in one release is being straightforward about the fact that this surface has had problems. Deletion of outlines and deduplicated error messages when chaining round out the non security entries.

The minimum Rust version is where the repository contradicts itself. The README states Rust 1.85 or later, required for edition 2024 features and object streams support, and Cargo.toml declares edition 2024 with `rust-version = "1.88"`. The README also tells you to check with `rustc --version` and to update with `rustup update`. Both cannot be satisfied by the same toolchain: a build pinned to 1.85 will be rejected by the manifest, so the README's number looks like it predates a bump. Treat the manifest as authoritative. The repository is not archived, with the last push on 2026-09-28, so this line is moving.

Where the crate ends and rendering begins

It is worth being plain about what lopdf does not do, because the name invites the wrong comparison. There is no rasteriser, no text layout engine, no line breaking, no table of contents rendering and no font shaping. `font_embedding` brings TrueType data into a document through `skrifa`, which is what you need to embed a font, not what you need to decide where a word goes on the page. Laying text out means computing positions yourself and emitting `Td`, `TJ` and `Tm` operators, which is exactly what the README example does.

The same applies to reading. Pulling text out of a PDF means walking content streams and interpreting text operators, which is what `extract_text.rs` does. That is tractable for a document the crate also produced, because the operators are the ones it emitted. Recovering a sensible reading order from a file produced by a word processor is a much larger problem, and the file name promises more than the operation is likely to deliver.

The alternatives are worth naming as categories rather than competitors. Full renderer toolkits such as MuPDF and Pdfium give you pixels or a layout tree at the cost of a C dependency. Higher level Rust generators produce finished documents without exposing the object model. lopdf sits below both of those, which is precisely the reason to reach for it when you need to merge, encrypt, annotate or template existing files, and the reason to skip it when you want a page image or a laid out paragraph. The README points outward to the specifications for the details the crate documentation does not settle.

Editorial conclusion

lopdf suits the job of producing a correct PDF when you already know what the file should contain: report generators, forms, anything assembling a document from a template. It also suits the job of taking a PDF apart, because the objects are addressable rather than hidden behind a page API. What it is not is a renderer, and the 2,255 stars on the repository should not be read as demand for one. The Cargo feature table, the twenty-five files in examples/ and the dependency list in Cargo.toml are where the real decisions sit, and the mismatch between the README's stated minimum Rust version and the one declared in Cargo.toml is worth checking before you pin a toolchain.

Frequently asked questions

Is lopdf a PDF renderer, or just an object model?

It is an object model. lopdf exposes dictionaries, arrays, streams and indirect objects so you can build and edit files yourself, and it has no rasteriser, no text layout engine and no font shaping. If you need a page image or a laid out paragraph, a rendering toolkit is the right tool.

How do I extract text from an existing PDF with lopdf?

You walk the page content streams and interpret the text operators, which is what the repository's examples/extract_text.rs does. This works well on files the crate produced, because the operators are known. Recovering reading order from a document written by another producer is a considerably harder problem.

Does lopdf handle encrypted and password protected PDFs?

Yes. The dependency list carries aes with cbc and ecb modes, md-5 and sha2 for the revision hashes, stringprep and rand, which is the standard security handler surface, and the examples directory includes both examples/encrypt.rs and examples/decrypt.rs.

What minimum Rust version does lopdf need?

The README says Rust 1.85 or later, but Cargo.toml declares rust-version 1.88 alongside edition 2024. The manifest is the enforceable figure, so a toolchain pinned to 1.85 will be rejected. The repository appears to have raised the minimum without updating the README.

Official sources

  1. Issues
  2. J-F-Liu/lopdf on GitHub
  3. License: MIT
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/j-f-liu-lopdf.svg)](https://hysenlabs.com/projects/j-f-liu-lopdf)