Model or dataset
yfedoseev/pdf_oxide avatar
yfedoseev/pdf_oxide

pdf_oxide: a Rust PDF core with 19 language bindings

The fastest PDF library for Python and Rust. Text extraction, image extraction, markdown conversion, PDF creation & editing. 0.8ms mean, 5× faster than industry leaders, 100% pass rate on 3,830 PDFs. MIT/Apache-2.0.

1,052 stars131 forksRustApache-2.0

At a glance

What is it?
pdf_oxide puts text extraction, image extraction, markdown conversion and PDF editing behind one Rust core with bindings for 19 languages, a CLI and an MCP server. The benchmarks are self-reported; the binding breadth and the licence split are the parts you can check yourself.
Who is it for?
Adopt pdf_oxide when you want one extraction and editing engine behind several runtimes, or when the AGPL terms of PyMuPDF are a blocker for a commercial product. Do not adopt it if you need OCR or scanned-document parsing: the README lists no OCR capability, and the benchmark explicitly excludes it.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 8 days ago.
What is it written in?
Mainly Rust, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem pdf_oxide targets: one PDF engine, many runtimes

Most teams end up with two or three PDF dependencies. A Python service uses one library for text, a Node worker uses another for markdown, and a Go batch job uses a third, and the three disagree about reading order on the same file. pdf_oxide is an attempt to collapse that into a single Rust core with idiomatic bindings on top. The README states the toolkit covers 20 languages, counting the Rust core plus 19 bindings: Python, Go, JavaScript/TypeScript, C#/.NET, Java, Kotlin, Scala, Clojure, Ruby, PHP, C++, Objective-C, Swift, Dart, R, Julia, Zig, Elixir and WASM.

The intended audience is narrow but real. It is for engineers building RAG or LLM ingestion pipelines who need markdown or plain text out of large PDF batches, and for teams that already ship a Rust core and want the same extraction behaviour in their Python and JavaScript clients. The README also ships a CLI and an MCP server, so an assistant-driven workflow is a first-class target rather than an afterthought. If you only ever read PDFs from one language and you do not care about licence terms, the binding breadth buys you nothing.

How the Rust core, the C ABI and the bindings fit together

The repository layout makes the architecture legible. The root Cargo.toml declares a workspace with four members: the crate itself, pdf_oxide_mcp, pdf_oxide_cli and pdf_oxide_jni. The language directories (python/, go/, js/, java/, kotlin/, csharp/, cpp/, objc/, dart/, elixir/, julia/, php/, clojure/) sit outside that workspace and are excluded from it, which tells you the bindings are maintained as separate packages rather than as cargo targets.

The bindings are built over a stable C ABI. The Makefile contains a c-header target that runs cbindgen against cbindgen.toml and writes include/pdf_oxide_c/pdf_oxide.h, and a c-header-check target that regenerates the header to a temp path and diffs it against the committed one. The comment above that target is unusually candid: the committed header had drifted to v0.3.24, missing every v0.3.50 and v0.3.51 C symbol, because nothing regenerated it. That history is worth knowing if you depend on the C or C++ binding, because it means the header was stale for a stretch of releases before the drift guard was added.

On the Python side the build backend is maturin, declared in pyproject.toml, and the same file records version 0.3.78 and requires-python >=3.9. The Rust core exposes a page-oriented document model: PdfDocument, then per-page lazy properties for text, chars, words, lines, tables and images, plus markdown() and html() conversions. Lazy properties matter here. If you only need page text, you are not paying for table or image extraction.

Installing pdf_oxide and extracting text from a first document

The README gives pip install pdf_oxide as the Python install path. The package name on PyPI is pdf_oxide, with an underscore, while the PyPI badge in the README points at pdf-oxide with a hyphen. Both spellings appear in the repository, so if one fails, try the other.

bash
pip install pdf_oxide

The README's quick-start example opens a document as a context manager, prints the page count, and reads text and markdown per page. The text and chars properties are documented as lazy, so nothing is extracted until you touch them.

python
from pdf_oxide import PdfDocument

with PdfDocument("paper.pdf") as doc:
    print(len(doc))                          # number of pages
    for page in doc:
        text = page.text                     # lazy property
        chars = page.chars                   # lazy property
        md = page.markdown(detect_headings=True)

Direct index access is also documented, including negative indices, so doc[0] and doc[-1] both resolve to pages.

python
from pdf_oxide import PdfDocument

doc = PdfDocument("paper.pdf")
page = doc[0]
text = page.text

If you would rather not write code at all, the CLI covers the same ground. The README shows these four subcommands verbatim:

bash
pdf-oxide text document.pdf
pdf-oxide markdown document.pdf -o output.md
pdf-oxide search document.pdf "pattern"
pdf-oxide merge a.pdf b.pdf -o combined.pdf

On macOS the CLI installs through a Homebrew tap, and the README notes that the same tap includes the MCP server binary:

bash
brew install yfedoseev/tap/pdf-oxide

For Rust, the crate is pdf_oxide and the README suggests version 0.3 in Cargo.toml. The Rust API is index-based rather than iterator-based: extract_text(0), extract_images(0) and to_markdown(0, Default::default()) all take a page index. That is a real difference from the Python API, where you iterate pages.

The benchmark numbers are self-reported and the corpus is public

The README claims 0.8ms mean per document, 5x faster than PyMuPDF and 15x faster than pypdf, with a 100% pass rate across 3,830 PDFs. Those figures come from the project's own benchmark, run on three public suites: veraPDF (2,907 PDFs), Mozilla pdf.js (897) and SafeDocs (26). The methodology line says single-thread, 60-second timeout, no warm-up, and text-extraction libraries only, with no OCR.

Treat the headline as a starting hypothesis, not a result. The corpus is public, which is the good part: you can reproduce the run against the same three suites. The comparison table mixes licences freely, and that matters more than the millisecond column. PyMuPDF is AGPL-3.0, pypdfium2 is Apache-2.0, pdftext is GPL-3.0, pdfminer, pdfplumber and markitdown are MIT, pypdf is BSD-3. A 4.6ms mean under AGPL-3.0 and a 0.8ms mean under MIT are not interchangeable numbers for a commercial product, and the table does not pretend otherwise, but it also does not draw the conclusion for you.

The text-quality claim is harder to check: 99.5% text parity against PyMuPDF and pypdfium2 across the corpus, and text extracted from 7 to 10x more hard files than it misses versus any competitor. Parity is a two-sided measure. Being 99.5% identical to another library means you inherit roughly the same reading-order decisions on the files where both agree, and diverge only on the remainder.

Where pdf_oxide is the wrong choice

There is no OCR. The benchmark methodology states plainly that it covers text-extraction libraries only, no OCR. The Cargo.toml lists an ocr feature gated behind tract-onnx and pdfium-render, but the README does not document an OCR workflow, and the pyproject.toml classifier list includes the keyword ocr without describing a scanned-document path. If your input is scans, this is not the tool, and the feature flag existing in Cargo.toml is not the same as a documented pipeline.

The project is also young. The Python package classifies itself as Development Status 4 - Beta, and the release history shows why that label is honest. v0.3.78, dated 2026-09-08, is titled "Correctness under audit" and covers 144 defects across rendering, text, reading order, files and colour. v0.3.76 fixed CCITT /ImageMask XObjects that were misreading compressed data as raw stencil rows, and corrected spatial table-cell ownership at singleton-span and boundary cases. Those are correctness fixes in extraction paths, not cosmetic changes. A library that is still closing 144 defects in one release will keep changing extraction output between minor versions, and if you snapshot markdown into a search index you should expect to re-index after an upgrade.

One more boundary: the README does not document rollback or a migration guide between minor versions. There is a CHANGELOG.md in the repository root, and the release notes are detailed, but the README itself is silent on downgrade paths. Pin your version.

pdf_oxide compared with PyMuPDF and pypdfium2

The most direct alternative is PyMuPDF, and the difference is not only speed. PyMuPDF is AGPL-3.0, which for a closed-source network service means either complying with the AGPL or buying a commercial licence from Artifex. pdf_oxide is MIT/Apache-2.0, and the README states the licence as MIT while the repository carries both LICENSE-MIT and LICENSE-APACHE and pyproject.toml declares "MIT OR Apache-2.0". If licence compatibility is your blocker, that is the real difference, and the 5x speed claim is secondary.

pypdfium2 is the closer comparison on terms: Apache-2.0, and the benchmark lists it at 4.1ms mean with a 99.2% pass rate. It wraps PDFium, the Chromium PDF engine, so its rendering behaviour tracks a browser rather than a purpose-built extractor. pdf_oxide's own rendering is written in Rust and is the part still being audited, per the v0.3.78 notes. If you need pixel-accurate rendering more than text extraction, a PDFium-backed library has a longer track record behind it.

For pure text, pypdf is MIT and BSD-3-adjacent in practice, and the benchmark puts it at 12.1ms. It is slower and less accurate on the corpus, but it is a mature, widely deployed dependency with no Rust toolchain in the build path. If your environment cannot compile or install a native extension, that constraint decides the choice before any benchmark does.

Maintenance, releases and what the licence split means for you

The repository is not archived, and the last push was on 2026-09-08, the same day as the v0.3.78 release. Release cadence in the recent window is uneven: v0.3.76 and v0.3.77 landed on 2026-07-27 and 2026-07-28, then v0.3.78 on 2026-09-08. The v0.3.77 notes describe prepare_search() and clear_search_index() being extended from the Rust core to every first-party binding, and a fix so that extract_text(), to_markdown() and to_plain_text() no longer silently drop /Artifact-tagged content such as running headers and footers. That last change is a behaviour change, not a bug fix from the caller's perspective: if you were relying on headers being stripped, they now appear.

Upgrade cost is the thing to budget for. The C header drift described in the Makefile is the cautionary example: the committed header fell behind by dozens of releases before c-header-check was added. If you consume the C ABI directly, run the equivalent of make c-header-check in your own pipeline rather than trusting the committed header.

On licensing, the repository ships LICENSE-MIT, LICENSE-APACHE, a NOTICE file, a CLA.md and a TRADEMARKS.md, and pyproject.toml declares "MIT OR Apache-2.0". Dual MIT/Apache-2.0 is the standard permissive Rust arrangement and permits commercial and closed-source use, but the presence of a CLA and a trademark file means the project reserves some rights around the name. Read TRADEMARKS.md before you ship a product called anything resembling pdf_oxide. This is a description of the files present, not legal advice.

Editorial conclusion

Adopt pdf_oxide when you want one extraction and editing engine behind several runtimes, or when the AGPL terms of PyMuPDF are a blocker for a commercial product. Do not adopt it if you need OCR or scanned-document parsing: the README lists no OCR capability, and the benchmark explicitly excludes it. Before committing, verify the pass rate on your own corpus rather than the 3,830-file aggregate, and check that the Python wheel for your interpreter version is published on PyPI.

Frequently asked questions

What is a good library to generate PDFs with pdf_oxide?

The README lists PDF creation as a feature alongside text and image extraction, and the repository ships examples such as create_pdf_from_markdown.rs, create_pdf_with_images.rs, html_to_pdf.rs and create_pdf_with_custom_font.rs. The feature table groups creation into documents, tables, graphics, templates and images.

Is there a way to extract only text from a PDF with pdf_oxide?

Yes. The Python API exposes a lazy text property on each page, and the CLI has a pdf-oxide text document.pdf subcommand. The Rust API calls extract_text(page_index), and there is also to_plain_text() alongside to_markdown().

Does pdf_oxide support OCR for scanned PDFs?

The README does not document an OCR workflow. The benchmark methodology states it covers text-extraction libraries only, with no OCR, and while Cargo.toml lists an ocr feature gated behind tract-onnx and pdfium-render, no usage is described.

Which languages can call pdf_oxide?

The README states the toolkit covers 20 languages: the Rust core plus 19 bindings, including Python, Go, JavaScript/TypeScript, C#/.NET, Java, Kotlin, Scala, Clojure, Ruby, PHP, C++, Objective-C, Swift, Dart, R, Julia, Zig, Elixir and WASM.

What licence is pdf_oxide released under?

The repository carries LICENSE-MIT and LICENSE-APACHE, and pyproject.toml declares "MIT OR Apache-2.0". The README describes it as MIT licensed and notes it can be used in commercial and open-source projects.

Official sources

  1. License: Apache-2.0
  2. Project website
  3. README
  4. Releases
  5. yfedoseev/pdf_oxide on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/yfedoseev-pdf-oxide.svg)](https://hysenlabs.com/projects/yfedoseev-pdf-oxide)