# firecrawl/pdf-inspector: four document types, one confidence score, and a benchmark run by version 0.2.6

> pdf-inspector decides whether a PDF is worth OCR-ing before you pay for it, then extracts positioned text and Markdown from the pages that do not need it. The routing decision is the product: the default path stays pure extraction, OCR is a second opt-in call, and the Python and Node bindings do not return the same type strings.

**firecrawl/pdf-inspector** — GitHub describes it as Fast Rust library for PDF inspection, classification, and text extraction. Intelligently detects scanned vs text-based PDFs to enable smart routing decisions.. The repository metadata lists Rust as its primary language. The metadata lists the MIT license. This article stays within the project description and details documented in the GitHub repository README.

- Repository: https://github.com/firecrawl/pdf-inspector
- Website: https://firecrawl.github.io/pdf-inspector/
- Stars: 19,444 · Forks: 1,315
- Language: Rust
- License: MIT
- Published: 2026-08-13 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/firecrawl-pdf-inspector

## The default path never loads PDFium, and OCR is a separate call

The design goal is stated in the second paragraph: handle text-based PDFs locally in under 200ms and skip expensive OCR services for the roughly 54% of PDFs that do not need them. The implementation of that is two entry points rather than one.

```python
import pdf_inspector

result = pdf_inspector.process_pdf("document.pdf")
print(result.pdf_type)   # "text_based", "scanned", "image_based", "mixed"
print(result.markdown)   # Markdown string or None

# Selective OCR; clean text PDFs do not load the external OCR runtime.
ocr = pdf_inspector.process_pdf_with_ocr("document.pdf")
print(ocr.pages_routed_to_ocr)
```

The comment in that snippet is the contract. The default Rust and browser builds stay pure extraction, and although the native Python and Node packages include the OCR integration, PDFium, ONNX Runtime and the model files remain external and are touched only when a page is actually routed to OCR. The consequence is that installing the Python package does not give you a working OCR path: something has to supply PDFium and ONNX Runtime, and pages_routed_to_ocr is the only place that shows whether they were used. A deployment that assumes pip install is the whole setup will discover the missing pieces on the first scanned document.

## The same four document types come back spelled differently in Python and Node

The four classifications are TextBased, Scanned, ImageBased and Mixed, detected in roughly 10 to 50 milliseconds by sampling content streams, and returned with a confidence score between 0.0 and 1.0 plus per-page OCR routing. How they are spelled depends on the binding.

Python returns them lowercased with underscores, for example "text_based", "scanned", "image_based" and "mixed". Node returns them in the same capitalization as the feature list, "TextBased", "Scanned", "ImageBased", "Mixed".

```javascript
import { readFileSync } from 'fs';
import { processPdf, processPdfWithOcr } from '@firecrawl/pdf-inspector';

const pdf = readFileSync('document.pdf');
const result = processPdf(pdf);
console.log(result.pdfType);   // "TextBased", "Scanned", "ImageBased", "Mixed"
console.log(result.markdown);  // Markdown string or null
```

So the same decision, expressed twice, with a naming convention that is not merely camel versus snake. Any comparison you write against one binding silently never matches in the other, and since a routing rule that fails open just sends everything to OCR, the symptom is a bill rather than an exception. The docs are separate files for this reason, docs/python.md and napi/README.md, so a port has to be read twice rather than translated once.

## The published comparison was produced by 0.2.6, and the crate is at 1.25.2

The benchmark table is specific about its inputs: the opendataloader-bench corpus of 200 PDFs, only local engines without model-based PDF parsing, OCR disabled, scores from 0 to 1. It was refreshed on July 31, 2026 on an Apple M4 Pro, and speed is the median of five alternating complete corpus runs after an excluded warm-up run, each parser handling documents sequentially in a single process. The engine versions are named, and pdf-inspector is listed as 0.2.6, while both pyproject.toml and Cargo.toml now declare version 1.25.2 and the newest release, v1.25.2, is dated 2026-09-28.

That gap is the thing to carry away from the table. The ranking is directional evidence about a build from the summer, not a measurement of what you would install today, and the paired harness described in docs/benchmarking.md exists so you can compare two local builds against the same corpus and evaluator revision yourself. The full configuration, per-document predictions and evaluator output live in a reproducible results branch.

The three newest tags also tell you where the work went: v1.25.0 is async positions and underline detection speed, v1.25.1 is font decoding fixes, and v1.25.2 is script sizes and ToUnicode repair checks, all three published on 2026-09-28.

## Headings stop at H4, and heading detection is the one row another engine wins

The Markdown converter handles headings from H1 to H4 derived from font size ratios, bullet, numbered and letter lists, code blocks detected from monospace fonts, tables both rectangle-based and heuristic, bold and italic, URL linking, and page breaks. Those are the whole feature set of the converter.

Two ceilings follow directly. There is no H5 or H6, so a document nested deeper than four heading levels comes out flattened, and the only signal that it happened is the shape of the Markdown. And code block detection is a font test, not a structural one, so code that a PDF styles in a proportional font arrives without fences.

The benchmark table is honest about the same weakness. In the Headings column, scored MHS, pdf-inspector is at 0.788 while liteparse is at 0.811, so on that one metric another engine is ahead, and it is ahead on the same table where pdf-inspector leads overall at 0.875, on reading order at 0.915, and on tables at 0.814, with the fastest complete run at 0.470s. If heading fidelity is your requirement, that row is the one to argue with.

## Broken font encodings get flagged for the caller, not repaired in place

Encoding problems are treated as a routing signal. Encoding issue detection automatically flags broken font encodings so callers can fall back to OCR, and the CID font support that makes detection possible covers ToUnicode CMap decoding for Type0 and Identity-H fonts across UTF-16BE, UTF-8 and Latin-1 encodings. TrueType parsing is pulled in specifically to extract cmaps from Identity-H fonts.

The consequence is about error handling. A flagged encoding is not an exception and not repaired text, so a caller that ignores the flag receives whatever bytes the font maps to, which is the kind of output that looks like a decoding bug in your own pipeline rather than a signal to try something else. Any integration has to treat the flag as an input to routing, not as a warning to log.

This is also the area the recent tags point at, with v1.25.1 named font decoding fixes and v1.25.2 named ToUnicode repair checks. Two examples ship for it, examples/rtl_fixtures.rs and examples/symbolic_font_fixtures.rs, which is a fair signal that right-to-left lines and symbolic fonts are where the hard cases live. Read those fixtures before you assume a font-heavy PDF will come out clean.

## Rotated text keeps its rotation angle, which breaks consumers that assume zero

Extraction is position aware and returns font information, X and Y coordinates, and an automatic multi-column reading order, with automatic detection of newspaper-style columns, sequential reading order, and RTL support. One detail is worth singling out because it is a deliberate refusal to normalize: rotated runs, the margin stamps and chart axis titles that show up in real documents, keep a true axis-aligned box and report their rotation angle instead of collapsing to zero width.

That refusal is correct and it still costs you. The CLI exposes exactly this through a flag:

```bash
# Convert PDF to Markdown
pdf2md document.pdf

# JSON output (for piping), with the document information (title, author, producer, dates, ...)
pdf2md document.pdf --json

# Positioned TextItem JSON (coordinates relative to the visible page box): axis-aligned box, rotation, font, paint (fill and stroke colour, render mode), underline metadata
pdf2md document.pdf --items-json

# Raw markdown only (no headers)
pdf2md document.pdf --raw
```

A consumer of --items-json has to read rotation, and has to remember that coordinates are relative to the visible page box rather than the media box. Code that assumes every run is horizontal, or that assumes width is always positive, misplaces exactly the stamps and axis titles that a layout-aware pipeline was supposed to handle.

## The browser build ships extraction, and the crate carries its own CMaps

The WebAssembly package runs the same Rust parser locally in browsers and Web Workers, with embedded CMaps and no server round trip. The documented usage is three steps: install `@firecrawl/pdf-inspector-wasm`, call init, then hand a fetched buffer to processPdf. The four type names come back capitalized, as in the Node binding.

What is absent from that example is the OCR call. The selective OCR line in the feature list names Rust, the CLI, Python and Node, and not the browser, and the browser example shows only processPdf. So a client-side application gets classification and Markdown conversion and nothing for a scanned page, which is the one case where a scanned PDF needs help.

The reason the CMaps are embedded rather than read at runtime is spelled out in Cargo.toml. A cargo-installed binary's CARGO_MANIFEST_DIR is the registry checkout, which is not a stable runtime location and does not exist at all for a published wheel, so the include_dir crate embeds them on every target. The same file notes the cost on the other side: crates.io caps uploads at 10 MiB, tests/fixtures alone exceeds that, so the include allowlist ships src, external/bcmaps, the Rust API doc, the license and the type stub, and nothing else.

## Consumers can build on Rust 1.88 while the project pins 1.98 for itself

Two version numbers, and they mean different things. Cargo.toml sets rust-version to 1.88, annotated as the oldest compiler the crate builds on, tied to a language feature the code uses. rust-toolchain.toml pins 1.98, and the comment there says that pin is what the project develops and releases with. So a downstream project on an older toolchain can still depend on the crate, while the maintainers build against something newer.

The Python side is configured for the same kind of reach. pyproject.toml requires Python 3.8 or newer and the pyo3 dependency is configured with the abi3-py38 feature, which is how a single wheel covers that whole range instead of one wheel per interpreter. The build backend is maturin, and the maturin feature set is just "python". The library target is declared as both lib and cdylib, which is what lets the same source serve the Rust crate and the extension module.

Version numbers are kept in two files that have to agree, with a comment naming the command that synchronizes them and noting that CI publishes automatically when that change lands on main. So a release is a coordinated edit, not a single tag.

## Conclusion

Use pdf-inspector if your pipeline is full of native-text PDFs and the thing you want to avoid is sending all of them to an OCR service, because classification and Markdown conversion are the parts that work without any external runtime. Do not treat it as a general PDF tool, and do not read the benchmark as a measurement of the current build. Before you wire it in, check four things: which of the four type strings your binding returns, since Python and Node disagree on spelling and a comparison written for one will never match in the other; that the OCR runtime is actually present, since PDFium, ONNX Runtime and the model files are external and only touched when a page is routed; that your consumers handle the flagged-encoding case, because a broken font encoding is reported rather than repaired; and that heading-heavy documents still meet your bar, since headings stop at H4 and the published table is the one row where another engine is ahead. The library is MIT licensed, the newest release is v1.25.2 dated 2026-09-28, and the last push was 2026-09-30.

## FAQ

### Is there a way to inspect a PDF?

With pdf-inspector, yes. It classifies a document as TextBased, Scanned, ImageBased or Mixed with a confidence score, extracts positioned text with font information and coordinates, and converts to Markdown. The CLI can also emit positioned TextItem JSON through the --items-json flag.

### What are the signs of a PDF virus?

The library makes no claims about malicious files, but it does expose document metadata you can look at. The --json output of the pdf2md command includes the document information, among them title, author, producer and dates, and --items-json exposes paint details, meaning fill and stroke colour and render mode, for each text item.

### How to make a PDF unreadable?

pdf-inspector cannot do that, since it only reads PDFs and writes Markdown or JSON. Its outputs are a Markdown string, JSON document information, positioned TextItem JSON, and raw Markdown with no headers. Making a PDF unreadable would require a writer, which is outside what this project does.

### How to download a PDF using Inspect?

Nothing here downloads a PDF from a web page. Every entry point takes a file you already have: process_pdf("document.pdf") in Python, readFileSync followed by processPdf in Node, and in the browser package a fetch of the file followed by processPdf on the resulting buffer.

## Sources

- [Official documentation](https://firecrawl.github.io/pdf-inspector/)
- [Official README](https://github.com/firecrawl/pdf-inspector#readme)
- [Project repository](https://github.com/firecrawl/pdf-inspector)
- [Release notes](https://github.com/firecrawl/pdf-inspector/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/firecrawl-pdf-inspector
