pdf-inspector: a Rust library that decides whether a PDF needs OCR before you pay for it
GitHub describes it as Fast Rust library for PDF inspection, classification, and text extraction. Intelligently detects scanned vs text-based PDFs to enable smart routing decisions.. The repository metadata lists Rust as its primary language. The metadata lists the MIT license. This article stays within the project description and details documented in the GitHub repository README.
At a glance
- What is it?
- pdf-inspector classifies PDFs as text-based, scanned, image-based or mixed, then extracts structured Markdown with tables and reading order. It is built for pipelines that want to skip OCR on the majority of documents, and it is honest about being a local text engine rather than a full OCR stack.
- Who is it for?
- Adopt pdf-inspector if your pipeline ingests native-text PDFs (reports, papers, invoices, contracts) and you want classification plus Markdown without paying OCR latency on every file. Do not adopt it as a replacement for a full OCR service: scanned and image-based documents are detected and routed, not solved, unless you opt into the selective OCR path and supply PDFium, ONNX Runtime and model files yourself.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 3 days ago.
- What is it written in?
- Mainly Rust, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem: most PDFs do not need OCR, but you cannot tell which ones by looking
A PDF is a container, not a format. The same file extension covers a digitally generated invoice with embedded fonts and a 300-page scan of a fax. Pipeline builders usually solve this by sending everything through OCR, which is slow and expensive, or by guessing from file size and page count, which fails quietly. pdf-inspector exists to make that decision explicit. The README states the project was built by Firecrawl to handle text-based PDFs locally in under 200ms and to skip expensive OCR services for the roughly 54% of PDFs that do not need them. That 54% figure is the project's own framing, not an independent measurement, and it will not match every corpus. The library is aimed at teams running document ingestion at volume: RAG pipelines, search indexes, archive digitisation, and anyone who has watched an OCR bill scale with documents that were already machine-readable. It returns a classification (TextBased, Scanned, ImageBased or Mixed), a confidence score between 0.0 and 1.0, and per-page OCR routing, so the caller decides what to do rather than the library deciding for them.
How classification and extraction actually work
The design choice that matters most is stated plainly in the feature list: the document is parsed once and shared between detection and extraction. Classification samples content streams rather than rendering pages, which is why the README quotes 10-50ms for detection. Sampling is cheap but it is also the source of the library's main failure mode, because a PDF whose first pages are text and whose later pages are scans can be classified from an unrepresentative sample. The Mixed category and per-page routing exist to handle that case, and the confidence score is the signal to check. Extraction is position-aware: text items carry font information and X/Y coordinates, and the library derives reading order from those coordinates, including newspaper-style multi-column layouts and RTL text. Markdown conversion is built on top of that geometry rather than on a text dump: headings are inferred from font size ratios (H1 through H4), lists from bullet, numbered and letter markers, code blocks from monospace font detection, and tables from two separate paths. Rectangle-based table detection reads PDF drawing operations, while a heuristic path infers tables from text alignment. That dual approach is why the benchmark table shows a strong TEDS score, but it also means table quality depends on how the source PDF was produced. Financial tables drawn as rectangles behave very differently from tables that are just aligned text. CID font handling covers Type0 and Identity-H fonts via ToUnicode CMap decoding, with UTF-16BE, UTF-8 and Latin-1 encodings, and the crate ships an external/bcmaps directory that tounicode.rs loads at runtime relative to CARGO_MANIFEST_DIR. Encoding issue detection flags broken font mappings so the caller can fall back to OCR instead of silently emitting garbage.
Installing pdf-inspector and running a first extraction
The Python package is the shortest path to a working result. Install it with pip, then call process_pdf on a file path. The result object exposes pdf_type and a markdown string, which is None when extraction produces nothing usable. The README gives this example:
pip install pdf-inspectorimport pdf_inspector
result = pdf_inspector.process_pdf("document.pdf")
print(result.pdf_type) # "text_based", "scanned", "image_based", "mixed"
print(result.markdown) # Markdown string or NoneIf you want the library to route only the pages that need it, call process_pdf_with_ocr instead. The README notes that clean text PDFs do not load the external OCR runtime at all, so the cost is only paid when a page is actually routed:
ocr = pdf_inspector.process_pdf_with_ocr("document.pdf")
print(ocr.pages_routed_to_ocr)The CLI is installed separately from crates.io and is useful for inspection before you write any integration code. Note that the CLI flags are documented in the README but the detection-only flag is truncated there, so check the CLI help output for the full set:
cargo install pdf-inspector
pdf2md document.pdf --json
pdf2md document.pdf --select-pages 1,3,5-10
pdf2md document.pdf --compactThe --json flag gives machine-readable output for piping, --select-pages restricts processing to specific pages, and --compact collapses long dot leaders and similar source padding for token-efficient output. For Node.js, the package is @firecrawl/pdf-inspector and processPdfWithOcr is asynchronous, which the README describes as running off the event loop. The browser build is a separate package, @firecrawl/pdf-inspector-wasm, with embedded CMaps and no server round trip.
Where pdf-inspector is the wrong tool
The README is unusually direct about the boundary: the default Rust and browser builds remain pure extraction, and OCR is opt-in for Rust and CLI consumers. Native Python and Node packages include the OCR integration, but PDFium, ONNX Runtime and model files remain external and are touched only when a page is routed to OCR. That means you are responsible for supplying those dependencies. If your documents are predominantly scanned, you are paying for a classifier and a text extractor and then still running OCR on nearly every page. The selective OCR path runs PP-OCRv6 Small locally, which the release notes for v1.15.0 describe as selective OCR for scanned and mixed PDFs, but the model files are not bundled. A second boundary is the benchmark itself. The README explicitly says only local engines without model-based PDF parsing are shown and that OCR was disabled. So the 0.875 overall score says nothing about how the library compares to a model-based parser on messy scans. Third, the benchmark was run on an Apple M4 Pro with pdf-inspector 0.2.6, while the current crate version is 1.19.0 and the Python package is 1.19.0. The version skew between the benchmarked build and the current release is large enough that the numbers should be treated as directional. Finally, if you need layout fidelity for complex scientific figures or vector graphics, text and table extraction will not give you that.
pdf-inspector compared with Docling and PyMuPDF4LLM
People searching for pdf-inspector alternatives usually land on Docling or PyMuPDF4LLM, and the difference is architectural. PyMuPDF4LLM is a Python wrapper over a mature C PDF engine; it is convenient and widely deployed, but the README's benchmark table lists it at 17.117s for 200 documents against pdf-inspector's 0.470s, with a tables score of 0.401 versus 0.814. Those numbers come from the project's own harness, refreshed on July 31, 2026, so treat them as a comparison run by an interested party rather than a neutral result. The mechanism behind the gap is that pdf-inspector is written in Rust with lopdf's parallel parser enabled on native targets, whereas PyMuPDF4LLM processes sequentially in Python. Docling takes a different route entirely: it is a model-based document conversion toolkit, and the README's benchmark deliberately excludes model-based parsers, so the two projects are not measured on the same axis. If your documents are clean native-text PDFs, the model-based approach adds weight you may not need. If your documents are scanned or heavily graphical, the model-based approach is doing work that pdf-inspector explicitly routes elsewhere. The other engines in the table, liteparse and opendataloader, sit closer to pdf-inspector in approach: liteparse scores higher on headings (0.811 versus 0.788) while pdf-inspector leads on tables and speed, and opendataloader is roughly five times slower on the same corpus.
Maintenance, licence and what upgrading involves
pdf-inspector is MIT licensed, and the licence field appears consistently across the Cargo manifest, the pyproject metadata and the repository root. MIT is permissive: you can embed the crate or the bindings in a commercial product, but you inherit no patent grant and no warranty, and the licence text itself is the authority rather than anything written here. On maintenance, the last push to main was on 2026-08-17, the same day v1.15.0 landed, and v1.14.2 shipped four days earlier on 2026-08-13. The release cadence over that window is frequent. The upgrade cost is dominated by version synchronisation rather than API churn. The pyproject.toml carries an explicit comment that package versions must be kept in sync with scripts/version.py and that CI publishes automatically when the synchronised change lands on main. The packages-2026-08-10 release shows the four artefacts versioning independently: Rust 0.1.8, Python 0.2.7, Node 1.13.0, WASM 0.1.4. The current manifests show Rust and Python both at 1.19.0, so the numbering has since converged, but the fact that it can diverge means you should pin the binding you depend on and check that its underlying Rust version matches what you tested. The crate requires Rust 1.88 or newer, and the development toolchain is pinned separately at 1.98 in rust-toolchain.toml. The crates.io upload has an explicit include allowlist because the 10 MiB cap would otherwise be exceeded by test fixtures alone.
Editorial conclusion
Adopt pdf-inspector if your pipeline ingests native-text PDFs (reports, papers, invoices, contracts) and you want classification plus Markdown without paying OCR latency on every file. Do not adopt it as a replacement for a full OCR service: scanned and image-based documents are detected and routed, not solved, unless you opt into the selective OCR path and supply PDFium, ONNX Runtime and model files yourself. Before committing, verify three things on your own corpus: the confidence scores your documents return, whether your tables are rectangle-based or heuristic, and whether the Python or Node package version matches the Rust crate you benchmarked against.
Frequently asked questions
Is there a way to inspect a PDF with pdf-inspector?
Yes. The library classifies a PDF as text-based, scanned, image-based or mixed, returns a confidence score between 0.0 and 1.0, and can extract position-aware text as Markdown. The CLI exposes this directly through pdf2md, including a --json mode for machine-readable output.
What are the signs of a PDF virus, and does pdf-inspector detect them?
pdf-inspector does not scan for malware. It inspects PDF structure for classification and text extraction, and its encoding issue detection flags broken font encodings so callers can fall back to OCR. The README and repository files describe no antivirus or threat-detection capability.
How do I make a PDF unreadable, and can pdf-inspector still read it?
No redaction or obfuscation feature is documented. pdf-inspector reads what the PDF exposes: text-based files are extracted directly, while scanned and image-based pages are classified and routed to OCR only if the caller opts in.
How to download a PDF using Inspect?
pdf-inspector does not download PDFs. It operates on bytes or a file path that you supply, for example process_pdf("document.pdf") in Python or a Uint8Array in the browser WebAssembly build. Fetching the file is left to your own code.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/firecrawl-pdf-inspector)