OpenDataLoader PDF: a Java PDF parser for RAG pipelines and auto-tagging
OpenDataLoader PDF parses PDFs into Markdown, JSON with bounding boxes, and HTML for AI and RAG pipelines, with OCR for 80+ languages in its hybrid AI mode.
At a glance
- What is it?
- OpenDataLoader PDF extracts Markdown, JSON with bounding boxes and HTML from PDFs, and auto-tags untagged files into Tagged PDFs. The Apache-2.0 core is real; PDF/UA export is not.
- Who is it for?
- Adopt it if you need deterministic, bounding-box-annotated Markdown or JSON for a retrieval pipeline, or if you want to try auto-tagging untagged PDFs without buying a remediation tool. Do not adopt it expecting free PDF/UA-1 or PDF/UA-2 output, Word or Excel parsing, or a pure-Python install with no JVM.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Java, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
DEEP OPEN-SOURCE ANALYSIS
The two jobs OpenDataLoader PDF actually does
Most PDF libraries pick one job. OpenDataLoader PDF picks two, and they share a layout-analysis core. The first job is feeding retrieval pipelines: the README describes output as Markdown, JSON with bounding boxes, and HTML, with a Python, Node.js or Java SDK on top. The second job is accessibility remediation. The README states that accessibility regulations such as the EAA, ADA and Section 508 demand Tagged PDFs, that manual remediation costs 50 to 200 dollars per document, and that auto-tagging untagged PDFs into Tagged PDFs is the free Apache-2.0 part. The project calls itself the first open-source tool to generate Tagged PDFs end to end.
The audience follows from that split. If you are chunking annual reports for a RAG index, you care about reading order and table cells. If you are a public-sector or enterprise team sitting on a backlog of untagged PDFs, you care about how many documents a script can push through without a human in the loop. The capability matrix puts both under one roof, but the tier column is where the honesty lives: extraction and auto-tagging are free, PDF/UA-1 and PDF/UA-2 export are marked enterprise.
XY-Cut++ reading order, bounding boxes and the hybrid AI fallback
The mechanism the README names is XY-Cut++ for reading order, paired with bounding boxes attached to every detected element. That combination is what makes the JSON output usable for source citations: you can point a generated answer back at a rectangle on a page instead of at a whole document. Heading hierarchy detection, list detection including nested lists, and image extraction with coordinates are all listed as free capabilities.
The interesting architectural decision is the split between local and hybrid mode. Local mode is deterministic and, per the README, runs at 0.015s per page. Hybrid mode routes what the README calls complex pages to an AI backend and is where OCR for scanned PDFs, borderless table extraction, LaTeX formula extraction and AI-generated chart descriptions live. The project reports 0.907 overall accuracy in hybrid mode and 0.928 table accuracy across 200 real-world PDFs. Those numbers come from the project's own benchmark page, so treat them as a claim to reproduce on your own corpus rather than a settled fact.
Two smaller choices are worth noting. The README lists prompt-injection filtering as a free capability, which matters because hybrid mode sends page content to a model. Header, footer and watermark filtering is also free, and anyone who has chunked a 300-page report knows why that matters more than it sounds.
Installing OpenDataLoader PDF from PyPI and converting a first batch
The README gives Java 11+ and Python 3.10+ as the requirements, and tells you to run java -version before anything else, installing a JDK from Adoptium if it is missing. That check is not decoration: the Python package spawns a JVM, so a machine without Java produces a failure rather than a Python traceback you can act on.
Install from PyPI:
pip install -U opendataloader-pdfThen convert. The README's own example passes a list of files and a folder in one call, and explicitly warns that each convert() spawns a JVM process, so repeated calls are slow:
import opendataloader_pdf
opendataloader_pdf.convert(
input_path=["file1.pdf", "file2.pdf", "folder/"],
output_dir="output/",
format="markdown,json"
)What you should see is an output/ directory containing Markdown and JSON for each input, with the JSON carrying bounding boxes and semantic element types. If you only need Node.js or Java, the README points to separate quick-start pages for those rather than repeating the steps here.
The repository layout explains the packaging. It is a monorepo with java/, node/ and python/ directories, and the root package.json is a private pnpm workspace pinned to [email protected] with node >=22.13. The options.json and schema.json files at the root are generated: the sync script exports options from the Java CLI jar and regenerates the schema, which is why the configuration surface is documented in one place across three language bindings.
Auto-tagging is free, PDF/UA export is not
This is the boundary that decides whether the project fits you. The README is explicit that auto-tagging untagged PDFs into Tagged PDFs is free under Apache 2.0, and that converting a Tagged PDF to PDF/UA-1 or PDF/UA-2 is an enterprise add-on, with a visual accessibility studio also marked enterprise. A Tagged PDF is the foundation of a PDF/UA workflow, not the finished article.
So if your compliance requirement is a veraPDF pass against PDF/UA, the free tier gets you partway and the last step is a commercial conversation. The README does describe collaboration with Dual Lab, the veraPDF developers, and states that auto-tagging follows the PDF Association's Well-Tagged PDF specification and is validated with veraPDF. That is a meaningful signal about how the tagging is designed, but it is not the same as a compliance certificate for your documents.
A second limitation is scope. The capability matrix answers No to processing Word, Excel or PowerPoint files. If your ingestion pipeline is mostly .docx with a few PDFs mixed in, this is the wrong tool for the majority of your inputs. It also answers No to requiring a GPU, which is a point in its favour for on-premise deployment.
How it differs from Docling and MarkItDown
People searching for this project usually arrive comparing it to Docling or MarkItDown, and the differences are structural rather than cosmetic. MarkItDown is a lightweight converter: it targets Markdown for LLM consumption and does not attempt accessibility output. If all you need is text out of a PDF for a prompt, a converter of that shape is less machinery to operate.
Docling is a Python document-conversion framework with its own model stack. The practical difference with OpenDataLoader PDF is where the intelligence sits. OpenDataLoader PDF runs a deterministic local mode by default and only routes complex pages to an AI backend in hybrid mode, which the README frames as a deliberate split between speed and accuracy. It also emits bounding boxes for every element as a first-class output, and it ships a Java core with Python and Node.js bindings rather than being Python-native. That Java core is the reason the install requires a JDK, and it is also why the same engine is reachable from a Java service without a Python sidecar.
The accessibility track has no real open-source competitor in this comparison. Docling and MarkItDown produce extracted content; OpenDataLoader PDF additionally produces Tagged PDF. If that second output is what you need, the comparison is mostly over before it starts.
Licence, maintenance and upgrade cost
The repository is licensed Apache-2.0, which permits commercial use, modification and redistribution provided you keep the licence and notice files. The repository root contains both LICENSE and NOTICE, plus a THIRD_PARTY/ directory, and Apache-2.0 obligations around attribution and third-party notices are the practical thing to check before shipping it inside a product. None of this is legal advice; if you are redistributing a modified build, have counsel read the NOTICE and THIRD_PARTY contents.
The commercial boundary is the part licence text will not tell you: the README marks PDF/UA-1 and PDF/UA-2 export and the accessibility studio as enterprise add-ons, so an Apache-2.0 grant on the core does not imply a free path to a PDF/UA deliverable.
On maintenance, the last push to main was on 2026-08-25, and the releases list shows v2.5.5 on the same date, with v2.5.3 and v2.5.2 landing in the preceding days. That is a tight release cadence around a single date rather than a long tail of steady commits, and the README does not document a support window or a deprecation policy for the option schema. The upgrade cost you should plan for is the generated options.json and schema.json: because they are produced from the Java CLI, a version bump can change the configuration surface that your Python or Node.js calls depend on. Pin your version and diff options.json between releases rather than tracking the latest tag in production.
Editorial conclusion
Adopt it if you need deterministic, bounding-box-annotated Markdown or JSON for a retrieval pipeline, or if you want to try auto-tagging untagged PDFs without buying a remediation tool. Do not adopt it expecting free PDF/UA-1 or PDF/UA-2 output, Word or Excel parsing, or a pure-Python install with no JVM. Before committing, verify that Java 11+ is present on every machine that will run convert(), check that your licence budget covers hybrid mode and any PDF/UA export, and run your own worst-case PDFs through both local and hybrid mode, because the 0.907 overall figure comes from the project's own 200-document benchmark and not from your corpus.
Frequently asked questions
How do I use OpenDataLoader PDF from Python?
Install the PyPI package, then call opendataloader_pdf.convert() with input_path, output_dir and format. The README recommends batching files and folders into a single call because each convert() spawns a JVM process.
Which LLM is the best for PDF parsing?
The repository does not answer this. It reports its own hybrid mode at 0.907 overall accuracy across 200 real-world PDFs and does not name or rank the AI backend it routes complex pages to.
What is the best open source, free PDF reader?
OpenDataLoader PDF is not a reader; the README describes it as a parser that outputs Markdown, JSON and HTML, and as an auto-tagging tool. It does not document a viewer interface, so it cannot be recommended as a replacement for a PDF reader application.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/opendataloader-project-opendataloader-pdf)
Community notes