open-parse: layout-aware PDF chunking for LLM pipelines
Improved file parsing for LLM’s
At a glance
- What is it?
- open-parse is a Python library that groups PDF content the way a human reads it, so retrieval pipelines get whole sections and clean tables instead of sliced text. It is MIT licensed, but its best table extraction path pulls in ML weights and PyMuPDF's separate licence.
- Who is it for?
- Adopt open-parse if you are building a RAG pipeline over PDFs with real structure: headings, bullets, multi-column pages, tables, and you want chunk boundaries that follow the document rather than a token counter. Skip it if your corpus is plain text or HTML, if you cannot add PyMuPDF's licence terms on top of MIT, or if you need an actively developed dependency (the last push was on 2026-05-17 and the latest release is v0.7.0 from 2024-11-13).
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 135 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The chunking problem open-parse is aimed at
Most retrieval pipelines start by converting a PDF to a flat string and then slicing that string every N tokens. open-parse's README calls this out directly: text splitting means you lose the ability to overlay a chunk back onto the original page, you ignore headings, sections and bullets, and you get no support for tables, images or markdown. The library is built for the people who hit that wall: engineers assembling RAG systems over contracts, manuals, financial filings or research papers, where a chunk that starts mid-sentence under a heading it no longer carries is close to useless as a retrieval unit.
The target user is narrow on purpose. If your corpus is already clean markdown or HTML, open-parse adds a PDF stack you do not need. If your documents are scanned images with no text layer, the library can route through OCR, but the README's OCR checklist shows that path depends on a separate Tesseract installation and a TESSDATA_PREFIX environment variable, so the setup cost is real.
How open-parse groups nodes instead of slicing text
The core object is DocumentParser. Calling parser.parse(path) returns a result whose .nodes attribute is an iterable of content nodes, and the README's basic example simply prints each node. Nodes are the unit of output, not pages and not fixed-size text windows. That is the architectural difference from a token splitter: the parser first identifies layout elements on the page, then decides which of them belong together, and only then hands you the grouping.
On top of that base, open-parse ships a semantic pipeline. The README describes the idea plainly: chunking is about grouping similar semantic nodes, so each node's text is embedded and nodes are clustered by similarity. The SemanticIngestionPipeline takes an OpenAI key, a model name, and min_tokens and max_tokens bounds. That means the default grouping can be replaced by an embedding-driven one, at the cost of an API call per node and a dependency on the openai package, which is already in the core dependency list.
Results are pydantic models, so parsed_content.dict() and parsed_content.json() are available for serialization. Tables are a separate concern: the ML path uses deep learning models, and the README points at PyMuPDF's built-in table detection, Microsoft's Table Transformer, and unitable, describing unitable as a transformers-based approach with state-of-the-art performance. Those are the project's claims, not independent measurements.
Installing open-parse and parsing your first PDF
The core install is a single pip command, and the README lists Python 3.8+ as the requirement. The core dependency set includes PyMuPDF, pypdf, pdfminer.six, pydantic, tiktoken, openai, pillow and numpy.
pip install openparseAfter that, the README's basic example is the shortest real use. Point basic_doc_path at your own PDF and iterate the nodes; each printed node is one grouped content unit rather than a raw page dump.
import openparse
basic_doc_path = "./sample-docs/mobile-home-manual.pdf"
parser = openparse.DocumentParser()
parsed_basic_doc = parser.parse(basic_doc_path)
for node in parsed_basic_doc.nodes:
print(node)If you want table extraction through the deep learning route, the README specifies an extra install and a weights download step. The optional extra is named ml, and it pulls torch, torchvision, transformers and tokenizers. The download command is a console script registered in pyproject.toml as openparse-download.
pip install "openparse[ml]"
openparse-downloadTable parsing is then selected through the table_args argument on DocumentParser, with parsing_algorithm set to "unitable" in the README's example. The README does not document the full set of accepted parsing_algorithm values in the excerpt available here, so treat unitable as the confirmed one and check the docs site for the rest.
Where open-parse is the wrong tool
The heaviest limitation is the table stack. The ML extra brings torch, torchvision, transformers and tokenizers into your environment, which is a large footprint for a library whose core install is otherwise modest. The README also flags the licence position rather than hiding it: PyMuPDF, which the core dependency list requires, has its own licensing page linked under Requirements. open-parse itself is MIT, but that does not relicense PyMuPDF, and the README explicitly points readers to mupdf.com's licensing page. If your legal review cannot accept those terms, the core dependency is a blocker, not a configuration detail.
OCR is the second sharp edge. The README states that PyMuPDF contains the OCR logic but still needs Tesseract's language data, and that the language folder must be communicated through TESSDATA_PREFIX or as a function parameter. It goes further: on Windows, this must happen outside Python, before starting the script, and manipulating os.environ will not work. That is an unusual constraint and a common source of silent failures.
The third is maintenance. The last push to the repository was on 2026-05-17, but the most recent release is v0.7.0 from 2024-11-13, with v0.6.1 and v0.6.0 before it in 2024. There is a gap between commit activity and tagged releases. If your team pins versions and expects frequent releases, that gap matters more than the commit date.
How open-parse compares with ML layout parsers and commercial APIs
The README draws two comparisons itself. Against ML layout parsers such as layout-parser, the argument is about scope: those tools identify text blocks, images and tables, but the README says they are not built to group related content effectively, that they focus strictly on layout parsing, and that you would need another model to extract markdown from images, parse tables and group nodes. open-parse's answer is to ship the grouping step, the semantic pipeline and the table algorithms in one library, with a processing_pipeline hook for your own post-processing. The trade-off is that you inherit open-parse's grouping heuristics instead of composing your own stack, and the README's performance remarks about other parsers are the maintainer's assessment, not a published benchmark.
Against commercial document APIs, the README gives a price anchor of roughly $10 per 1,000 pages and names Google Document AI, AWS Textract and Reducto as examples. The second point is data residency: commercial services require sending documents to a vendor. open-parse runs locally, which is the reason many teams pick it. That advantage disappears if you enable the semantic pipeline, because SemanticIngestionPipeline calls OpenAI embeddings with your API key, so node text leaves your machine anyway. If local-only processing is the requirement, the default parser plus the ML table path keeps everything on-premises; the semantic pipeline does not.
Licence, upgrade cost and what to check before pinning
open-parse is MIT licensed, per pyproject.toml and the LICENSE file at the repository root. That covers the library's own code. It does not cover PyMuPDF, which the core dependencies require at version 1.23.2 or above and which the README links to a separate commercial licensing page. pdfminer.six is described in the README as fully open source. The ML extra adds torch and transformers, whose licences are their own. This is a description of what the repository states, not legal advice; anyone shipping a product should have the PyMuPDF terms reviewed.
Upgrade cost is mostly dependency-driven. The core list pins minimum versions rather than upper bounds, so a fresh install today can resolve newer PyMuPDF, pydantic or openai releases than the ones the 0.7.0 code was written against. pydantic 2.0 or above is required, which means code written for pydantic v1 serialization patterns needs adjusting. The ml extra is the expensive part to re-resolve, since torch and torchvision carry platform-specific wheels. Model weights are fetched by openparse-download rather than bundled, so CI images need that step cached or repeated.
Editorial conclusion
Adopt open-parse if you are building a RAG pipeline over PDFs with real structure: headings, bullets, multi-column pages, tables, and you want chunk boundaries that follow the document rather than a token counter. Skip it if your corpus is plain text or HTML, if you cannot add PyMuPDF's licence terms on top of MIT, or if you need an actively developed dependency (the last push was on 2026-05-17 and the latest release is v0.7.0 from 2024-11-13). Before committing, verify three things on your own documents: whether the default parser's node grouping matches your section boundaries, whether your OCR path works with TESSDATA_PREFIX set outside Python on Windows, and whether the unitable table weights download successfully via openparse-download on your target machine.
Frequently asked questions
What does open-parse do?
It parses PDFs and groups their content into nodes that follow the document's layout, so an LLM pipeline gets whole sections, headings and tables instead of text sliced at a fixed token count. The README describes it as visually analyzing documents for LLM input.
Is it possible to parse a PDF file with open-parse?
Yes. The README's basic example calls openparse.DocumentParser().parse() on a .pdf path and iterates the resulting nodes. PDF handling in the core install relies on pdfminer.six and PyMuPDF.
How do you install open-parse?
The README gives pip install openparse for the core library, which requires Python 3.8 or higher. Table parsing through deep learning models needs the extra install pip install "openparse[ml]" followed by the openparse-download command to fetch model weights.
Can you install open-parse on an iPad or iPhone?
The README documents installation through pip with Python 3.8 or higher, which is a desktop and server workflow; it does not describe an iOS or iPad installation path.
What does it mean to parse a file?
In open-parse's case it means reading a document and returning structured content nodes rather than raw text. Calling parser.parse(path) yields a result whose .nodes attribute you can iterate, and the result serializes through pydantic with dict() or json().
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/filimoa-open-parse)