docling vs unstructured: local document parsing against a broad ingestion toolkit
Docling converts documents into one typed representation on your own hardware; unstructured partitions many formats into labelled elements. They overlap on PDFs, but the first optimises for layout fidelity and local execution, the second for breadth of formats and connectors. Many teams can run one and keep the other for the formats it covers better.
At a glance
| Project | docling-project/docling | Unstructured-IO/unstructured |
|---|---|---|
| Licence | MITPermissive: commercial use allowed | Apache-2.0Permissive: commercial use allowed |
| Maintenance | Commits in the last dayLast push September 29, 2026 | Commits in the last six monthsLast push September 15, 2026 |
| Language | Python | HTML |
| GitHub stars | 68,180 | 15,442 |
| Read more | Our analysisGitHub | Our analysisGitHub |
Which one to choose
Choose docling if you process scanned or layout-heavy PDFs, need reading order and table structure preserved, must keep documents on your own machines or in an air-gapped network, and want a single DoclingDocument that exports to Markdown, HTML, DocTags or lossless JSON before feeding LangChain, LlamaIndex, Haystack or an MCP agent.
Choose unstructured if your corpus spans many file types and you mainly need clean, labelled elements (titles, narrative text, tables, list items) for chunking and embedding, and you accept a heavier dependency set and a library that is updated frequently.
Two different centres of gravity: one representation against labelled elements
Docling, from IBM Research Zurich, parses PDF, DOCX, PPTX, XLSX, HTML, EPUB, email, audio and video into a single document representation, the DoclingDocument, and then exports Markdown, HTML, DocTags or lossless JSON. The representation is the product. Layout models decide page layout, reading order, table structure, code blocks and formulas, and everything downstream reads that one object. The README also lists application-specific XML schemas (DocLang, USPTO patents, JATS articles, XBRL financial reports) and a VLM pipeline, for example docling --pipeline vlm --vlm-model granite_docling, which swaps the default pipeline for a visual language model. The practical consequence is that a table survives conversion as a table, not as a run of text with pipes in it. Our analysis of docling says the main reason to pick it over a hosted parser is that it runs locally, and the main reason to check your hardware first is the same: the layout models need memory.
Unstructured starts from the other end. Its library ingests PDFs, HTML, Word documents, images and many other types, and its partition functions cut each file into typed elements that a language model pipeline can chunk and embed. The README frames the project as open-source pre-processing tools for unstructured data, with modular functions and connectors forming a system for ingestion. There is no claim of one canonical document object; the output is a stream of elements with types and metadata. The README also points to an Unstructured Transform MCP server that turns 60+ file types into agent-consumable output, and to an enterprise Platform product for production workflows, partitioning, enrichment, chunking and embedding.
The split matters when you choose. If your pipeline needs to reason about page geometry, Docling gives you the geometry. If your pipeline needs a predictable element stream across twenty formats, unstructured gives you the stream and leaves geometry to the partitioner.
Getting each one running: pip install against dependency weight
Docling installs with pip install docling, and the README notes that Python 3.9 support was dropped in version 2.70.0, so Python 3.10 or higher is required. It runs on macOS, Linux and Windows for x86_64 and arm64. The shortest path is the CLI: docling https://arxiv.org/pdf/2206.01062 writes a Markdown file into the current directory. The Python path is three lines: construct a DocumentConverter, call convert on a local path or URL, then call export_to_markdown on the resulting document. Model weights are downloaded on first use, which is the first thing to plan for in a locked-down environment.
Unstructured is also installed from PyPI, but its scope is wider: the README describes components for ingesting and pre-processing images and text documents, and the partition functions for different formats pull in different dependencies. That breadth is the point and also the cost. A team that only needs PDFs still carries the installation surface of a multi-format toolkit unless it installs selectively. The README points to documentation for partitioning, and the project's release cadence (0.27.5 on 2026-08-28, 0.27.1 and 0.27.0 on 2026-08-21) means pinning a version and reading the changelog before upgrading is normal practice rather than caution.
Neither install is hard. The difference is what arrives with it: Docling brings layout and OCR models, unstructured brings format handlers and their transitive dependencies. Our analysis of unstructured says to adopt it only if you can manage the dependency weight and maintenance burden.
Operations: local inference, model weights and where the service boundary sits
Docling runs locally, which the README states as a capability for sensitive data and air-gapped environments. That is an operational commitment as much as a feature. You size machines for the layout models, you cache model weights, and you decide whether to redistribute them: our analysis of docling says to verify which model weights your environment can legally redistribute if you ship a container. For batch work, the API server (docling-serve) turns the library into a service, and the MCP server connects an agent to the same conversion path. Scaling is therefore a question of how many conversions per minute your hardware sustains with the pipeline you chose, and whether the VLM pipeline is worth its cost on your documents.
Unstructured's open-source library is a component you host yourself; the README directs readers to the enterprise Platform product for production-grade workflows. That is the operational fork. If you need guaranteed uptime and support, our analysis of unstructured says not to adopt the open-source library for that purpose, because the Platform exists for it. If you self-host the library, you own dependency upgrades, partitioner behaviour changes and the throughput of the partition functions on your files. The release history shows frequent minor releases, so an upgrade policy is part of operations, not an afterthought.
One asymmetry is worth stating plainly. Docling's local execution is the stated reason to choose it; unstructured's local execution is simply how the open-source library works, with the managed path sold separately.
Where each one falls short
Docling is the weaker choice when the job is plain text extraction from clean, text-layer PDFs. Our analysis of docling says a smaller extractor costs less to install and run in that case, and that you should skip Docling if you cannot give the process enough memory for the layout models. The format list is broad, but the depth is concentrated in PDF understanding: page layout, reading order, table structure, code, formulas, image classification. Teams that need many formats handled with equal depth will find the PDF path the most developed one. The README also lists metadata extraction, including title, authors, references and language, under Coming soon, so that is not something to depend on yet.
Unstructured is the weaker choice when output fidelity on complex tables or reading order is the requirement, because its contract is labelled elements rather than a typed document graph. It is also the weaker choice when you want a managed service with uptime guarantees: our analysis of unstructured says the enterprise Platform product exists for that, and the open-source library is not a substitute. Dependency weight is the third limit. A large dependency tree is not a defect by itself, but it changes upgrade risk, image size and the number of places a build can break.
Both projects leave a gap for teams that want a fully managed parser with no self-hosting at all. Docling's answer is docling-serve on your own infrastructure; unstructured's answer is the Platform. Neither open-source repository is that managed service.
Licence and maintenance: MIT against Apache-2.0, and what the push dates say
Docling is MIT licensed. Unstructured is Apache-2.0. Both are permissive and both permit commercial use; the difference that matters for distribution is the patent grant language in Apache-2.0 and the notice and attribution expectations that come with it. Teams shipping containers should read the licence text rather than assume the two are interchangeable, and in Docling's case should also check the terms attached to model weights, which are separate from the code licence.
On maintenance, the facts are simple. Docling's last push was 2026-09-25 and its most recent release was v2.123.1 on 2026-08-28. Unstructured's last push was 2026-09-15 and its most recent release was 0.27.5 on 2026-08-28. Neither repository is archived. Both are being pushed to and released from within the last two months, so neither carries a staleness warning. What differs is release rhythm: Docling's recent releases land days apart (v2.123.1 on 2026-08-28, v2.123.0 on 2026-08-26, v2.122.0 on 2026-08-25), and unstructured's recent releases are also close together (0.27.5 on 2026-08-28, 0.27.1 and 0.27.0 on 2026-08-21). Frequent releases are good for fixes and bad for pinning. Plan to pin either one and to test upgrades against your own corpus.
Licence choice should follow distribution model, not popularity. If you embed the parser in a product you ship, Apache-2.0's explicit patent grant may be easier to explain to legal than MIT's silence on patents, and MIT's brevity may be easier to satisfy in an internal tool.
Choosing for concrete scenarios
A team digitising scanned contracts with tables and multi-column layout should start with docling, because layout, reading order and table structure are the stated strengths, and local execution keeps the documents inside the network. Verify the default pipeline on your worst scan before committing, as our analysis of docling advises, and check where the CLI output lands relative to your pipeline's expectations.
A team building a RAG ingestion service over a mixed corpus of PDFs, Word files, HTML pages, images and email should start with unstructured, because the partition functions cover that breadth and produce elements that chunk cleanly for embeddings. Verify that your specific file types are handled well by the partition functions, test output quality on your own documents, and confirm that Apache-2.0 fits your distribution model, which is what our analysis of unstructured recommends.
A team that needs both can run them together without much friction. Use unstructured to partition the long tail of formats into elements, and docling where layout fidelity on PDFs decides whether the extracted table is usable. The two do not share a document model, so the combination means two output shapes and two upgrade cycles; that cost is real, but it is smaller than forcing one tool into a job it does not do well.
A team that only needs text from clean, text-layer PDFs should pick neither by default. Our analysis of docling says a smaller extractor costs less to install and run, and the same logic applies to unstructured for that narrow case. Reach for these two when structure, not just text, is the deliverable.
Bottom line
Docling is the better default when documents stay on your own hardware and PDF layout decides whether the output is usable; unstructured is the better default when format breadth and a labelled element stream matter more than geometry. Before committing, run your worst scan through docling's default pipeline and check the model weight terms, and run your specific file types through unstructured's partition functions while confirming Apache-2.0 fits how you ship.