Docling parses video and email too, but not your metadata
Get your documents ready for gen AI. Convert a document (CLI) This generates a .md file in the current directory containing structured document content.
At a glance
- What is it?
- An MIT-licensed Python library that turns PDF, office documents, EPUB, video and email into one document representation, then exports Markdown, DocTags or lossless JSON. It runs locally, and the install is one pip command against a pyproject that declares a different package name.
- Who is it for?
- Adopt Docling when your input is messy documents and your output is a model, since the unified representation plus lossless JSON is the part worth paying for, and run it locally if the documents are sensitive. Do not adopt it as a metadata extractor: title, authors, references and language sit on the coming soon list, so that stage needs a second tool.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The README installs docling, the pyproject declares docling-slim
Begin with a discrepancy that will confuse anyone who reads the build files alongside the quickstart. The documented install is one command.
pip install doclingThe pyproject.toml in the repository declares its name as docling-slim, describes itself as a modular version of the Docling package, and marks that dependency set as a minimal base of eight packages at roughly 50MB. So the file a contributor reads to understand the install describes a smaller artifact than the one the documentation asks for. The gap matters because a working layout model pipeline needs weights rather than just Python packages, which means the 50MB figure is a floor for the library and not an estimate of what a first conversion costs in time and disk. The slim package exists so an environment can assemble its own model set, which is a sound reason for the split and a poor basis for capacity planning.
Title, authors and language are on the coming soon list
The feature list runs to a dozen bullets and the roadmap list runs to two, and the distance between them is the part to read carefully. Under coming soon there is metadata extraction covering title, authors, references and language, plus complex chemistry understanding for molecular structures. The first of those is a live constraint for the most common use of this tool. A retrieval pipeline over a folder of papers usually wants title and authors first, to label and cite a chunk, before it wants body text. Docling will hand you the body, the reading order and the table structure, and it will not hand you the citation metadata, so that has to come from somewhere else: the filename, a sidecar file, or a second tool. Anyone budgeting a metadata stage should assume it is not in this release. The chemistry gap is narrower, but it means a chemistry document parses as text and figures rather than as structure.
Every export except the JSON is lossy
The export list is where the central design decision becomes visible. Docling writes out Markdown, HTML, WebVTT, DocLang, DocTags, and lossless JSON, and the word lossless is attached to exactly one of them. Everything else is a projection of the DoclingDocument representation, which is where the expensive work sits: page layout, reading order, table structure, code and formula handling, and image classification. The rule that follows is about when you convert. If a script writes Markdown to disk as its first step and your indexer then reads that file, the table structure and the reading order have already been thrown away and no later stage can recover them. Keep the JSON and export at the point of consumption. Of the projections, DocTags is the one aimed at feeding a model directly, while DocLang and the schema targets sit closer to structured publishing than to a chat prompt.
docling https://arxiv.org/pdf/2206.01062Video, audio and email turned a PDF parser into a general converter
The recent additions change what the tool is for. The new list covers video in MP4, AVI, MOV, MKV and WebM, parsed with an ASR transcript and representative keyframes; OpenDocument files for text, spreadsheets and presentations; XBRL for financial reports; email in .eml and .msg; EPUB; Apple Pages and Keynote; plain text and Markdown supersets; and chart understanding that converts bar, pie and line plots into tables or code with detailed descriptions. A separate axis runs alongside it, the application schemas: DocLang, USPTO patents, JATS articles and XBRL reports. The price of that breadth is model surface area, because one invocation can now reach for OCR, speech recognition and a vision language model, and the feature list already named GraniteDocling alongside several others. Keep two lists apart when you plan capacity: the file formats, which are cheap to add, and the models, which are what a container image has to carry.
The container installs CPU-only torch, then downloads the weights
The Dockerfile shows what a reproducible install really costs. It starts from python:3.11-slim-bookworm, installs a short system package list including libgl1 and libglib2.0-0, then runs pip install with an extra index URL pointing at the PyTorch CPU wheel index. The comment in the file states that this installs torch with only CPU support, and that removing the extra index URL is how you pick up the GPU requirements. Weight download is a separate step, docling-tools models download, which bakes the models into the image rather than fetching them at runtime. Two settings after that matter in a deployment. HF_HOME and TORCH_HOME both point at /tmp/, so the model cache lives in the container filesystem instead of a mounted volume and disappears with the container. OMP_NUM_THREADS is set to 4, with a comment about avoiding thread congestion in container environments. The default image is therefore a CPU converter, and a vision language pipeline will run inside it without a GPU.
The shipped image turns off SSH host key checking
One line in that Dockerfile deserves a decision rather than a copy. The image sets GIT_SSH_COMMAND to ssh with StrictHostKeyChecking turned off. That is a common convenience in a build file, and it explains why a dependency fetched over SSH does not stop on an unknown host key. It is equally the reason the image will not notice if it is later pointed at a host whose key has changed. For a local conversion container the setting costs nothing, because the image is not fetching anything once it is built. For a shared or long-lived image that later runs git over SSH, the setting travels with it, and the fix belongs in your deployment rather than in your copy of this file. It is the clearest case here of a build convenience that quietly becomes an operational decision, and the only line in the Dockerfile with that reach.
Contributing means uv, prek hooks and a cap on file length
The developer workflow is unusually explicit for a project this size, and the Makefile is where it lives. The setup target runs uv sync frozen against the dev group with all extras, excluding the docs and examples groups, so the environment is pinned rather than resolved on every clone. Git hooks are installed through prek. The read-only check target then runs a chain: ruff format in check mode, ruff check, the ty type checker, tach for module boundaries, a script that verifies tach module coverage, a separate script named check_max_lines, dprint for the non-Python files, and finally uv lock in locked mode to prove the lockfile is current. That last check and the max lines script describe the project's priorities better than any prose would. Module boundaries are enforced in CI rather than described in a document, and file length is capped, which means a large new feature is expected to arrive as several small modules instead of one long file.
Editorial conclusion
Adopt Docling when your input is messy documents and your output is a model, since the unified representation plus lossless JSON is the part worth paying for, and run it locally if the documents are sensitive. Do not adopt it as a metadata extractor: title, authors, references and language sit on the coming soon list, so that stage needs a second tool. Two things to verify before you commit. First, pip install docling does not install the slim package described in the repository pyproject, so size your disk for weights rather than for the fifty megabyte base. Second, the shipped container installs CPU-only torch, which makes any vision language pipeline a CPU job unless you change the extra index URL.
Frequently asked questions
What is docling used for?
Docling parses documents into one unified representation so downstream tools can use them. It handles PDF, DOCX, PPTX, XLSX, HTML, EPUB, Apple Pages, WAV, MP3, WebVTT, email, images, LaTeX and plain text, with advanced PDF work on layout, reading order, table structure, code, formulas and image classification.
Does Docling run locally?
Yes. Local execution for sensitive data and air-gapped environments is a listed feature, and the quickstart install is a single pip command. Python 3.9 support was dropped in version 2.70.0, so you need Python 3.10 or higher.
How do you use docling in Python?
Import DocumentConverter, give it a local path or a URL, and call convert. The result exposes a document object with an export_to_markdown method, and the same object can be exported to HTML, WebVTT, DocLang, DocTags or lossless JSON.
How do you install docling in Docker?
The repository ships a Dockerfile built on python:3.11-slim-bookworm. It installs the package with an extra index URL for the PyTorch CPU wheels, runs docling-tools models download to bake in the weights, and sets OMP_NUM_THREADS to 4. Removing that extra index URL is how you get the GPU requirements.
Official sources
Where this project is recommended
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/docling-project-docling)