# Docling parses video and email too, but not your metadata

> An MIT-licensed Python library that turns PDF, office documents, EPUB, video and email into one document representation, then exports Markdown, DocTags or lossless JSON. It runs locally, and the install is one pip command against a pyproject that declares a different package name.

**docling-project/docling** — Get your documents ready for gen AI. Convert a document (CLI) This generates a .md file in the current directory containing structured document content.

- Repository: https://github.com/docling-project/docling
- Website: https://docling-project.github.io/docling
- Stars: 68,180 · Forks: 4,955
- Language: Python
- License: MIT
- Published: 2026-08-08 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/docling-project-docling

## The README installs docling, the pyproject declares docling-slim

Begin with a discrepancy that will confuse anyone who reads the build files alongside the quickstart. The documented install is one command.

```bash
pip install docling
```

The pyproject.toml in the repository declares its name as docling-slim, describes itself as a modular version of the Docling package, and marks that dependency set as a minimal base of eight packages at roughly 50MB. So the file a contributor reads to understand the install describes a smaller artifact than the one the documentation asks for. The gap matters because a working layout model pipeline needs weights rather than just Python packages, which means the 50MB figure is a floor for the library and not an estimate of what a first conversion costs in time and disk. The slim package exists so an environment can assemble its own model set, which is a sound reason for the split and a poor basis for capacity planning.

## Title, authors and language are on the coming soon list

The feature list runs to a dozen bullets and the roadmap list runs to two, and the distance between them is the part to read carefully. Under coming soon there is metadata extraction covering title, authors, references and language, plus complex chemistry understanding for molecular structures. The first of those is a live constraint for the most common use of this tool. A retrieval pipeline over a folder of papers usually wants title and authors first, to label and cite a chunk, before it wants body text. Docling will hand you the body, the reading order and the table structure, and it will not hand you the citation metadata, so that has to come from somewhere else: the filename, a sidecar file, or a second tool. Anyone budgeting a metadata stage should assume it is not in this release. The chemistry gap is narrower, but it means a chemistry document parses as text and figures rather than as structure.

## Every export except the JSON is lossy

The export list is where the central design decision becomes visible. Docling writes out Markdown, HTML, WebVTT, DocLang, DocTags, and lossless JSON, and the word lossless is attached to exactly one of them. Everything else is a projection of the DoclingDocument representation, which is where the expensive work sits: page layout, reading order, table structure, code and formula handling, and image classification. The rule that follows is about when you convert. If a script writes Markdown to disk as its first step and your indexer then reads that file, the table structure and the reading order have already been thrown away and no later stage can recover them. Keep the JSON and export at the point of consumption. Of the projections, DocTags is the one aimed at feeding a model directly, while DocLang and the schema targets sit closer to structured publishing than to a chat prompt.

```bash
docling https://arxiv.org/pdf/2206.01062
```

## Video, audio and email turned a PDF parser into a general converter

The recent additions change what the tool is for. The new list covers video in MP4, AVI, MOV, MKV and WebM, parsed with an ASR transcript and representative keyframes; OpenDocument files for text, spreadsheets and presentations; XBRL for financial reports; email in .eml and .msg; EPUB; Apple Pages and Keynote; plain text and Markdown supersets; and chart understanding that converts bar, pie and line plots into tables or code with detailed descriptions. A separate axis runs alongside it, the application schemas: DocLang, USPTO patents, JATS articles and XBRL reports. The price of that breadth is model surface area, because one invocation can now reach for OCR, speech recognition and a vision language model, and the feature list already named GraniteDocling alongside several others. Keep two lists apart when you plan capacity: the file formats, which are cheap to add, and the models, which are what a container image has to carry.

## The container installs CPU-only torch, then downloads the weights

The Dockerfile shows what a reproducible install really costs. It starts from python:3.11-slim-bookworm, installs a short system package list including libgl1 and libglib2.0-0, then runs pip install with an extra index URL pointing at the PyTorch CPU wheel index. The comment in the file states that this installs torch with only CPU support, and that removing the extra index URL is how you pick up the GPU requirements. Weight download is a separate step, docling-tools models download, which bakes the models into the image rather than fetching them at runtime. Two settings after that matter in a deployment. HF_HOME and TORCH_HOME both point at /tmp/, so the model cache lives in the container filesystem instead of a mounted volume and disappears with the container. OMP_NUM_THREADS is set to 4, with a comment about avoiding thread congestion in container environments. The default image is therefore a CPU converter, and a vision language pipeline will run inside it without a GPU.

## The shipped image turns off SSH host key checking

One line in that Dockerfile deserves a decision rather than a copy. The image sets GIT_SSH_COMMAND to ssh with StrictHostKeyChecking turned off. That is a common convenience in a build file, and it explains why a dependency fetched over SSH does not stop on an unknown host key. It is equally the reason the image will not notice if it is later pointed at a host whose key has changed. For a local conversion container the setting costs nothing, because the image is not fetching anything once it is built. For a shared or long-lived image that later runs git over SSH, the setting travels with it, and the fix belongs in your deployment rather than in your copy of this file. It is the clearest case here of a build convenience that quietly becomes an operational decision, and the only line in the Dockerfile with that reach.

## Contributing means uv, prek hooks and a cap on file length

The developer workflow is unusually explicit for a project this size, and the Makefile is where it lives. The setup target runs uv sync frozen against the dev group with all extras, excluding the docs and examples groups, so the environment is pinned rather than resolved on every clone. Git hooks are installed through prek. The read-only check target then runs a chain: ruff format in check mode, ruff check, the ty type checker, tach for module boundaries, a script that verifies tach module coverage, a separate script named check_max_lines, dprint for the non-Python files, and finally uv lock in locked mode to prove the lockfile is current. That last check and the max lines script describe the project's priorities better than any prose would. Module boundaries are enforced in CI rather than described in a document, and file length is capped, which means a large new feature is expected to arrive as several small modules instead of one long file.

## Conclusion

Adopt Docling when your input is messy documents and your output is a model, since the unified representation plus lossless JSON is the part worth paying for, and run it locally if the documents are sensitive. Do not adopt it as a metadata extractor: title, authors, references and language sit on the coming soon list, so that stage needs a second tool. Two things to verify before you commit. First, pip install docling does not install the slim package described in the repository pyproject, so size your disk for weights rather than for the fifty megabyte base. Second, the shipped container installs CPU-only torch, which makes any vision language pipeline a CPU job unless you change the extra index URL.

## FAQ

### What is docling used for?

Docling parses documents into one unified representation so downstream tools can use them. It handles PDF, DOCX, PPTX, XLSX, HTML, EPUB, Apple Pages, WAV, MP3, WebVTT, email, images, LaTeX and plain text, with advanced PDF work on layout, reading order, table structure, code, formulas and image classification.

### Does Docling run locally?

Yes. Local execution for sensitive data and air-gapped environments is a listed feature, and the quickstart install is a single pip command. Python 3.9 support was dropped in version 2.70.0, so you need Python 3.10 or higher.

### How do you use docling in Python?

Import DocumentConverter, give it a local path or a URL, and call convert. The result exposes a document object with an export_to_markdown method, and the same object can be exported to HTML, WebVTT, DocLang, DocTags or lossless JSON.

### How do you install docling in Docker?

The repository ships a Dockerfile built on python:3.11-slim-bookworm. It installs the package with an extra index URL for the PyTorch CPU wheels, runs docling-tools models download to bake in the weights, and sets OMP_NUM_THREADS to 4. Removing that extra index URL is how you get the GPU requirements.

## Sources

- [Official documentation](https://docling-project.github.io/docling)
- [Official README](https://github.com/docling-project/docling#readme)
- [Project repository](https://github.com/docling-project/docling)
- [Release notes](https://github.com/docling-project/docling/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/docling-project-docling
