# docTR: A PyTorch OCR Library With a Two-Stage Detection and Recognition Pipeline

> docTR is an Apache-2.0 Python library that combines text detection and text recognition into one predictor. It is maintained by t2k GmbH, and its last push was on 2026-09-01.

**mindee/doctr** — docTR (Document Text Recognition) - a seamless, high-performing & accessible library for OCR-related tasks powered by Deep Learning. Ongoing development and maintenance by t2k.

- Repository: https://github.com/mindee/doctr
- Website: https://mindee.github.io/doctr/
- Stars: 6,369 · Forks: 683
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/mindee-doctr

## What docTR does that a plain OCR call does not

Most OCR wrappers hand you a string. docTR hands you a document model. The library splits optical character recognition into two stages: text detection, which localizes words on the page, and text recognition, which identifies the characters inside each localized box. The README states this explicitly: "End-to-End OCR is achieved in docTR using a two-stage approach: text detection (localizing words), then text recognition (identify all characters in the word)."

The output of that pipeline is a Document object with a nested structure of Page, Block, Line, Word and Artefact, and it can be flattened with result.export() into a nested dict that maps cleanly onto JSON. That structure is the reason to pick docTR over a library that only returns text: you keep the geometry, the reading order and the confidence values, so you can post-process by region rather than by string offset.

The audience is developers who are comfortable in Python and want OCR as a component inside a larger application. The classifiers in pyproject.toml list Intended Audience :: Developers, Education and Science/Research, which matches the shape of the API: there is no command line entry point described in the README, and the quick tour is written as Python snippets. If you want a binary you can point at a PDF and get a text file back, this is not that.

## The two-stage pipeline and the Document object it returns

The predictor is assembled from two architecture names. The README gives this example: ocr_predictor(det_arch="db_resnet50", reco_arch="crnn_vgg16_bn", pretrained=True). The detection architecture and the recognition architecture are chosen independently from the lists in the documentation, which means you can trade accuracy against speed by swapping one side without touching the other.

Input handling is separated from inference. DocumentFile.from_pdf reads a PDF, DocumentFile.from_images accepts a single path or a list of page images, and DocumentFile.from_url fetches a webpage, though the README notes that the URL path requires weasyprint to be installed. Multi-page image input is supported through the list form, so a scanned batch can be passed as one document.

Page orientation is a configuration decision rather than an automatic one. Passing assume_straight_pages=True fits straight boxes directly and the README calls this the fastest option. Passing export_as_straight_boxes=True keeps rotated pages but converts the final localizations to straight boxes. With both left at False, the predictor fits and returns rotated boxes, potentially with an angle of 0 degrees. That is a real fork in behaviour, and the README does not describe how the predictor decides orientation when rotation handling is enabled.

A layout detection stage can be added by passing detect_layout=True. The detected regions carry types such as Title, Text, Table, Page-header and Page-footer, and they are attached to every page. The README shows iterating result.pages[0].layout and printing region.type, region.confidence and region.geometry.

## Installing docTR and running a first page through it

The package name on PyPI is python-doctr, not doctr, and pyproject.toml requires Python 3.11.0 or newer and below 4. The Dockerfile installs the package from a git checkout with a framework extra, using the command shown in the repository file:

```dockerfile
RUN pip3 install -U pip setuptools wheel && \
    pip3 install "python-doctr[$FRAMEWORK]@git+https://github.com/$DOCTR_REPO.git@$DOCTR_VERSION"
```

The Dockerfile builds from nvidia/cuda:12.2.0-base-ubuntu22.04 and sets FRAMEWORK to torch by default, so a container route exists if you would rather not manage CUDA and the OpenCV system libraries yourself. For a first run, load a PDF and hand it to the default pretrained predictor. The README gives this example:

```python
from doctr.io import DocumentFile
from doctr.models import ocr_predictor

model = ocr_predictor(pretrained=True)
# PDF
doc = DocumentFile.from_pdf("path/to/your/doc.pdf")
# Analyze
result = model(doc)
```

After this, result is a Document. Calling result.export() gives the nested dict form. To look at the predictions rather than read them, result.show() renders an interactive view, and the README notes that this requires matplotlib and mplcursors. There is also result.synthesize(), which rebuilds page images from the predictions and returns a list of arrays that matplotlib can display. If your pages are known to be upright, adding assume_straight_pages=True to the predictor call is the cheaper configuration.

## The KIE predictor and where the README stops short

The KIE predictor is the more interesting extension. Instead of one generic text class, the detection model can be trained to detect multiple classes, and the README gives dates and addresses as examples. The kie_predictor wires a multi-class detector to a recognition model and sets up the pipeline for you, which saves the glue code you would otherwise write between two separate models.

The README's description of the KIE predictor is thin. It shows the import and the beginning of a constructor call, but the example is cut off before the arguments are complete, so the exact signature is not visible in the README itself. Anyone planning to use KIE should read the documentation rather than the README for the full parameter list.

The same caution applies to training. The repository contains a references/ directory and a Makefile with targets for tests and docs, but the README excerpt does not walk through fine-tuning a detector on a custom class set. The multi-class detector in the KIE example has to come from somewhere, and the README does not say where. That gap is the main thing to resolve before you design around KIE.

## Cost, orientation and the beta label are the real limits

Two models run per page. Detection runs first, then recognition runs on each detected box. That is the price of the structured output, and it is why the straight-page option is described as the fastest one: every page that needs rotated box fitting adds work before recognition even starts.

The dependency list is not small. pyproject.toml pulls in torch and torchvision, onnx, numpy, scipy, h5py, opencv-python, pypdfium2, pyclipper and shapely, among others, with upper bounds on most of them. In a constrained environment, opencv-python and the CUDA base image are the parts most likely to fight with what you already have.

The most concrete limitation is the status label. pyproject.toml sets the classifier to Development Status :: 4 - Beta even though the project has reached v1.1.0. The setup.py default build version is 1.1.1a0. A beta classifier on a 1.x release means you should expect API movement between minor versions, and you should pin the version you deploy rather than track main. The Dockerfile's default DOCTR_VERSION argument is main, which is the opposite of what a production build wants.

Finally, docTR is the wrong tool if your documents are dense tables where cell structure matters more than word text. Layout regions are detected and labelled, but the README does not describe reconstructing a table grid from those regions.

## docTR compared with Tesseract's approach

Tesseract is the obvious alternative and the comparison is not close in design. Tesseract is a C++ engine with a long history, shipped as a standalone binary with language data files, and it is typically driven from the command line or through a thin binding. Its recognition is built around a line-oriented model, and its training tooling is a separate stack.

docTR is a Python library whose models are neural networks loaded through PyTorch. There is no language data file to download separately; the weights arrive through the pretrained flag on the predictor. Detection and recognition are separate, swappable architectures, so you can change the recognizer without changing the detector. The output is a geometry-aware Document object rather than a text stream with bounding boxes bolted on.

The trade-off runs the other way too. Tesseract runs without a Python interpreter and without a GPU, which makes it easier to drop into a shell script or a small container. docTR's dependency set and its beta classifier are the costs you accept in exchange for the structured output and the ability to plug in a custom multi-class detector. If your requirement is extracting plain text from clean scans on a machine with no Python, Tesseract is the simpler answer.

## Licence and the cost of keeping up

docTR is Apache-2.0, and setup.py carries the header "This program is licensed under the Apache License 2.0." The pretrained weights are distributed through the same package, so the licence covers the code you install. This is a permissive licence that does not impose copyleft on your application, but it is not legal advice and the obligations that apply to your distribution are for your own counsel to review.

Upgrade cost is driven by the dependency bounds. torch is pinned to >=2.0.0,<3.0.0 and opencv-python to >=4.5.0,<6.0.0, so major-version jumps in either will require a docTR release before you can move. The project's own quality gates are visible in the Makefile: ruff check and mypy for static checks, and pytest with coverage report --fail-under=80, which tells you the maintainers hold a coverage floor.

The last push to the default branch was on 2026-09-01, and v1.1.0 was released on 2026-08-21. The repository is not archived. That is the maintenance picture you have; the README does not publish a support window or a deprecation policy for the model architectures, so pinning a version and reading the release notes before each bump is the practical approach.

## Conclusion

Adopt docTR if you need a Python-native OCR pipeline that returns a structured Document object and can be extended with a KIE detector for multi-class regions. Do not adopt it if you need a plain-text CLI with no Python runtime, or if you expect a stable API, because pyproject.toml still classifies the package as Development Status 4 - Beta. Before committing, verify that the detector and recognizer pair you pick is listed in the documentation, that your Python is 3.11 or newer, and that the GPU or CPU cost of running two models per page fits your throughput budget.

## FAQ

### What is docTR OCR?

docTR is a Python library for OCR-related tasks powered by deep learning, originally created by Mindee and now developed and maintained by t2k GmbH. It performs end-to-end OCR with a two-stage approach: text detection to localize words, then text recognition to identify the characters in each word.

### How do I install docTR?

The package name on PyPI is python-doctr, and the Dockerfile installs it from a git checkout with a framework extra, using pip3 install "python-doctr[$FRAMEWORK]@git+https://github.com/$DOCTR_REPO.git@$DOCTR_VERSION". pyproject.toml requires Python 3.11.0 or newer and below 4.

### Which docTR models can I choose for detection and recognition?

The predictor takes a det_arch and a reco_arch argument, and the README gives db_resnet50 for detection and crnn_vgg16_bn for recognition as an example. The documentation lists the available detection and recognition implementations, and the two sides can be selected independently.

### How does docTR handle rotated documents?

Passing assume_straight_pages=True fits straight boxes directly and is described as the fastest option. Passing export_as_straight_boxes=True returns straight boxes even when pages are rotated, and with both options False the predictor fits and returns rotated boxes, potentially with an angle of 0 degrees.

### How does docTR compare with Tesseract?

Tesseract is a standalone C++ engine typically driven from the command line, while docTR is a Python library built on PyTorch whose detection and recognition models are loaded through the pretrained flag. docTR returns a structured Document object with Page, Block, Line, Word and Artefact levels, which Tesseract's text output does not provide.

## Sources

- [License: Apache-2.0](https://github.com/mindee/doctr/blob/main/LICENSE)
- [mindee/doctr on GitHub](https://github.com/mindee/doctr)
- [Project website](https://mindee.github.io/doctr/)
- [README](https://github.com/mindee/doctr/blob/main/README.md)
- [Releases](https://github.com/mindee/doctr/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/mindee-doctr
