# tesserocr: a Cython wrapper whose Dockerfile is four years behind its own README

> 2,173 stars for calling the Tesseract C++ API directly from Python, with the GIL released so threads actually help. The bundled container image still targets Python 3.6 and Tesseract 3.04.

**sirfz/tesserocr** — A Python wrapper for the tesseract-ocr API

- Repository: https://github.com/sirfz/tesserocr
- Stars: 2,176 · Forks: 260
- Language: Python
- License: MIT
- Published: 2026-10-08 · Updated: 2026-10-08 · Language: en
- Canonical page: https://hysenlabs.com/projects/sirfz-tesserocr

## A wrapper thin enough to expose the C++ API directly

tesserocr is a Python wrapper around the `tesseract-ocr` C++ API, and it states its design plainly: it integrates directly with that API using Cython, which is what allows the Python-facing source to stay simple and readable. The library is described as Pillow friendly, meaning `PIL.Image` objects are accepted directly, and equally usable with plain image files.

The concurrency claim is the substantive one. It says the library enables real concurrent execution when used with Python's `threading` module by releasing the GIL while an image is processed. That is a differentiator, because a wrapper that held the GIL for the duration of a recognition call would make threads pointless no matter how they were configured. Version 2.10.0, published 2026-01-13, is consistent with that claim: its single change was placing long running API calls in cysignals blocks.

The scale is steady rather than explosive. The project reports 2,173 stars and 260 forks, 46 open issues, MIT licensed, Python, with Cython listed as a topic. The newest release is 2.11.0 on 2026-08-04, the same day as the last recorded push, and it is small: catch an exception in `get_languages()`, fix a test fixture, and fix a memory leak in the internal pix-to-image conversion.

## System libraries are the install story on Linux

Requirements are Python 3.9 or newer, libtesseract 3.04 or newer, libleptonica 1.71 or newer, and Cython 0.23 or newer for building, with Pillow optional but needed to accept `PIL.Image` objects. The README notes plainly that Python 2 is no longer supported.

On Debian and Ubuntu the dependencies are a single apt line.

```
apt-get install tesseract-ocr libtesseract-dev libleptonica-dev pkg-config
```

Installation then goes through pip, and the setup script attempts to detect the include and library directories, using pkg-config when it is available. Both escape hatches are documented for people whose libraries live somewhere unusual.

```
pip install tesserocr
```

```
CPPFLAGS=-I/usr/local/include pip install tesserocr
```

The README also warns that a more recent engine than the distribution ships may require compiling Tesseract yourself, and that a machine with more than one Tesseract or Leptonica installation will need `LD_LIBRARY_PATH` pointed at the right one. That warning is the practical crux of installing this library: the pip install is the easy part, and matching the native library versions is where builds fail.

## Windows relies on wheels someone else builds

The Windows story is the least self-contained part of the documentation. The recommended route is conda, using a third party channel.

```
conda install -c simonflueckiger tesserocr
```

The conda-forge channel is offered as an alternative. Beyond conda, the README describes proposed stand-alone downloads that bundle the Windows libraries needed for execution, so that no separate Tesseract install is required, and points at a separate build project for wheel files to install with pip.

If your Windows Python version is not covered there, the README points at a step by step guide for 64-bit Windows in a file named `Windows.build.md` at the repository root.

The heading markup in that section is worth noting if you read the raw reStructuredText. The Conda and pip subsections are underlined with runs of backticks of differing lengths rather than the conventional punctuation such characters should use, which means the source does not render as clean section levels even though it renders readably on the hosting site. Small thing, but it is the kind of thing that tells you the file is maintained by hand in a hurry.

The overall shape is that Linux and the BSDs get a first class install story and Windows gets a pointer to someone else's build artifacts. That is a normal division for a wrapper around a native library, and it is worth checking before committing to it on a Windows-only team.

## The Dockerfile in the repository root contradicts the README

The root `Dockerfile` and the README describe different software. The README requires Python 3.9 or newer. The Dockerfile starts from a Python 3.6.5 image on Debian stretch, a combination whose upstream support has been over for years, and it still uses the deprecated `MAINTAINER` instruction rather than a label.

It then pins its dependencies to versions that match the README's stated minimums rather than anything current: it downloads and compiles Leptonica 1.73 and Tesseract 3.04.01 from source archives, with a build cache argument dated 2016-01-01, and fetches version 4.00 of the trained data. The README's floor is libtesseract 3.04 and libleptonica 1.71, so this image sits exactly at the minimum and not above it.

The build itself is instructive in an old-fashioned way. It installs the autotools and image library headers, compiles the two native libraries in sequence with explicit library and include flags, runs ldconfig, then copies a fixed list of files into a temporary directory before building a wheel and installing it.

```
RUN pip install numpy Pillow opencv-python
```

That last line installs OpenCV, which the library does not depend on, and it does not install the `cysignals` package that `requirements.txt` lists, even though 2.10.0 moved long running calls into cysignals blocks. The honest reading is that this file is a historical snapshot kept for reference rather than a maintained build definition. The project's real current builds run through `cibuildwheel`, configured in `pyproject.toml` for manylinux on both x86-64 and aarch64, musllinux and macOS, each with a before-all step that installs the native dependencies.

## Trained data has to match the engine version

The tessdata section is short and the problem it describes is a recurring one. If the data directory cannot be detected automatically, you set the `TESSDATA_PREFIX` environment variable or pass a path to the API class directly, and the directory must contain the `.traineddata` files.

The critical instruction is the version match: the README says to make sure the trained data version corresponds to the engine version reported by the engine's version flag. Mismatched data produces recognition that is wrong rather than absent, which is the worst failure mode for a text extraction pipeline because it looks like success.

To see what is available, the README shows a call that prints the languages for a given path.

```python
from tesserocr import get_languages

print(get_languages('/usr/share/tessdata'))  # or any other path that applies to your system
```

The most recent release made exactly this function more defensive, catching the exception instead of propagating it. For a function whose job is to tell you whether your installation is configured correctly, raising an unhandled exception and returning an empty list are both unhelpful, so that fix is more interesting than a patch number suggests.

## Reuse one API instance, because the helpers are not the point

The usage section makes a distinction that should shape how you write against this library. It recommends initialising the API instance once and reusing it across multiple images, and it notes that the object is finalised automatically when used as a context manager, with an explicit end call otherwise.

```python
from tesserocr import PyTessBaseAPI

images = ['sample.jpg', 'sample2.jpg', 'sample3.jpg']

with PyTessBaseAPI() as api:
    for img in images:
        api.SetImageFile(img)
        print(api.GetUTF8Text())
        print(api.AllWordConfidences())
```

The convenience functions exist too, for text output from an image object or a file path, and the README says they can be used with threading to process multiple images concurrently. But it also says that the API class exposes several methods whose docstrings you should read, which is the tell that the higher level calls are conveniences rather than the real interface.

The advanced examples confirm that. One walks component images at the text line level, setting a rectangle from the returned box and re-running recognition per box, then printing a per-box confidence alongside the text. Another asks for automatic orientation and script detection, runs layout analysis, and prints orientation, writing direction and line order with a deskew angle.

Things like per-word confidence and text line geometry are exactly what you need for a document processing pipeline and exactly what a simple text extractor throws away, and that is the argument for a direct wrapper rather than a convenience one.

## Conclusion

tesserocr earns its place when you need layout analysis, confidence scores or bounding boxes rather than a string of text, because it exposes the Tesseract API rather than wrapping one convenience function, and because it releases the GIL so a thread pool genuinely parallelises recognition. The install story is the part that needs care: pip on Linux and the BSDs builds against system libraries, while Windows relies on wheels from a third party build project or a conda channel, and the Dockerfile at the root of this repository is a historical artifact rather than something to copy. Point TESSDATA_PREFIX at the trained data that matches your engine version, initialise the API once and reuse it, and read the docstrings on the exposed methods rather than guessing at parameters.

## FAQ

### How to use tesserocr?

Initialise `PyTessBaseAPI` once and reuse it, pass it a file or a Pillow image, then call the recognition methods and read their docstrings. The README shows a context manager over a list of images, and separately a set of helper functions that return plain text from a file or image object for cases where that is enough.

### Does tesserocr use threads effectively?

That is the project's main argument for existing. It states it enables real concurrent execution with Python's threading module by releasing the GIL while an image is processed, and version 2.10.0 moved long running API calls into cysignals blocks to that end. The helper functions are documented as usable with threading for concurrent image processing.

### Why does tesserocr fail to build on my machine?

Almost always the native libraries. It needs libtesseract 3.04 or newer and libleptonica 1.71 or newer, and the setup script detects include and library directories via pkg-config, overridable with flags such as `CPPFLAGS`. A machine with several Tesseract or Leptonica installations also needs `LD_LIBRARY_PATH` pointed at the right one.

### How do I get trained data for languages other than English?

Point `TESSDATA_PREFIX`, or pass a path to the API class, at a directory containing the `.traineddata` files, and make sure that data version matches your engine version. Call `get_languages()` with that path to list what is installed. Mismatched data yields wrong text rather than an error, so verify it.

## Sources

- [Issues](https://github.com/sirfz/tesserocr/issues)
- [License: MIT](https://github.com/sirfz/tesserocr/blob/master/LICENSE)
- [README](https://github.com/sirfz/tesserocr/blob/master/README.md)
- [Releases](https://github.com/sirfz/tesserocr/releases)
- [sirfz/tesserocr on GitHub](https://github.com/sirfz/tesserocr)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/sirfz-tesserocr
