Library / SDK
madmaze/pytesseract avatar
madmaze/pytesseract

pytesseract: a thin Python wrapper around the Tesseract OCR binary

A Python wrapper for Google Tesseract

6,393 stars747 forksPythonApache-2.0

At a glance

What is it?
pytesseract does not perform optical character recognition itself. It builds command line calls to a Tesseract executable you install separately, then parses what comes back. That distinction decides whether the package fits your project.
Who is it for?
Adopt pytesseract when Tesseract is already your OCR engine and you want it callable from Python without hand-building subprocess arguments. Do not adopt it expecting a self-contained OCR library; it is a wrapper, and a missing or misconfigured tesseract binary produces errors that the README addresses only with a tesseract_cmd example and a tessdata-dir note.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 3 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What pytesseract actually wraps, and who needs that

The README states plainly that Python-tesseract is a wrapper for Google's Tesseract-OCR Engine, and that it will recognize and read the text embedded in images. The wrapper part is the whole story. pytesseract contains no OCR model, no trained data and no recognition code. It constructs a command line invocation of the tesseract executable, runs it, and converts the output into Python objects or strings. Everything about recognition quality comes from the engine and the language data installed alongside it.

That makes the package useful to a specific kind of user. If you already run Tesseract, perhaps from a shell script or a batch pipeline, and you want the same behaviour inside a Python program, pytesseract saves you from assembling argument lists by hand and parsing stdout yourself. The README also notes it works as a stand-alone invocation script, printing recognized text instead of writing it to a file, which suits quick checks from a terminal.

It is the wrong choice when you want a single pip install to give you working OCR with no system dependency. There is no bundled engine. A container or a CI runner that lacks the tesseract binary will fail at call time, and the failure surfaces as an exception from the subprocess layer rather than an import error, which is why the module-not-found question below has a different answer than people expect.

How a call reaches Tesseract and what comes back

The data flow is one-directional and stateless. You pass an image, pytesseract writes it to a temporary file if it was a PIL or NumPy object, spawns the tesseract process with the options you supplied, reads the process output, and returns it in the shape the function promises. The README documents several shapes: unmodified text from image_to_string, character boxes from image_to_boxes, boxes with confidences and line and page numbers from image_to_data, orientation and script detection from image_to_osd, and ALTO XML from image_to_alto_xml.

The README gives an explicit reason to avoid the conversion step. Passing a file path directly, rather than a PIL Image, bypasses the image conversions pytesseract performs, with the caveat that you must supply a format Tesseract itself supports or the engine returns an error. That is a real trade-off: fewer moving parts and no temporary files, but the format check moves from Pillow to Tesseract, and Pillow is more permissive.

One function is worth singling out. run_and_get_multiple_output replaces the single extension argument with a list, and the README says it returns the corresponding data after only one tesseract call, supporting mix and match of txt, pdf, hocr, box and tsv. If you need both the plain text and the bounding boxes for the same image, this avoids paying the recognition cost twice. Recognition is the expensive part of the pipeline, so that is a meaningful design choice rather than a convenience wrapper.

Configuration passes through to the engine. The README shows oem and psm supplied through the config keyword, and also shows loading a predefined Tesseract config file by name, in its example the string words. Nothing in pytesseract interprets those values; they are forwarded.

Installing pytesseract and running a first recognition

The README does not give pip or conda install commands, but the badge targets it references are the PyPI page and the conda-forge package, and the repository includes setup.py and setup.cfg, so the package is distributed through both channels. The order matters: install the Tesseract engine first, then the Python package. If the executable is not on your PATH, the README shows how to point the wrapper at it.

python
import pytesseract

pytesseract.pytesseract.tesseract_cmd = r'<full_path_to_your_tesseract_executable>'
# Example tesseract_cmd = r'C:\Program Files (x86)\Tesseract-OCR\tesseract'

Setting tesseract_cmd is the documented escape hatch for a binary that is installed but not discoverable. The README's own example is a Windows path, which is where this comes up most often.

A minimal first run takes an image and prints the text. The README's quickstart uses a test image from the tests/data folder of the repository.

python
from PIL import Image
import pytesseract

print(pytesseract.image_to_string(Image.open('test.png')))

If that prints nothing, the call succeeded and the engine found no text, which is a different problem from an exception. Before debugging recognition, confirm the engine is reachable and which languages are present.

python
print(pytesseract.get_tesseract_version())
print(pytesseract.get_languages(config=''))

get_tesseract_version returns the Tesseract version installed on the system, and get_languages returns the languages Tesseract currently supports. If the second call does not list the language you passed to lang, the engine will not recognize that script regardless of what your Python code does.

The README also covers the tessdata error, which reads Error opening data file. Its fix is a config string pointing at the data directory, with double quotes around the path.

python
tessdata_dir_config = r'--tessdata-dir "<replace_with_your_tessdata_dir_path>"'
pytesseract.image_to_string(image, lang='chi_sim', config=tessdata_dir_config)

The README stresses that the double quotes around the directory path are important, which is the kind of detail that costs an afternoon when omitted.

Timeouts, OpenCV arrays and the limits of the wrapper

Long-running recognition is handled with a timeout argument, and the README is explicit that Tesseract processing is terminated when it fires. The exception raised is RuntimeError, and the README's example catches exactly that. This matters in services: without a timeout, a pathological image can hold a worker indefinitely, and pytesseract gives you a documented way to bound it.

OpenCV input is supported, but with a colour-order trap. OpenCV stores images in BGR, while pytesseract assumes RGB, so the README converts with cv2.cvtColor before calling image_to_string. Skipping that conversion does not raise; it feeds the engine wrong channel data and degrades the text. The README offers a second route using Image.frombytes with a raw BGR mode, which avoids a copy for large arrays.

The genuine limitation is architectural. Every call spawns a process. There is no persistent engine handle, no batching API beyond the multi-extension call, and no way to keep a model resident across requests. For a handful of images this is invisible. For a queue of thousands, process startup and file I/O become a measurable share of the work, and pytesseract offers no mechanism to avoid it. The README does not document connection pooling, worker reuse or a server mode, because the wrapper has none.

A second limitation is diagnostic. When recognition is poor, pytesseract gives you the engine's output and nothing more. Choosing page segmentation modes, preprocessing images, or tuning the engine is outside the wrapper's scope; the config keyword forwards your choices, it does not make them.

pytesseract against EasyOCR, and how the approaches differ

The comparison people search for is pytesseract versus EasyOCR, and the difference is not quality tuning, it is packaging. pytesseract is a wrapper around a system binary that you install and configure separately. EasyOCR is a Python package built on deep learning models that it downloads and runs in-process, with no external executable to locate.

That changes the deployment story completely. With pytesseract, your Dockerfile or CI image must install the Tesseract engine and the language data, and the wrapper must find the binary, either on PATH or through tesseract_cmd. With an in-process library, the dependency arrives with the package. The trade is weight and startup behaviour: model-based recognizers pull large weights and consume memory and compute per process, while Tesseract is a comparatively small native program that starts fast and reads trained data from disk.

There is a second axis, which is control. pytesseract forwards engine options such as oem and psm and predefined config files, so anyone who has tuned Tesseract from the command line can carry that tuning into Python unchanged. A model-based library exposes its own parameters instead, and your existing Tesseract configuration does not transfer.

Neither is universally better, and the README makes no claim about accuracy against other engines. The honest framing is that pytesseract is the right shape when Tesseract is already your engine or your environment can install native packages, and the wrong shape when you need a pure-Python dependency with no system-level install step.

Maintenance, release cadence and the Apache-2.0 licence

The repository is not archived, and the last push was on 2026-07-13. The most recent tagged release in the list is v0.3.13 from 2023-10-15, preceded by v0.3.12 and v0.3.11 in the two months before it. That is the maintenance picture: commits continue on master, while tagged releases have been infrequent since late 2023. Anyone pinning a version should expect to track master for fixes that have not been cut into a release, and should check the commit history rather than the release list when judging whether a specific bug is addressed.

The repository carries pre-commit configuration and a tox setup, and its CI runs through GitHub Actions, so contributions are gated by automated checks. The upgrade surface is small by design. The package is a wrapper; its public functions map to engine capabilities, and the main upgrade risk is not API churn inside pytesseract but a change in the Tesseract binary it calls, since options like psm and tessdata layout are the engine's, not the wrapper's.

The licence is Apache-2.0, as listed for the repository. That is a permissive licence with an explicit patent grant, and it applies to the wrapper code. It does not cover Tesseract itself, which is a separate project with its own licensing, nor the language data files you install. If you redistribute a container image that bundles the engine and trained data, the wrapper's licence tells you nothing about those components. This is not legal advice; check the terms of each component you ship.

Editorial conclusion

Adopt pytesseract when Tesseract is already your OCR engine and you want it callable from Python without hand-building subprocess arguments. Do not adopt it expecting a self-contained OCR library; it is a wrapper, and a missing or misconfigured tesseract binary produces errors that the README addresses only with a tesseract_cmd example and a tessdata-dir note. Before writing code against it, verify three things on the target machine: that get_tesseract_version() returns without raising, that get_languages() lists the language packs you need, and that a single image_to_string call on a representative file gives usable text. If the third check fails, the problem is the engine or the image, not the wrapper, and no amount of Python changes will fix it.

Frequently asked questions

What is pytesseract used for?

It is used to run Google's Tesseract OCR engine from Python. The README describes it as a wrapper that recognizes and reads the text embedded in images, and it can also be used as a stand-alone invocation script that prints recognized text instead of writing it to a file.

Why does Python say No Module Named pytesseract?

That error means the pytesseract package is not installed in the Python environment you are running, not that the Tesseract engine is missing. The README references PyPI and conda-forge for distribution, so installing through one of those channels for the same interpreter resolves it.

What are the differences between Tesseract and Pytesseract?

Tesseract is the OCR engine; pytesseract is a Python wrapper that calls it. The README states that pytesseract is a wrapper for Google's Tesseract-OCR Engine, so recognition is performed by the engine and the wrapper handles invocation and output parsing.

Which is better, pytesseract or EasyOCR?

The README makes no accuracy comparison between the two. The structural difference is that pytesseract wraps a separately installed Tesseract binary, while EasyOCR is a Python package that runs its own models in-process, so the choice depends on whether your environment can install a native OCR engine.

How do I install pytesseract?

The README does not spell out an install command, but its badges point to the PyPI page and the conda-forge package, and the repository ships setup.py and setup.cfg. Install the Tesseract engine first, then the Python package, and set tesseract_cmd if the executable is not on your PATH.

How do I use pytesseract in Python?

Import pytesseract, open the image with Pillow, and call image_to_string, as the README quickstart does. If the tesseract executable is not on your PATH, set pytesseract.pytesseract.tesseract_cmd to its full path first.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. madmaze/pytesseract on GitHub
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/madmaze-pytesseract.svg)](https://hysenlabs.com/projects/madmaze-pytesseract)