Open-source project
tesseract-ocr/tesseract avatar
tesseract-ocr/tesseract

Tesseract OCR: two recognition engines behind one binary, and no window to click

Tesseract Open Source OCR Engine (main repository)

76,679 stars10,817 forksC++Apache-2.0

At a glance

What is it?
The package is libtesseract plus a command line program, with an LSTM engine and a legacy character-pattern engine selected per run. The output formats, the traineddata dependency and the absent GUI are where an integration is actually decided.
Who is it for?
Tesseract earns its place in a batch pipeline: it is a headless engine with an output format list that includes invisible-text-only PDF and TSV, and the format you pick determines what your downstream consumer has to parse. Two things rule it out for other readers.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 19 days ago.
What is it written in?
Mainly C++, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.

Editorial analysis

One binary, two OCR engines, and a flag that picks between them

Tesseract 4 introduced a neural net engine based on LSTM, focused on line recognition, and that engine is what a current install uses. The older engine from Tesseract 3 is still present and works a different way, by recognizing character patterns. Both sit behind the same `tesseract` program, and the switch between them is the OCR engine mode flag, where `--oem 0` selects the legacy engine for Tesseract 3 compatibility.

The choice is therefore a per-invocation decision rather than a property of your build, and turning compatibility mode on costs more than the flag itself. The legacy engine also needs traineddata files that support it, drawn from the tessdata repository, so a language that works under the LSTM engine can be missing under the other. Anything comparing accuracy between two setups has to record which engine produced the text, because the same image and the same language can yield different answers under each.

Seven output formats, and the pick is a pipeline decision

Recognition gets the attention, but the output list is what decides your integration. Input covers PNG, JPEG and TIFF. On the way out the engine produces plain text, hOCR as HTML, PDF, invisible-text-only PDF, TSV, ALTO and PAGE, and those are not cosmetic variants of each other. Plain text drops the geometry. TSV keeps positions in columns, hOCR carries the same information as markup, and invisible-text-only PDF places a recognized text layer behind a scanned page so a search index can read it without altering the image.

Invocation follows a fixed shape, with the output base name carrying the format choice:

bash
tesseract imagename outputbase [-l lang] [--oem ocrenginemode] [--psm pagesegmode] [configfiles...]

Page segmentation mode is the other parameter worth understanding before tuning anything, since it tells the engine how the page is structured before recognition starts. `tesseract --help` and `man tesseract` cover the remainder, with worked examples in the command line usage documentation.

The repository ships no GUI and says so outright

What the package contains is an OCR engine, `libtesseract`, and a command line program, `tesseract`. That is the entire user interface. The README states plainly that the project does not include a GUI application, and sends anyone who needs one to a third-party projects page.

For most pipelines that is a virtue, because a headless binary is what a batch job wants. For a person who downloads the tool expecting a window to drop an image into, it is the first surprise, and no flag in the command line resolves it. Packaging exists around the engine without amounting to an interface: the tree carries an `nsis/` directory, a `snap/` directory, an `appveyor.yml` and a `Makefile.am`, none of which add a front end. The consequence is that an evaluation on a desktop means pairing this binary with separate viewer software, and any impression of the engine arrives through a program this repository does not ship.

More than 100 languages, each one a traineddata file

Tesseract has UTF-8 support and is described as recognizing more than 100 languages out of the box, which sounds like a build-time feature and is not. Language capability lives in traineddata files, and the legacy engine needs its own set drawn from the tessdata repository. The repository itself carries a `tessdata/` directory, so the engine and its data files are versioned in the same place.

The consequence is that adding or switching a language is a file operation on your side, and the file you need depends on the engine you selected. A deployment standardizing on `--oem 0` for compatibility has to standardize the legacy data set as well, or it discovers missing languages at run time instead of at install time. Tesseract can also be trained to recognize other languages outright, with training documentation covering that route, which is a far larger commitment than dropping in a prepared file.

Poor scans produce poor text, and the repair happens upstream

Among the least technical sentences in the README is also the one that decides most disappointing first runs: in many cases, to get better OCR results you need to improve the quality of the image you give Tesseract. Input is limited to PNG, JPEG and TIFF, and nothing on the command line repairs a low-resolution, skewed or badly lit capture. A dedicated page in the documentation covers improving image quality, which is a fair signal about where the project expects effort to go.

The consequence is diagnostic rather than technical. Garbled characters, dropped lines and wrong separators all read as engine defects, and the instinct is to change engines, page segmentation modes or language files when the cause is often the scan itself. Teams adopting Tesseract for document ingestion should treat capture quality as part of the interface contract and measure it, instead of treating recognition accuracy as a property of the binary alone.

CMake and autotools both build the same tree

Two build systems are maintained side by side. The CMake path is `CMakeLists.txt` with a `cmake/` directory, and the autotools path is `configure.ac` with `Makefile.am`, `m4/` and an `autogen.sh` that generates the configure script. Installation guidance points at a separate document for each route, and both `INSTALL` and `INSTALL.GIT.md` sit at the top level, a sensible way to cover a pre-built binary package and a source checkout independently.

Downstream packaging also gets what it needs for linking, since `tesseract.pc.in` and `tesseract.pc.cmake` produce pkg-config metadata and let a consumer ask the build system where `libtesseract` lives rather than hardcoding a path. The one precondition worth not skipping is unglamorous: check that your compiler is one of the supported compilers before building from source. The consequence of ignoring it is a build failure that says nothing about Tesseract itself.

The first-party API is C and C++, so Python arrives as a wrapper

Embedding Tesseract means linking `libtesseract` through one of two headers, `include/tesseract/capi.h` for C and `include/tesseract/baseapi.h` for C++. Those are the surfaces this repository maintains. Bindings for other languages are not part of it, and the README points you to the wrapper section of the AddOns documentation instead. The top-level tree carries a `java/` directory and no first-party Python package, which makes the boundary concrete.

The consequence lands hardest on the most common scripting use. In a Python pipeline, the quality and update cadence of your OCR live in a wrapper maintained elsewhere, while the native library underneath it is the part you control. Pinning therefore matters twice: for the engine, where the 5.x line has been stable since 5.0.0 on November 30, 2021, and for the wrapper, which can lag behind it. API documentation is generated from the source with doxygen and published on the project site, so read the headers before designing around either.

Ownership moved from HP to Google and then to named maintainers

Tesseract was developed at Hewlett-Packard Laboratories Bristol UK and at Hewlett-Packard Co in Greeley, Colorado between 1985 and 1994, with a Windows port in 1996 and the C++ conversion in 1998. HP open sourced it in 2005, and Google developed it from 2006 until August 2017. Ray Smith was lead developer until 2017, Stefan Weil is the current lead developer, and Zdenko Podobny is the maintainer, with `AUTHORS` and `CITATIONS.bib` at the top level of the tree.

The repository is not archived, and its last push was 2026-09-11, with 5.5.3 released on 2026-07-24 following 5.5.2 on 2025-12-26 and 5.5.1 on 2025-05-25. Support runs through the documentation and its FAQ first, then two mailing lists, one for users and one for developers, under an explicit rule that issues are for bugs rather than for asking questions. For an adopter that structure has a practical effect: answers accumulate in archives others can read, and the issue tracker tells you about known defects rather than whether yours is known.

Editorial conclusion

Tesseract earns its place in a batch pipeline: it is a headless engine with an output format list that includes invisible-text-only PDF and TSV, and the format you pick determines what your downstream consumer has to parse. Two things rule it out for other readers. A team that needs a desktop viewer is looking at third-party software, and a team that needs a first-party Python binding is depending on a wrapper maintained outside this repository. Verify three things before committing: which engine mode your accuracy target was measured under, whether your languages have traineddata files for that same mode, and whether your captures clear the image quality bar the project points at.

Frequently asked questions

How do I use Tesseract OCR from the command line?

Invocation takes an image name and an output base name, with optional language, OCR engine mode and page segmentation mode flags, plus any config files. `tesseract --help` and `man tesseract` list the remaining options, and the command line usage documentation carries worked examples.

How do I install Tesseract?

You can install a pre-built binary package or build from source, and each route has its own page in the documentation. Before compiling, check that your system compiler is one of the supported compilers listed there.

How do I use Tesseract OCR in Python?

There is no first-party Python binding in this repository. Python code reaches the engine through a wrapper around `libtesseract`, and the wrapper section of the AddOns documentation is where those are listed, with the C API in `include/tesseract/capi.h` as the underlying surface.

Does Tesseract come with a graphical interface on Windows?

No. The package contains the `libtesseract` engine and the `tesseract` command line program, and the project does not include a GUI application. If you need one, the documentation points to a page of third-party user projects.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
  4. Release notes
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/tesseract-ocr-tesseract.svg)](https://hysenlabs.com/projects/tesseract-ocr-tesseract)