OCRmyPDF writes the text layer under the scan, and leaves the recognizing to Tesseract
OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched
At a glance
- What is it?
- OCRmyPDF is a pure Python command line program that makes scanned PDFs searchable and emits PDF/A, writing recognized text below the original page image without redrawing it. It cannot run on a host that has only the Python package: Ghostscript, Tesseract 4.1.1 or newer and Python 3.11 are all hard requirements, and the language data behind -l is a separate install.
- Who is it for?
- OCRmyPDF suits an archive or records team holding thousands of page scans, with a reason to want PDF/A, and hosts where an administrator can install Ghostscript and Tesseract. Check three things before committing: that tesseract on the PATH reports 4.1.1 or newer, that the language packs for every script in the corpus are installed as Tesseract data rather than assumed, and that Python is 3.11 or newer with the pikepdf, Pillow and pdfminer.six floors satisfied.
- Can I use it commercially?
- Yes, with conditions. MPL-2.0 is a weak copyleft licence: you can use it inside commercial and closed-source software, but if you distribute changes to its own files, you must publish those changes under the same licence.
- Is it still maintained?
- Yes. The repository last received commits 2 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The text layer goes under the image, which is the whole point
OCRmyPDF does not redraw the page. It opens a PDF that already contains a page image, runs recognition, and writes the recognized text into the same page as a layer positioned below the image. That placement is what makes a copy from a search hit paste back in reading order instead of as scattered fragments, and the author puts that defect first on his list of reasons to write his own tool: the free command line tools he tried put the text in the wrong place. The rest of that complaint list maps onto concrete promises. Accents and multilingual characters are handled through Tesseract's language packs, changed image resolution is avoided by keeping the exact resolution of the original embedded images, oversized output is addressed by optimizing PDF images, crashes and invalid files by validating input and output, and the absence of an archival format by producing PDF/A.
Two hedges deserve a close read before you promise anyone a smaller archive. Inserting OCR data is a "lossless" operation only "when possible", the best case rather than a guarantee about every input. And image optimization "often" yields a file smaller than the input, a tendency rather than a number to plan storage around. The project also describes itself as battle-tested on millions of PDFs, a claim about its own history that carries no measurement behind it.
Ghostscript and Tesseract sit outside the wheel, so pip alone is not enough
A Python install of OCRmyPDF does not produce a working program. The project is pure Python, but it requires external installations of Ghostscript and Tesseract OCR, so on a clean machine the command resolves and then fails until both binaries exist on the system. The packaging metadata sets requires-python to >=3.11, which shuts out older interpreters entirely, and the dependency floors are set high: pikepdf[pdfa]>=10.15, Pillow>=12, pdfminer.six>=20260107, pypdfium2>=5.0.0 and uharfbuzz>=0.53.2, alongside fpdf2>=2.8.0, img2pdf>=0.5, packaging, pydantic, rich and pluggy, with typing-extensions pulled in below Python 3.13.
The pdfminer.six floor is not arbitrary. The comment beside it in pyproject.toml ties the minimum to parsing of tokens split across the read buffer and streams, reference gh #1361, so an older copy already present in a shared environment will not serve. For a long-term-support base image, those floors mean the work is upgrading Python before anything OCR related. The build backend is hatchling, the license is MPL-2.0, the version in the metadata is 17.13.0, and James R. Barlow is the listed author.
PATH order picks the Tesseract, not you
Recognition cannot be pinned to a specific engine binary. Support starts at Tesseract 4.1.1, and OCRmyPDF takes whichever version it meets first on the PATH environment variable. On Windows, if PATH yields no Tesseract binary at all, it falls back to the highest version number installed according to the Windows Registry. Both rules hand the decision to the machine rather than to the operator. Put a tesseract older than 4.1.1 earlier in PATH and that is the binary under test, and a registry fallback on Windows can select a version nobody chose deliberately.
The practical consequence is diagnostic rather than theoretical. When recognition quality looks wrong on a batch, check which tesseract answers first on PATH and what version it reports, before suspecting the PDF handling, because the tool's own answer to that question is not in your hands. The rest of the option set is not summarised anywhere in the repository, and the built-in help is where the reader is pointed:
ocrmypdf --helpThe language packs ship separately, and -l is only a hint
The claim of more than 100 languages belongs to Tesseract's tessdata, not to OCRmyPDF. Nothing in an install of ocrmypdf places a language pack on disk, so passing -l fra on a host that carries only the English data yields a valid PDF/A file with nothing recognized in French, and no error reports the missing pack. The argument is worded as a hint about which languages to search for, and several can be joined with a plus sign, which means a bilingual page is recognized against two models. The cost of that hint is time on every page, multiplied by the number of languages named.
Packs are installed per platform, and the names differ: apt-get install tesseract-ocr-chi-sim for Simplified Chinese on Debian and Ubuntu, pacman -S tesseract-data-eng tesseract-data-deu on Arch, pkg_add tesseract-cym on OpenBSD, brew install tesseract-lang on macOS, and a search for tesseract-langpack on Fedora. A corpus with mixed scripts therefore has an install-time checklist, not a command-line one, and nothing in the tool verifies that the codes passed to -l correspond to packs that are actually present.
Skew and rotation stay off until a flag asks for them
Both page geometry fixes are opt-in in the project's own signature example, and the feature list is explicit that it deskews and/or cleans the image before performing OCR if requested. A plain conversion of a tilted phone scan leaves the tilt where it was, the recognizer reads a skewed baseline, and the text layer inherits those errors even though its placement below the image is accurate. Correct placement buys back nothing that damaged recognition lost, which is the failure pattern that produces a searchable PDF full of plausible but wrong words.
Two flags in the same block matter on a first run. --title sets output metadata, and --jobs narrows a fan-out that otherwise spreads across all available CPU cores, which is what you want on a shared host rather than on a dedicated workstation. A second gap is worth naming: the README does not document what happens when the input already carries a selectable text layer, so treat that input as undefined rather than assuming detection and skipping.
Passing one path twice overwrites the only copy of the scan
One of the documented examples hands the tool the same filename as both input and output, described as adding OCR to a file in place, and it modifies that file only on success. Read carefully, the guarantee covers the failure case and nothing more. On a successful run there is no surviving copy of the original scan, because the write landed on the same path, and validating input and output makes the success path trustworthy without creating a backup of anything. The French-language example does the same thing to a non-English file, which makes the pattern easy to absorb by habit.
When a scan is the only copy of a document, the two-path form is the one to use, and the header example shows the intended shape: a distinct output name after the input, with the output type spelled out. On large files the same rule applies, because the project says it scales to documents with thousands of pages, and an in-place rewrite is a single write with no partial output to recover. A pipeline that runs unattended over a shared volume should decide this per file, not by habit.
The plugin list is two entries and one of them needs macOS
A pluggy-backed plugin interface allows the capabilities to be extended or replaced, and two plugins are named: OCRmyPDF-AppleOCR, which replaces the standard Tesseract engine with the Apple Vision Framework and requires macOS, and OCRmyPDF-EasyOCR. The framing is careful enough to matter for planning, because these are the plugins the project is aware of rather than a curated registry with compatibility ranges attached. A pipeline decision should not rest on that list being current or complete.
Substituting the engine also invalidates part of the language story told above. The tessdata packs, the -l codes and the more than 100 languages figure all describe the Tesseract path, so a build that swaps in Apple Vision should not assume the same language workflow or the same pack names still apply, and the README does not document a replacement for that workflow. Portability is the second cost: an engine that requires macOS ends the ability to run the same pipeline on the Linux hosts that normally absorb batch OCR work.
Ten install rows, two container architectures, and a release every few weeks
The install table is the first thing to read, because it encodes where the dependencies come from as much as it encodes the install. Debian and Ubuntu use apt, and the Windows Subsystem for Linux takes the same route. Fedora uses dnf, macOS has three choices, LinuxBrew reuses the Homebrew line, FreeBSD ships py-ocrmypdf, OpenBSD uses pkg_add, and the Ubuntu Snap is a separate row. Docker images exist for both x64 and ARM, which is the practical answer for a fleet where each host would otherwise need its own Ghostscript and Tesseract pairing. On the rate of change, the main branch was last pushed on 2026-09-27, v17.13.0 was published on 2026-09-28, and v17.12.0 and v17.12.1 both landed on 2026-09-16, a feature release and its patch in the same afternoon. Nothing in the repository is marked archived. Four of the routes cover most readers:
apt install ocrmypdf
brew install ocrmypdf
nix-env -i ocrmypdf
snap install ocrmypdfEditorial conclusion
OCRmyPDF suits an archive or records team holding thousands of page scans, with a reason to want PDF/A, and hosts where an administrator can install Ghostscript and Tesseract. Check three things before committing: that tesseract on the PATH reports 4.1.1 or newer, that the language packs for every script in the corpus are installed as Tesseract data rather than assumed, and that Python is 3.11 or newer with the pikepdf, Pillow and pdfminer.six floors satisfied. A team that needs a browser service, per-document engine selection, or a Windows-native binary outside the Windows Subsystem for Linux should look at the other options first.
Frequently asked questions
How do I use OCRmyPDF?
It is a scriptable command line program that takes an input PDF or image and writes an output PDF, for example ocrmypdf --output-type pdfa input.pdf output.pdf. Options cover languages, page rotation, deskewing, output metadata, job count and output type, and ocrmypdf --help prints the full syntax.
What does ocrmypdf do?
It adds an OCR text layer to scanned PDF files so their contents can be searched and copied, placing the recognized text below the embedded image while keeping the exact resolution of the original images. Output is PDF/A, and the input and output files are both validated.
How do I install ocrmypdf on mac?
Three routes are listed for macOS: brew install ocrmypdf, port install ocrmypdf and nix-env -i ocrmypdf. Ghostscript and Tesseract OCR still have to be installed separately, and the Homebrew route to Tesseract language data is brew install tesseract-lang.
How do I install ocrmypdf on Windows?
The Windows support in the install table runs through the Windows Subsystem for Linux, using the same apt install ocrmypdf command as Debian and Ubuntu. Tesseract 4.1.1 or newer is required, and if PATH provides no Tesseract binary the highest version installed according to the Windows Registry is used instead.
Is OCRmyPDF free?
It is published under the MPL-2.0 license and distributed through platform package managers and Docker images, with no hosted or paid tier in the repository. The project states that Docker images are available for x64 and ARM, and points anyone not covered by its install table to the documentation page.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/ocrmypdf-ocrmypdf)