llm_aided_ocr: LLM Correction on Top of Tesseract for Scanned PDFs
Enhances Tesseract OCR output using LLMs (local or API) for error correction, smart chunking, and markdown formatting of scanned PDFs
At a glance
- What is it?
- A Python pipeline that runs Tesseract, then sends the text through a local or API LLM to fix OCR errors and emit markdown. It is useful when Tesseract output is nearly right but not quite, and it costs you a second model pass over every chunk.
- Who is it for?
- Adopt it if you have a scanned PDF where Tesseract is already close and you want markdown with fewer character-level errors, and you can accept a second model pass over every chunk. Skip it if your scans are low quality, if you need a deterministic byte-identical output, or if you have no GPU and no API budget, since the correction step is the whole point of the tool.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 58 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The gap between Tesseract output and a usable document
Tesseract is good at recognising glyphs and bad at knowing what a word should have been. A scanned letter comes out with 'rn' where 'm' belongs, broken hyphenation, stray page numbers, and inconsistent spacing. The result is readable to a human and wrong to a parser.
llm_aided_ocr targets that second layer. The README describes the goal as transforming raw OCR text into accurate, well-formatted and readable documents, and the pipeline is explicitly two-stage: Tesseract extracts, an LLM corrects and optionally formats. The audience is anyone with a scanned document who needs clean text rather than an image of text: researchers digitising archives, engineers building retrieval over old PDFs, people who want markdown instead of a text dump.
The project does not replace Tesseract and does not claim to. It assumes the OCR pass already produced something close enough that a language model can repair it. That assumption is the whole design, and it is also where the tool breaks down.
How the pipeline moves a PDF from page image to corrected markdown
The flow is linear and visible in the function names the README lists. `convert_pdf_to_images()` uses `pdf2image` to turn pages into images, with `max_pages` and `skip_first_n_pages` parameters for working on a subset. `ocr_image()` calls `pytesseract`, and before it does, `preprocess_image()` converts to grayscale, applies binary thresholding with Otsu's method, and dilates to sharpen text.
From there `process_document()` splits the text into chunks on sentence boundaries with an overlap between them so context survives the cut. Each chunk goes through `process_chunk()`, which first asks the LLM to fix OCR errors while preserving structure, then optionally asks it to convert to markdown. Duplicate paragraph removal happens inside the markdown step. Header, footer and page number suppression is optional and configurable.
LLM calls are abstracted behind three functions: `generate_completion_from_local_llm()` for `llama_cpp` inference with optional grammars, `generate_completion_from_claude()`, and `generate_completion_from_openai()`. API calls run concurrently under `asyncio`, and the README states chunk order is maintained so the final document stays coherent. Token handling is explicit: `estimate_tokens()` prefers model-specific tokenizers and falls back to `approximate_tokens()`, while `TOKEN_BUFFER` and `TOKEN_CUSHION` adjust `max_tokens` against model limits. A final `assess_output_quality()` call asks the LLM to score the output against the original OCR text.
Installing llm_aided_ocr and running your first PDF
The README requires Python 3.12 or newer, the Tesseract engine, and the packages in `requirements.txt`. It recommends pyenv to pin the interpreter. On Ubuntu the Tesseract engine itself comes from the system package manager, on macOS from Homebrew, and on Windows from the UB-Mannheim build linked in the README.
Clone the repository, pin the interpreter, and install into a virtual environment:
git clone https://github.com/Dicklesworthstone/llm_aided_ocr
cd llm_aided_ocr
pyenv local 3.12
python -m venv venv
source venv/bin/activate
python -m pip install --upgrade pip
pip install -r requirements.txtThen install the OCR engine if it is not already present:
sudo apt-get install tesseract-ocrConfiguration lives in a `.env` file, read through `python-decouple`. The README gives this starting point, and you only need the key for the provider you actually use:
USE_LOCAL_LLM=False
API_PROVIDER=OPENAI
OPENAI_API_KEY=your_openai_api_key
ANTHROPIC_API_KEY=your_anthropic_api_keyUsage is file-based rather than argument-based. Place the PDF in the project directory and point the input filename variable at it in the script, then run `llm_aided_ocr.py`. The repository also contains `llm-aided-ocr-cli.py` for a command-line entry point. Two outputs appear next to the input: `{base_name}__raw_ocr_output.txt` with the untouched Tesseract text, and `{base_name}_llm_corrected.md` or `.txt` with the corrected version. The repository ships a worked example, the Warren Buffett and Katharine Graham letter, with both the raw OCR text and the corrected markdown checked in, which is the fastest way to see what the correction pass actually changes.
Where the correction pass makes things worse
The failure mode is inherent to the architecture. An LLM asked to correct text will also rewrite it. The README's instruction is to maintain original structure and content, but there is no diff gate, no confidence threshold, and no rule that rejects a chunk whose corrected form diverges too far from the input. If Tesseract mangled a proper noun into something that looks like a different real word, a language model has no way to know which one was on the page. It will pick the plausible one.
That makes the tool a poor fit for anything where the exact characters matter: legal citations, part numbers, financial figures, names in a genealogical record. The `assess_output_quality()` step produces a score and an explanation, but it is itself an LLM judgement against the original OCR text, so it cannot detect a confident hallucination that reads well. Treat the score as a signal, not a verification.
There are practical constraints too. Python 3.12 or newer is a hard floor, which rules out older environments. Local inference pulls in `llama-cpp-python` and a compatible GGUF model, and the README notes GPU acceleration is needed for local LLM inference to be practical. API mode moves the cost to per-token billing and sends your document text to a third party, which is a non-starter for confidential scans. Neither path is free.
Finally, the README does not document rollback, resumption after a failed run, or a way to reprocess only the chunks that failed. On a long PDF with an API rate limit or a dropped connection, that gap matters more than the correction quality.
llm_aided_ocr versus Tesseract alone and versus a multimodal model
The honest comparison is against plain Tesseract, because that is what the LLM OCR versus Tesseract question is really asking. Tesseract alone is deterministic, offline, fast, and free. Run it twice on the same page and you get the same bytes. llm_aided_ocr gives that up: the same input can produce different output across runs, and the text now depends on a model you did not train and cannot audit. What you buy is formatting, markdown structure, duplicate removal, and repair of errors that a language model can infer from context. If your Tesseract output is already clean, the correction pass adds risk without adding much.
The other alternative is a multimodal model that reads the page image directly instead of reading Tesseract's text. That is a different architecture: the model sees the layout, the tables, and the handwriting, and there is no separate OCR stage to correct. It also means no Tesseract dependency and no preprocessing step. The trade-off runs the other way. You lose the deterministic raw text that llm_aided_ocr writes to `__raw_ocr_output.txt`, which is your only ground truth for checking what the model changed, and you pay image tokens rather than text tokens for every page. For a clean printed document, correcting Tesseract is cheaper. For a messy one, reading the image directly is more likely to work, and llm_aided_ocr's correction step has less to work with.
Licence, maintenance and what an upgrade actually costs
The repository's licence is reported as NOASSERTION, meaning GitHub could not match the LICENSE file to a known template. Read that file directly before shipping the code in a product; nothing here is legal advice, and an unrecognised licence is a question for whoever owns the compliance decision at your organisation.
The last push to the default branch was on 2026-08-03, which is recent enough that the project is not abandoned, and the repository is not archived. There are no tagged releases, so there is no version number to pin and no changelog of breaking changes to read. The repository does carry a CHANGELOG.md and an UPGRADE_LOG.md, and those files are where you would look before pulling new commits.
Upgrade cost is dominated by the dependency list rather than the project's own code. `llama-cpp-python`, `transformers`, `tiktoken`, `opencv-python-headless` and `nvgpu` all move independently, and a local GGUF model has to stay compatible with whichever `llama-cpp-python` build is installed. Since there are no releases, an upgrade means tracking `main`, and the practical safeguard is the bundled example: rerun the Warren Buffett PDF after an upgrade and compare the generated markdown against the checked-in corrected file to see whether behaviour shifted.
Editorial conclusion
Adopt it if you have a scanned PDF where Tesseract is already close and you want markdown with fewer character-level errors, and you can accept a second model pass over every chunk. Skip it if your scans are low quality, if you need a deterministic byte-identical output, or if you have no GPU and no API budget, since the correction step is the whole point of the tool. Before trusting it on a document that matters, run the bundled Warren Buffett and Katharine Graham PDF and diff the generated _llm_corrected.md against the checked-in 160301289-Warren-Buffett-Katharine-Graham-Letter__raw_ocr_output.txt to see exactly where the model changed the text rather than fixed it.
Frequently asked questions
Does ChatGPT have OCR?
That is outside this project's scope; llm_aided_ocr does not use ChatGPT for OCR. It uses Tesseract for text extraction and calls an LLM, which can be OpenAI, Anthropic, or a local GGUF model, only to correct and format the text Tesseract already produced.
Which LLM is best for OCR correction in llm_aided_ocr?
The README does not rank models. It documents three paths: local inference through `llama_cpp` with a compatible GGUF model, plus OpenAI and Anthropic through their APIs, selected with `USE_LOCAL_LLM` and `API_PROVIDER` in the `.env` file.
What is OCR in AI models?
In this project OCR is not done by the model at all. Tesseract extracts the text through `pytesseract`, and the AI model sits downstream, correcting errors and optionally converting the result to markdown.
Which AI tool is best for OCR?
The README does not compare tools. llm_aided_ocr is a pipeline rather than a single tool: it combines Tesseract with an LLM correction stage and writes both `{base_name}__raw_ocr_output.txt` and `{base_name}_llm_corrected.md` so you can compare the two.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/dicklesworthstone-llm-aided-ocr)