olmocr: a GPU pipeline for turning PDFs into LLM training text
Toolkit for linearizing PDFs for LLM datasets/training
At a glance
- What is it?
- olmocr is Allen AI's Python toolkit for converting PDFs, PNGs and JPEGs into clean Markdown using a 7B vision language model. It is built for bulk dataset work, not for a laptop without a GPU.
- Who is it for?
- Adopt olmocr if you have a CUDA GPU, a corpus of scanned or image-based documents, and a downstream LLM pipeline that wants Markdown rather than a PDF reader. Do not adopt it for a one-off contract PDF on a laptop, for forms and receipts where field extraction matters more than reading order, or if you cannot accept the Apache-2.0 model weights being pulled from Hugging Face at runtime.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Activity is slowing. The repository last received commits 6 months ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 17, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem olmocr solves: PDFs that no parser can read
Most PDF text extraction fails on exactly the documents that matter for training data. A born-digital paper with embedded fonts extracts fine. A 1990s scan, a two-column journal page with a figure inset, or a table with merged cells does not. The text layer is either absent or wrong, and the reading order comes out scrambled because the extractor follows the PDF content stream rather than the visual layout.
olmocr targets that gap. The README describes a toolkit for converting PDFs and other image-based document formats into clean, readable, plain text, with features listed as equations, tables, handwriting, complex formatting, automatic header and footer removal, and natural reading order in multi-column layouts with insets. The audience is narrow and specific: people assembling LLM pretraining or fine-tuning corpora, or anyone who needs a large document collection as Markdown instead of as page images. The README quotes a cost of less than $200 USD per million pages converted, which only makes sense at that scale.
The project is not a general document AI product. It does not extract structured fields, it does not return bounding boxes as its primary output, and it does not try to preserve the original layout. It produces text in reading order, and it says so plainly: the model is a 7B parameter VLM, so it requires a GPU.
How the olmocr pipeline actually works
The pipeline is a rendering step feeding a vision language model, wrapped in a queue that retries bad pages. That shape is visible in the repository layout and the dependency list rather than in a prose architecture section.
Rendering is handled by pypdfium2 and pypdf, with poppler-utils installed at the system level and extra fonts installed so that documents render with the glyphs they expect. A page that renders with substituted fonts produces a different image, and the model reads the image, not the text layer. That is why the installation instructions spend as much space on fonts as on Python packages.
Inference runs through vLLM. The pyproject.toml declares the gpu extra as torch>=2.7.0, transformers==4.57.3 and vllm==0.11.2, and the Dockerfile builds from vllm/vllm-openai:v0.11.2. The changelog entry for v0.1.75 records the switch from sglang to a vllm based inference pipeline, so the serving layer has changed once already.
The output is Markdown. The dependency list includes markdown2, markdownify and bleach, which is consistent with a pipeline that produces Markdown and then cleans it. Header and footer removal is described as automatic, meaning the model is asked to drop repeated page furniture rather than a separate rule engine detecting it.
The queue behaviour is the part worth understanding before you commit. The release notes for v0.2.1 state that the newer model scores 3 points higher on olmOCR-Bench and needs much fewer retries per document. Retries are a first-class part of the design: a page whose output fails validation is re-run, and the retry count is a quality signal as much as a performance one. Throughput planning has to include that, not just raw pages per second.
Installing olmocr and running a first conversion
The README splits installation into system dependencies and a Python install. On Ubuntu or Debian, the system packages are listed explicitly, and they include poppler-utils plus a set of fonts used for rendering:
sudo apt-get update
sudo apt-get install poppler-utils ttf-mscorefonts-installer msttcorefonts fonts-crosextra-caladea fonts-crosextra-carlito gsfonts lcdf-typetoolsThe ttf-mscorefonts-installer package normally prompts for a EULA. The Dockerfile shows how the project handles that in automation, by pre-seeding the answer with debconf-set-selections and running with DEBIAN_FRONTEND=noninteractive. If you script the install, expect to do the same.
The Python package is named olmocr and requires Python 3.11 or newer. The pyproject.toml declares the GPU dependencies as a separate extra, with exact pins rather than ranges:
gpu = [
"torch>=2.7.0",
"transformers==4.57.3",
"vllm==0.11.2"
]That block is the whole of what the repository states about the GPU install; the README's installation section is where the pip command itself appears, and it is not reproduced here. The console entry point is olmocr, wired to olmocr.pipeline:cli_main in pyproject.toml, and the README's usage section is where the actual invocation lives. Read it in full before running anything, because the flags are not repeated here.
The project also ships Docker images. The changelog entry for v0.1.70 states that official Docker support and images are available, and the repository contains both a Dockerfile and a Dockerfile.with-model. The second one exists because the first pulls weights at runtime; if you are running in an environment with restricted egress, the variant that bakes the model into the image is the one you want.
Before a full run, the sensible first check is a single document. The repository ships a PDF at the top level, olmOCR-2-Unit-Test-Rewards-for-Document-OCR.pdf, which is a natural smoke test input. What you should see is Markdown in reading order with headers and footers stripped. If you get page furniture in the output or text out of order, the model version and the rendering fonts are the first two things to check.
Where olmocr is the wrong tool
The GPU requirement is the hard boundary, and the README states it without hedging: the pipeline is based on a 7B parameter VLM, so it requires a GPU. There is no documented CPU path. If your workload is a few hundred pages a month, the operational cost of a GPU host plus the pinned vLLM and transformers versions will exceed the value of the output.
The second limitation is that the model reads rendered images. That is what makes it good at scans and bad at documents where the text layer is already perfect and cheap to extract. For a born-digital PDF with a clean text layer, running a 7B VLM is a large amount of compute for a result that pypdf alone would give you. A sensible pipeline routes those documents elsewhere and reserves olmocr for the ones that need it, but the project does not ship that router.
The third is retries. The release notes describe retry counts as a headline improvement, which tells you that some pages fail validation and get re-run. A corpus with unusual scripts, dense tables or poor scan quality will sit at the high end of that distribution, and the throughput numbers you plan around will be optimistic.
Finally, output is Markdown, not structure. If your downstream task needs field-level extraction, key-value pairs, or coordinates, olmocr's reading-order text is the wrong shape. It is a linearization tool, and the repository description says exactly that.
olmocr compared with Marker, MinerU and Mistral OCR
The README publishes an olmOCR-Bench table, which makes comparison concrete rather than rhetorical. The benchmark covers over 7,000 test cases across 1,400 documents, and the table reports per-category scores.
Against Mistral OCR API, the difference is deployment model before it is accuracy. Mistral OCR is a hosted API; olmocr is a local pipeline you run on your own GPU. The README's table gives olmOCR v0.4.0 an overall of 82.4 against 72.0 for Mistral OCR API, with a much wider gap on old scans (47.7 against 29.3). If your documents are clean and modern, that gap narrows and the hosted API removes the GPU from your problem.
Against Marker 1.10.1, the overall scores are 82.4 and 76.1. Marker is stronger on ArXiv (83.8 against 83.0), and olmOCR is far ahead on old scans (47.7 against 33.5) and headers and footers (96.1 against 86.6). Both are open and runnable locally, so the choice comes down to document mix rather than deployment.
Against MinerU 2.5.4, the split is the reverse of what you might expect. MinerU scores 84.9 on tables, the same figure the table gives olmocr, and 96.6 on headers and footers against 96.1, while olmocr leads on old scans math (82.3 against 54.6). If your corpus is table-heavy and modern, MinerU's table handling is the reason to look at it. If it is historical scans, olmocr's margin is large.
Two caveats on the table itself. Several rows are marked with an asterisk, which the README does not explain in the excerpt available, and one row reports 82.5±? with no confidence interval. Treat the unstarred, interval-bearing rows as the comparable ones, and remember that these are the project's own benchmark results on its own suite.
Maintenance, licensing and the cost of upgrading
The repository is not archived, and the most recent push recorded is 2026-03-25. The latest release in the list is v0.4.27 from 2026-03-12, following v0.4.25 and v0.4.24 in January 2026. The release cadence through 2025 was frequent, with model releases in July, August and October, and the changelog records infrastructure changes such as the sglang to vllm switch and the move to CUDA 12.8.
That cadence is the upgrade cost. Model weights are versioned separately from the package, and the news entries pair each model release with a benchmark score. Upgrading the package without upgrading the model, or the reverse, is a configuration the project does not test. The pins in pyproject.toml are exact for the GPU extra (transformers==4.57.3, vllm==0.11.2), so a package upgrade can drag a vLLM upgrade with it, and vLLM upgrades are not small.
The licence is Apache-2.0, listed in pyproject.toml as an OSI-approved Apache license and shipped as a LICENSE file. Note the badge at the top of the README points at the OLMo repository's licence file rather than olmocr's own, which is a documentation inconsistency rather than a licensing statement. Model weights are distributed separately on Hugging Face, and the README links to olmOCR-2-7B-1025-FP8 among others. Whether those weights carry the same terms as the code is a question for the model card, not the repository, and it is worth checking before you build a product on the output.
Editorial conclusion
Adopt olmocr if you have a CUDA GPU, a corpus of scanned or image-based documents, and a downstream LLM pipeline that wants Markdown rather than a PDF reader. Do not adopt it for a one-off contract PDF on a laptop, for forms and receipts where field extraction matters more than reading order, or if you cannot accept the Apache-2.0 model weights being pulled from Hugging Face at runtime. Verify first that your GPU has enough memory for the 7B FP8 model, that poppler-utils and the font packages are present on the host, and that your expected page volume fits the cost figure the README quotes of under $200 USD per million pages.
Frequently asked questions
What is olmocr?
It is a toolkit from the Allen Institute for AI for converting PDFs, PNGs and JPEGs into clean Markdown, using a 7B parameter vision language model. The README lists support for equations, tables, handwriting and multi-column reading order, plus automatic removal of headers and footers.
How do I install olmocr?
Install the system packages first (poppler-utils plus the font packages listed in the README), then install the Python package with the GPU extra, which requires Python 3.11 or newer. Official Docker images are also available according to the changelog entry for v0.1.70.
How do I use olmocr?
The console entry point is olmocr, wired to olmocr.pipeline:cli_main, and it runs the vLLM-based inference pipeline over your documents to produce Markdown. The README's usage section carries the actual invocation and flags, which are not reproduced in the installation section.
Is olmocr free and open source?
The code is licensed Apache-2.0, as declared in pyproject.toml and shipped in the LICENSE file. The model weights are distributed separately on Hugging Face, so check the model card for the terms that apply to them.
What is olmOCR-Bench?
It is the benchmark suite shipped in the repository, covering over 7,000 test cases across 1,400 documents, with per-category scores for ArXiv, old scans, tables, headers and footers, multi-column layouts and long tiny text. The README uses it to compare olmOCR against Mistral OCR API, Marker, MinerU and others.
How does olmocr compare with Mistral OCR?
The main difference is deployment: Mistral OCR is a hosted API while olmocr runs locally on a GPU. In the README's olmOCR-Bench table, olmOCR v0.4.0 scores 82.4 overall against 72.0 for Mistral OCR API, with the widest gap on old scans.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/allenai-olmocr)