TurboOCR: a TensorRT document parser for Linux and NVIDIA GPUs
TurboOCR, >200 img/s OmnidocBench. TensorRT FP16, PP-OCRv6, HTTP + gRPC
At a glance
- What is it?
- TurboOCR wraps PP-OCRv6 detection and recognition, layout, tables and formulas into one CUDA/TensorRT engine behind HTTP and gRPC. It is a strong fit if you already run NVIDIA hardware and want structured Markdown out of documents, and the wrong tool if you do not.
- Who is it for?
- Adopt TurboOCR if you run Linux with an NVIDIA Turing-or-newer GPU and want OCR plus layout, tables and formulas behind one HTTP or gRPC endpoint. Do not adopt it if your inference fleet is CPU-only, Apple Silicon or AMD, because the README states NVIDIA is the only backend shipped today and Metal, OpenVINO and ROCm are still in testing or development.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 3 days ago.
- What is it written in?
- Mainly C++, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What TurboOCR solves, and for whom
Most OCR stacks stop at text. You get boxes and strings, then you write your own layout heuristics, your own table reconstruction and your own reading-order pass if you want Markdown for a retrieval pipeline. TurboOCR bundles those stages into one process. The README describes it as a document parser rather than an OCR engine: PP-OCRv6 detection and recognition, layout via PP-DocLayoutV3 with 25 classes, class-aware XY-cut for reading order, SLANet+ tables rendered to HTML and PP-FormulaNet-S formulas rendered to LaTeX. Everything runs locally, with the README stating there is no VLM in the path.
The intended user is an infrastructure or ML platform engineer who already has NVIDIA GPUs and wants a document ingestion service that keeps up with a queue. The README frames the target as whole-page OCR at up to 559 images/s on receipts with one RTX 5090, and full structured parsing at roughly 20 pages/s, which it contrasts with VLM document parsers it says run near 1 page/s. Those are the project's own figures from its benchmark section, not independent measurements.
The secondary audience is anyone replacing a Python OCR service that has become the bottleneck in a RAG indexing job. TurboOCR is a compiled C++20 service with a stable HTTP surface, so the migration is mostly about deployment rather than about rewriting call sites.
The single-engine architecture behind the HTTP and gRPC endpoints
The core design decision is that all stages share one multi-stream CUDA/TensorRT engine rather than running as separate models behind separate services. The README states the whole pipeline runs on a single multi-stream engine. In practice that means detection, recognition and the optional layout, table and formula networks are all TensorRT engines loaded into one process, and requests are batched across CUDA streams.
On top of that sits a two-process serving layout. nginx listens on port 8000 and reverse-proxies to Drogon on port 8080, and the README says both start automatically inside the container. gRPC is exposed separately on port 50051. Prometheus metrics are exposed on /metrics.
Stages are opt-in per request even when the weights are loaded. Layout is on by default, and the README notes that requesting tables or formulas auto-enables layout, since a table region has to be found before it can be parsed. A GET /capabilities call reports which stages a running server actually loaded, which is the honest way to check a deployment rather than assuming an environment variable took effect.
PDF handling is native rather than bolted on. The README says pages are rendered and OCR'd in parallel, with optional page-image export and auto-rotation. That matters because a naive PDF path serialises rendering and inference, which would waste the throughput the engine is built for.
Installing TurboOCR with Docker and running a first request
The README's Quick Start assumes Docker with GPU access. Requirements are Linux, NVIDIA driver 595 or newer, and a Turing-or-newer GPU (RTX 20-series or GTX 16-series and up). Plan for roughly 4 GB of VRAM for text only and about 8 GB for the full pipeline. The command below starts the server with HTTP on 8000 and gRPC on 50051, and mounts a named volume so TensorRT engines survive restarts:
docker run --gpus all -p 8000:8000 -p 50051:50051 \
-v trt-cache:/home/ocr/.cache/turbo-ocr \
ghcr.io/aiptimizer/turboocr:latestFirst startup builds TensorRT engines from ONNX. The README puts that at about 90 seconds on a 5090 and up to an hour on older GPUs, and states that requests get a connection refused error from nginx until the backend is ready. Setting TRT_OPT_LEVEL=3 is documented as cutting build time 3 to 5 times with a small speed regression. Because the volume caches the engines, later starts skip the build.
Once the backend is up, a raw image POST returns text, confidence and a bounding box per region:
curl -X POST http://localhost:8000/ocr/raw \
--data-binary @document.png -H "Content-Type: image/png"The README shows the response shape as a results array, each entry carrying text, a confidence value and a four-point bounding box. To turn on the heavier stages, add the backend environment variables at container start and then ask for them per request. All weights are baked into the image, so no model paths are needed:
docker run --gpus all -p 8000:8000 -p 50051:50051 \
-e TABLE_BACKEND=slanext -e FORMULA_BACKEND=ppformulanet_s \
-v trt-cache:/home/ocr/.cache/turbo-ocr ghcr.io/aiptimizer/turboocr:latestThen combine the query flags. The README documents layout=1, tables=1 and formulas=1 as freely combinable, with tables and formulas pulling layout in automatically. PDF input goes to /ocr/pdf, PDF to Markdown to /ocr/pdf?markdown=1, and a single page to Markdown to /ocr/markdown. The OCR_MODEL variable selects tiny (the default), small, medium, or one of the retained script models: arabic, eslav, korean, thai, greek.
Where TurboOCR is the wrong tool
The largest constraint is hardware. The README states plainly that NVIDIA is the only backend shipped today, with Apple Metal and Intel OpenVINO in testing and AMD ROCm in development. If your inference fleet is CPU-only, or you run on Apple Silicon, or you have standardised on AMD, this project is not deployable for you right now. The README does not document a CPU fallback path.
VRAM is the second constraint, and it scales in a way that is easy to miss. The README says each additional PIPELINE_POOL_SIZE replica adds roughly another full set of the ~4 GB or ~8 GB working set, so a small card cannot simply be given more concurrency. On a 16 GB card running the full pipeline, that budget is tight after the first replica.
The accuracy picture is narrower than the throughput headline. The README describes the engine as accurate on forms and receipts and competitive with PaddleOCR-VL, PaddleOCR-Python, RapidOCR, EasyOCR and Tesseract, but the benchmark claims it leads with are receipt and form throughput. Dense documents are quoted at 200+ images/s rather than 559. If your corpus is handwritten, heavily degraded or outside the scripts listed, the README gives no evidence either way, and you should treat that as unmeasured rather than solved.
Operationally, the first-start engine build is a real failure mode rather than a footnote. On older GPUs it can take up to an hour, and during that window the documented behaviour is connection refused from nginx. Any orchestrator health check that treats a refused connection as a crash will restart the container in a loop. The README does not document a readiness endpoint separate from the backend itself.
How TurboOCR differs from running PaddleOCR yourself
The obvious alternative is PaddleOCR, which TurboOCR builds on: the README credits PP-OCRv6 detection and recognition, PP-DocLayoutV3, SLANet+ and PP-FormulaNet-S as the underlying models. The difference is what surrounds them. Running PaddleOCR directly means Python, PaddlePaddle and a set of separate model invocations that you schedule yourself. TurboOCR compiles those into TensorRT engines at FP16 and serves them from one process with batching across CUDA streams.
That trade is not free. You inherit a container that must build engines on first start, a CUDA and TensorRT version dependency, and a Linux-only deployment story. In exchange you get a fixed HTTP and gRPC contract and no Python runtime in the serving path. Teams that already have a Python inference platform with GPU scheduling may reasonably prefer to keep PaddleOCR and scale it horizontally, because their tooling already handles model loading and health checks.
The interesting middle ground is the model tiering. TurboOCR lets you pick tiny, small or medium from one image by setting OCR_MODEL, and retains PP-OCRv5 recognisers for Arabic, Cyrillic, Korean, Thai and Greek. That is a lighter way to cover multiple scripts than maintaining several PaddleOCR pipelines, provided the accuracy of the retained recognisers meets your bar.
Maintenance, build cost and the MIT licence
The repository is not archived, and the last push was on 2026-09-08, the same day as the v3.5.6 release. The release history shows a tight cadence in early September 2026: v3.5.3 on 2026-09-05, v3.5.5 on 2026-09-06 and v3.5.6 on 2026-09-08. The v3.5.5 notes mention aarch64 builds and OpenCV 5, and v3.5.3 mentions a single JPEG path, bounded memory and no silent fallbacks. Those notes suggest active bug-fixing rather than feature drift, but no roadmap is published, so long-term direction is not something you can plan against from the repository alone.
The upgrade cost is concentrated in two places. First, the v3.0 release introduced medium, small and tiny tiers and the README links a breaking-changes document at docs/build/upgrading-v3.md, so anyone on v2 has a migration to do. Second, TensorRT engine caches are tied to the build environment. A driver or TensorRT change can invalidate the cached engines in the trt-cache volume, which puts you back in the 90-second to one-hour first-start window. The README does not document a way to pre-build engines outside the container, so in a rolling deployment you should expect either a warm shared volume or a slow first pod.
On licensing: the repository is MIT, which is permissive and places few obligations on how you ship the service. That covers the TurboOCR code. It does not automatically answer questions about the upstream model weights or about NVIDIA's TensorRT redistribution terms, and the README does not address either. If you plan to redistribute a built image, check those separately rather than assuming the MIT badge settles it.
Editorial conclusion
Adopt TurboOCR if you run Linux with an NVIDIA Turing-or-newer GPU and want OCR plus layout, tables and formulas behind one HTTP or gRPC endpoint. Do not adopt it if your inference fleet is CPU-only, Apple Silicon or AMD, because the README states NVIDIA is the only backend shipped today and Metal, OpenVINO and ROCm are still in testing or development. Before committing, check GET /capabilities on a running container to confirm which stages actually loaded, and measure your own document mix rather than trusting the receipt and form numbers.
Frequently asked questions
Which OCR engine is better, PaddleOCR or RapidOCR?
TurboOCR's README places itself against both, describing its output as competitive with PaddleOCR-Python and RapidOCR on forms and receipts while claiming 15 to 90 times the speed of classic OCR engines on those document types. It does not publish a head-to-head accuracy table against RapidOCR, so treat the comparison as a throughput claim rather than an accuracy verdict.
What GPU and driver does TurboOCR require?
The README lists Linux, NVIDIA driver 595 or newer, and a Turing-or-newer GPU such as RTX 20-series or GTX 16-series and up. It states that NVIDIA is the only backend shipped today, with Metal and OpenVINO in testing and ROCm in development.
How much VRAM does TurboOCR need?
The README advises planning for about 4 GB of VRAM for text-only OCR and about 8 GB for the full pipeline with layout, tables and formulas. Each additional PIPELINE_POOL_SIZE replica adds roughly another full set, so lower it on smaller cards.
Why does TurboOCR return connection refused on first start?
The README states that first startup builds TensorRT engines from ONNX, which takes about 90 seconds on a 5090 and up to an hour on older GPUs, and that during the build requests return a connection refused error from nginx until the backend is ready. Setting TRT_OPT_LEVEL=3 is documented as cutting build time 3 to 5 times with a small speed regression.
How do I check which TurboOCR stages are loaded?
The README documents GET /capabilities as reporting which stages a running server has loaded. Stages are selected at container start with environment variables such as TABLE_BACKEND and FORMULA_BACKEND, and requested per call with query flags like layout=1, tables=1 and formulas=1.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/aiptimizer-turboocr)