franken_ocr: a pure-Rust, CPU-only OCR engine for five hand-ported VLMs
Pure-Rust, CPU-only OCR engine for Baidu Unlimited-OCR (a DeepSeek-OCR-derived 3B MoE VLM). Five-model zoo, custom int8 kernels, no ML framework, no Python, no GPU.
At a glance
- What is it?
- franken_ocr runs Baidu Unlimited-OCR and four other vision-language models on CPU from a single Rust binary, with custom int8 kernels instead of a ML framework. It targets the machines that actually need OCR: laptops, CI runners, agent hosts and edge boxes with no GPU.
- Who is it for?
- franken_ocr fits teams that need strong document OCR on CPU-only hosts: CI pipelines, agent hosts, laptops and edge boxes, where a single verified binary plus pulled weights beats a Python and CUDA dependency tree. It does not fit anyone who needs GPU throughput, model fine-tuning, or models outside the five ported families.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 1 day ago.
- What is it written in?
- Mainly Rust, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Five ported models, no Python, no GPU
The pitch is a direct answer to a distribution problem. Baidu Unlimited-OCR is, in the README's own summary, a strong document-parsing model covering Markdown, tables, LaTeX, reading order and many pages in one pass, but its official stack is Python plus CUDA. Most machines that need OCR, laptops, CI runners, agent hosts, edge boxes, have no usable GPU, and a Python plus CUDA dependency is heavy to ship and awkward to embed.
franken_ocr is a library plus a single-binary CLI named focr that runs a small family of hand-ported vision-language models on CPU with nothing but a Rust binary. The zoo has five ready engines with distinct jobs: Unlimited-OCR as the fast default for document OCR, GOT-OCR2 for specialized structured formats, SmolVLM2 for image description and VQA, OneChart for extracting chart data, and Polyphonic-TrOMR, which turns full scanned sheet-music pages or staff crops into MusicXML through the music task. Planned additions named in the README are TrOCR and pix2tex.
Current release binaries are about 13 to 17 MB, shipped raw for six targets: macOS Apple Silicon, macOS Intel, Linux x86-64, Linux ARM64, Windows x86-64 and Windows ARM64, each with a SHA256 sidecar. There is no Python, no CUDA, no FFI at inference and no GPU anywhere in the runtime story.
Hand-written kernels and runtime ISA dispatch
The engineering idea underneath is that these models are a fixed, known set, so they do not need a general framework. Checkpoints in bf16 are transformed into custom .focrq int8 artifacts, and inference runs through kernels written for each model's exact shapes. The v0.7.0 Unlimited-OCR artifact, for instance, is pinned by an embedded schema-v2 manifest at 4,157,448,783 bytes under the recipe unlimited-ocr-ffn-int8-attn-bf16-lmhead-bf16-v1, and focr pull verifies all three part hashes plus the reassembled file before installation.
Dispatch happens at runtime. One binary per architecture detects every available int8 ISA tier and reports the effective dense-GEMM route separately. On x86 it selects AVX-512-VNNI, AVX-VNNI, AVX2 or scalar; Apple Silicon currently defaults to a measured-faster LLVM autovec path while keeping forced SDOT and SMMLA proof paths. The default cached decode applies a conservative int8 FFN and expert cache regardless of source storage, and raw BF16 exists only as an all-high-precision native diagnostic paired with a stateless decode environment variable.
Two badges in the README say something about the codebase's posture: the toolchain is pinned to Rust nightly, and unsafe is forbidden. For a project doing SIMD-level work by hand, advertising the absence of unsafe blocks is a real constraint, not decoration.
Install, then pull the weights
The quick install is one script pipe:
curl -fsSL https://raw.githubusercontent.com/Dicklesworthstone/franken_ocr/main/install.sh | bashOn native Windows the equivalent is:
irm https://raw.githubusercontent.com/Dicklesworthstone/franken_ocr/main/install.ps1 | iexThe installer detects the platform, resolves the latest published GitHub binary release, currently v0.8.0, verifies the downloaded asset by SHA256, and puts focr on your PATH. What it does not install is the model weights, and that split is deliberate: weights arrive separately through the focr pull command from the model zoo, landing under ~/.cache/franken_ocr/models. Once compatible weights are present, inference never touches the network.
Two environment variables refine that default. Weight bytes are owned by default, meaning copied rather than shared, and trusted immutable deployments can opt into mmap instead. A stateless decode mode exists for the diagnostic path. The v0.7.0 default is the versioned, conservative exact-recipe Unlimited-OCR artifact, and TrOMR publishes both a 61 MB int8 default artifact and an 86 MB f32 reference artifact.
Documents, PDFs, figures and sheet music
As a library, the OcrEngine API exposes synchronous, blocking calls for Markdown output, structured layout, figure extraction, in-memory images and load-once batches. That synchronous shape is a fit for the CLI and embedding use cases the project targets, and it means there is no hidden runtime or event loop to ship alongside.
PDF handling is native: scanned PDFs are rasterized in process with pure Rust, page /Rotate and image-placement rotations are honored, and page selection picks exact PDF pages. A split-spreads option handles two-page book scans by separating them. Figure extraction saves chart and photo regions beside the Markdown or JSON output. For multi-page documents, a multi-page mode runs the Unlimited-OCR infer_multi contract over selected pages or an image list and produces one document with explicit page boundaries instead of unrelated per-page parses.
The odd one out is Polyphonic-TrOMR, which is optical music recognition rather than text OCR: full-page sheet music to MusicXML, including polyphony in the model's name, selectable via the music task. Few OCR engines bundle an OMR model at all, let alone a CPU-only one.
Measured evidence, and a claim the release refuses to make
The README is unusually specific about its own accuracy evidence, and it attributes everything. A historical real-page Unlimited-OCR run, recorded by the project, measured end-to-end character-error-rate 0.0094 and matched the reference decode to within a single token. The v0.7.0 corpus receipt records 20 of 20 pages within budget, with a still-high tail on one page where the exact-recipe artifact emits EOS early, at an aggregate normalized CER of 0.19307925.
More interesting is the claim the release declines to make. The README states plainly that it does not claim a strict three-party OpenPGP certificate, because the committed fail-closed finalizer requires three independently controlled registry-pinned signers and production audit receipts, and those are not available in this release process. Local corpus checks, model census, installer and binary checks are presented as evidence, not as a substitute for that governance claim.
This is the part reviewers rarely see: a project that separates measurement from ceremony, publishes the first, and refuses to fake the second. The measured numbers above come from the project's own records, not from an independent run, which is exactly the distinction the README itself keeps drawing.
Licence, cadence, and when to use the Python stack instead
The repository carries an MIT licence file and badge. Releases are frequent and recent: v0.9.0 was published on 2026-08-23, v0.8.0 on 2026-08-20, alongside a models-unlimited-wasm-v1 artifact, and the tree contains focr-wasm and focr-ios directories, which suggests browser and iOS targets beyond the six desktop binaries. The last push was on 2026-09-15.
The limits are structural. The model zoo is exactly the five ported families, so anything outside document OCR, structured formats, chart extraction, image QA and sheet music needs a different tool. The int8-first design is a conservative default, not a peak-throughput one. And a hand-ported kernel per model means each new model is real porting work, which is why the planned list is two entries long.
The alternative is the thing franken_ocr defines itself against: the upstream Python plus CUDA stacks. Running Unlimited-OCR through its own Python release gives you the full ecosystem, GPU throughput when you have one, and a path to fine-tuning, at the cost of a dependency tree that is miserable to embed and idle on CPU-only hosts. A second alternative for plain text OCR is Tesseract: battle-tested, everywhere, and much weaker on layout, tables and multi-page structure because it is not a vision-language model at all. Choose by host, not by model paper: no GPU and a binary to ship, focr; a CUDA box and training plans, Python.
Editorial conclusion
franken_ocr fits teams that need strong document OCR on CPU-only hosts: CI pipelines, agent hosts, laptops and edge boxes, where a single verified binary plus pulled weights beats a Python and CUDA dependency tree. It does not fit anyone who needs GPU throughput, model fine-tuning, or models outside the five ported families. Verify first: that focr pull completes and verifies all part hashes for your platform, that your accuracy bar survives the project's own published CER numbers rather than a demo, and that the missing three-party signing certificate matters or does not for your supply chain. The last push was on 2026-09-15.
Frequently asked questions
What is Unlimited-OCR?
Baidu Unlimited-OCR is a vision-language model for document OCR, handling Markdown, tables, LaTeX, reading order and many pages in one pass. It is a DeepSeek-OCR-derived 3B mixture-of-experts model, and franken_ocr is a pure-Rust CPU engine that runs hand-ported versions of it.
Does franken_ocr need a GPU?
No. It is CPU-only with no Python, no CUDA, no ML framework and no FFI at inference. Custom int8 kernels with per-architecture ISA dispatch do the work, selecting AVX-512-VNNI, AVX-VNNI, AVX2 or scalar on x86 and autovectorized or forced paths on Apple Silicon.
How are model weights installed?
Weights are installed separately from the binary through focr pull, landing under ~/.cache/franken_ocr/models. Once present, inference runs offline; the v0.7.0 default pins the versioned conservative exact-recipe Unlimited-OCR artifact.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/dicklesworthstone-franken-ocr)