# ocrs: A Rust OCR Engine That Skips the Preprocessing

> ocrs is a Rust library and CLI for extracting text from images, built on PyTorch models exported to ONNX and run through RTen. It is an early preview with Latin-alphabet support only, and it is aimed at developers who want OCR embedded in a Rust or WebAssembly program rather than a service.

**robertknight/ocrs** — Rust library and CLI tool for OCR (extracting text from images)

- Repository: https://github.com/robertknight/ocrs
- Stars: 1,892 · Forks: 91
- Language: Rust
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/robertknight-ocrs

## What ocrs solves, and who it is for

Most OCR pipelines inherited from the Tesseract era assume you will clean the image first: deskew it, binarize it, crop it, tune a page segmentation mode. The ocrs README states the goal directly: a modern engine that works on scanned documents, photos containing text and screenshots with "zero or much less preprocessing effort compared to earlier engines like Tesseract", achieved by using machine learning more extensively in the pipeline. That is the pitch. If you have ever written a chain of ImageMagick calls before feeding an image to an OCR binary, that chain is what ocrs is trying to make unnecessary.

The second audience is Rust developers. The workspace layout makes the intent clear: the repository holds ocrs (the library), ocrs-cli (the command line tool), ocrs-capi (a C ABI with a generated header), ocrs-extension (a browser extension), and a js directory built for WebAssembly. This is not a Python service you deploy next to your application. It is a crate you link, or a WASM module you ship to a browser. The README lists "easy to compile and run across a variety of platforms, including WebAssembly" as an explicit goal, and the justfile has a wasm target that builds the library for wasm32-unknown-unknown with SIMD enabled and runs wasm-bindgen into js/dist.

Two constraints come with that positioning and both are stated plainly. The README says ocrs is "currently in an early preview" and to "expect more errors than commercial OCR engines". It also says ocrs "currently recognizes the Latin alphabet only (eg. English)", with more languages planned and tracked in issue 8. Anyone evaluating this for production should treat those as the boundaries of the tool, not as caveats to work around.

## The pipeline: PyTorch models, ONNX export, RTen at runtime

The mechanism is a two-model neural pipeline. The README says the library uses "neural network models trained in PyTorch, which are then exported to ONNX and executed using the RTen engine". The workspace Cargo.toml pins the runtime dependencies: rten, rten-imageproc and rten-tensor, all at version 0.26.0, with rten declared default-features = false. RTen is a separate project by the same author, so the inference engine and the OCR engine are maintained in tandem rather than ocrs wrapping an external runtime like ONNX Runtime or tract.

The CLI exposes the pipeline's two stages through its output flags. Running ocrs on an image produces plain text. Adding --json produces text plus layout information, which is the detection stage's output: where words and lines were found. Adding --png produces an annotated image showing "the location of detected words and lines", per the README's example list. Those three outputs map onto the same underlying data, so a caller who needs coordinates should use --json rather than trying to recover positions from the text.

The models themselves live elsewhere. The README points to the ocrs-models repository for details about the models and datasets, for tools to train custom models, and notes that the models are also available in ONNX format for use with other runtimes. That is a meaningful detail: the ONNX files are portable artifacts, so the detection and recognition networks could in principle be run by a different engine. The justfile confirms how the tests get them, downloading into ocrs/examples and passing paths through the OCRS_DETECTION_MODEL and OCRS_RECOGNITION_MODEL environment variables when running the ocrs-capi integration test. The repository does not commit the model weights, which is why the first CLI run downloads them.

## Installing the ocrs CLI and running your first image

The README's install path assumes Rust and Cargo are already present. The CLI is published as a separate crate, ocrs-cli, not ocrs, so the install command names the CLI package and passes --locked to build against the committed dependency versions.

```bash
cargo install ocrs-cli --locked
```

If you want to read images straight off the system clipboard, the README documents an optional feature flag. This is the only build-time option the README mentions for the CLI.

```bash
cargo install ocrs-cli --locked --features clipboard
```

Then point it at an image. The README's first example is a single positional argument.

```bash
ocrs image.png
```

What you should see: extracted text on stdout. The first run behaves differently from later runs, because the README states that when the tool is run for the first time it downloads the required models automatically and stores them in ~/.cache/ocrs. If you are building in a sandbox with no network access, that first run will fail and you will need to arrange for the cache to be populated ahead of time.

Two more invocations cover the common follow-ups. Writing to a file uses -o, and requesting structured output adds --json.

```bash
ocrs image.png -o content.txt
ocrs image.png --json -o content.json
```

With the clipboard feature installed, the input can come from the clipboard instead of a path, using either the long or short form.

```bash
ocrs --clipboard
ocrs -c
```

If you would rather build from source, the README gives the clone-and-run sequence, which builds the CLI in release mode through the workspace.

```bash
git clone https://github.com/robertknight/ocrs.git
cd ocrs
cargo run -p ocrs-cli -r -- image.png
```

## Latin script only, and what that costs you

The single hardest limitation is language coverage. The README states that ocrs "currently recognizes the Latin alphabet only (eg. English)" and links to issue 8 for planned support. For an English-language screenshot pipeline this is fine. For a document archive containing Cyrillic, Greek, Arabic, Hebrew, Devanagari or any CJK script, ocrs cannot do the job at all, and no configuration flag changes that. The limitation is in the trained model, not in a setting you forgot to enable.

The second limitation is accuracy, and it is self-declared. The status section says ocrs is in an early preview and to expect more errors than commercial OCR engines. That is an unusual thing for a README to say, and it should be read literally. There is no published accuracy table in the README; it defers model evaluation to the ocrs-models repository. If you need a number before adopting, that is where to look, and if the number is not there, you do not have one.

The third is operational. Models are downloaded on first use into ~/.cache/ocrs rather than bundled. That is convenient for a developer laptop and awkward for hermetic builds, air-gapped deployments, or CI runners where the cache is cold on every job. The justfile works around this for its own tests by downloading models into ocrs/examples and pointing OCRS_DETECTION_MODEL and OCRS_RECOGNITION_MODEL at the resulting files, which is the pattern to copy if you need deterministic builds. Note also that the README does not document an offline install step, a model checksum, or a way to pin a model version, so a cache populated today may not be the cache you expect after an upstream model update.

## How ocrs differs from Tesseract and from LLM-based OCR

Tesseract is the obvious comparison and the README makes it explicitly, framing ocrs as needing "zero or much less preprocessing effort" than earlier engines. The difference in approach is where the machine learning sits. Tesseract's classical pipeline leans on hand-tuned image processing for layout analysis and character segmentation, which is why practitioners accumulate preprocessing recipes. ocrs moves more of that work into learned models, so the input tolerance is broader by design. The trade-off is the one already noted: Tesseract supports well over a hundred languages and has been deployed for decades, while ocrs supports Latin script and describes itself as an early preview. If your text is English and your images are messy, ocrs is the more modern bet. If your text is not English, Tesseract is the only one of the two that will run.

LLM-based OCR is a different shape of tool. A vision-language model handles arbitrary scripts and unusual layouts without a training step, but it runs as a large model behind an API or a GPU, and per-page cost and latency scale with model size rather than with the number of text lines on the page. ocrs runs a small ONNX model through RTen on the CPU, which is why the justfile can build it for wasm32-unknown-unknown and ship it to a browser. The two are not substitutes: one is a local library with a narrow script range, the other is a hosted or GPU-bound model with a broad one.

For Rust projects specifically, the alternative is usually a binding to a C or C++ engine, which reintroduces a native build dependency and a cross-compilation problem. ocrs is pure Rust in the workspace, which is what makes the WASM target and the ocrs-capi header generation straightforward. That is the concrete difference, and for a browser extension or an embedded tool it may matter more than accuracy.

## Maintenance, licence and what an upgrade actually costs

The repository is not archived and the last push was on 2026-09-02, so development is recent. There are no retrieved releases, which means there is no tagged version history to reason about here; the CHANGELOG.md file exists at the top level but its contents are not reproduced in the README, so the practical upgrade path is to track the main branch or the published crate versions. The workspace pins rten, rten-imageproc and rten-tensor at 0.26.0, and since ocrs and RTen share an author and are versioned together, a bump in the runtime is likely to arrive alongside library changes rather than independently. Budget for that coupling.

Upgrades carry a second cost that is easy to miss: the models. Because the CLI downloads models at first run and the README does not document a pinned model version or a checksum, an upgrade can change recognition behaviour without any change to your code. If output stability matters, mirror the model files yourself and pass explicit paths, the way the justfile does with OCRS_DETECTION_MODEL and OCRS_RECOGNITION_MODEL.

On licensing, the repository carries LICENSE-APACHE.txt and LICENSE-MIT.txt at the top level while the stated project licence is Apache-2.0. The models and datasets are described in the README as "open and liberally licensed", with the specifics living in the ocrs-models repository. That distinction matters if you redistribute: the code licence and the model licence are separate questions, and the answer to the second one is not in this repository. There is also an AI_POLICY.md at the top level, which is worth reading before contributing. None of this is legal advice; check the model repository's terms yourself if you plan to ship the weights.

## Conclusion

Adopt ocrs if you are writing Rust or WebAssembly and want an OCR engine you can read, modify and compile into your own binary, and if your images are Latin script. Do not adopt it if you need Cyrillic, CJK or Arabic text, or if you need accuracy on par with a commercial engine today, since the README states it is an early preview and expects more errors. Before committing, run `ocrs` on a sample of your own images and compare the `--json` output against what your downstream code expects, and check the model download step works on your build machine, because the models are fetched at first run into `~/.cache/ocrs` and the repository does not ship them.

## FAQ

### What is OCR and how does it work in ocrs?

OCR is optical character recognition, extracting text from images. In ocrs, the README states that neural network models are trained in PyTorch, exported to ONNX, and executed with the RTen engine, producing text plus optional layout information.

### What is the best OCR right now?

The README does not rank OCR engines, and ocrs describes itself as an early preview that expects more errors than commercial OCR engines. It positions itself as needing less preprocessing than earlier engines such as Tesseract, and it currently recognizes the Latin alphabet only.

### Which OCR engine is best?

That depends on the script and the deployment target, and the README only covers ocrs directly. ocrs is a Rust library and CLI that compiles to WebAssembly, which suits embedding in a Rust or browser application, but it supports Latin script only and the README states more languages are planned.

## Sources

- [Issues](https://github.com/robertknight/ocrs/issues)
- [License: Apache-2.0](https://github.com/robertknight/ocrs/blob/main/LICENSE)
- [README](https://github.com/robertknight/ocrs/blob/main/README.md)
- [robertknight/ocrs on GitHub](https://github.com/robertknight/ocrs)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/robertknight-ocrs
