# vorojar/Folio-OCR: local OCR that ships with the GPU path commented out

> A batch OCR workbench that turns scanned PDFs, book photos and exam papers into editable text using Ollama with GLM-OCR for recognition and PP-DocLayoutV3 for layout, exporting to Markdown, text, Word and EPUB. The GPU path ships disabled, the documented fix for a timeout on current Ollama is to roll back eleven versions, and the default listen address is every interface.

**vorojar/Folio-OCR** — Open-source batch OCR workbench — a free, local alternative to ABBYY FineReader. Powered by Ollama + GLM-OCR + PP-DocLayoutV3, ~0.5s/page on RTX 4090. Three-panel editor, layout-aware, PDF/image batch processing, Markdown/Word export. 批量OCR工作台，纯本地运行，免费平替ABBYY，适合书籍文档数字化。

- Repository: https://github.com/vorojar/Folio-OCR
- Website: https://vorojar.github.io/Folio-OCR/
- Stars: 473 · Forks: 62
- Language: Python
- License: MIT
- Published: 2026-09-15 · Updated: 2026-09-15 · Language: en
- Canonical page: https://hysenlabs.com/projects/vorojar-folio-ocr

## The GPU path is present in the compose file and commented out

The stack runs as two services. Ollama does the recognition, holding port 11434 and persisting its data to a named volume. The application holds port 3000, talks to Ollama over the internal network address, and keeps uploads, its SQLite file and the Hugging Face model cache in three more named volumes. The GPU reservation block is there in the compose file with every line commented, preceded by a note to uncomment it if you have an NVIDIA card. So acceleration is a documented manual edit rather than something the profile detects. That default is deliberate rather than an oversight: the configuration explains that the layout analysis model is kept on the CPU specifically to leave GPU memory for Ollama, and the device setting accepts three values, where the automatic mode moves layout to the GPU only when CUDA is present and the forced CUDA mode fails at startup rather than falling back.

## The fix for a current-Ollama timeout is to roll back eleven versions

The troubleshooting section names a specific failure: on Ollama v0.30.6, a PDF of roughly one megabyte times out. It first tells you to determine whether the browser request or the Ollama runner is stuck, then notes the application already raised its frontend OCR timeout to 300 seconds and offers a larger value for slow machines. The concrete remedy offered to Docker users, though, is not a tuning change but a version change, setting the Ollama image to 0.19.0 before bringing the profile up and re-pulling the model. That is a rollback of eleven minor releases. It sits awkwardly against the compose file, where the Ollama image is specified with a default of `latest`, so an unattended reinstall moves the backend version without anyone choosing to. Pinning that tag yourself is the single most useful thing you can do before the first run.

## None of the eleven endpoints checks who is asking

The API is a FastAPI surface with eleven routes and no authentication described anywhere. It exposes service status and Ollama connectivity, a route that starts Ollama and warms the model, an upload endpoint returning a streamed page flow, a route that serves page images by document and filename, single-page OCR, document listing and detail, deletion of a whole document, saving edited page text, and two export routes. The delete route is the one to think about, since a document id is the only thing standing between a caller and removing a document along with its pages and OCR results. That matters because the configuration table gives the listen address a default of all interfaces, while the documentation repeatedly tells you to open the service at localhost on port 3000. Those two defaults disagree, and the loopback assumption is the safer one.

## Four export formats advertised, two of them have no endpoint

The feature list names four export targets: Markdown, plain text, Word documents and EPUB. The API table lists export routes for exactly two, one producing DOCX and one producing EPUB. There is no server route for Markdown or plain text, so those two are produced somewhere else, most plausibly in the browser from the text already being edited in the right-hand panel. The Word path is the more substantial of the two that exist, since the documentation specifies it is produced with python-docx and is a real Word document containing section breaks and page numbers rather than renamed HTML. The EPUB path has no supporting library in the dependency file, which lists only the API framework, the multipart and HTTP libraries, the PDF renderer, imaging, the machine learning stack and the Word library. So the ebook is assembled without a dedicated package.

## Installing the package pulls in torch and transformers anyway

The project metadata declares its dependencies as dynamic and reads them from the requirements file, so what you install from a package index is exactly that file. It contains the web framework and server, a multipart parser, an async HTTP client, the PDF library, image handling, and then three heavy entries: transformers, torch and torchvision. Those three exist for the layout analysis model, not for the OCR, because recognition happens inside Ollama over HTTP. Anyone who installs this expecting a lightweight local service gets a deep learning stack downloaded and imported regardless. The requirements file also opens with a single comment labelling the whole list as Ollama backend dependencies, which misdescribes the torch and transformers entries, since those drive the layout partitioning that runs locally on the machine.

## The runtime image keeps its compiler and runs as root

The Docker image is a single stage built on the slim Python 3.11 base. It installs a compiler toolchain with apt, namely gcc, g++, and the library headers for FFI, then copies the requirements file, installs from it, copies the rest of the tree, exposes port 3000 and starts the server module. Three things follow from that shape. The build tools installed to compile wheels stay in the final image rather than being discarded in a separate stage, which makes it larger than it needs to be. There is no instruction creating an unprivileged user, so the process runs as root inside the container. And the copy of the source happens after the dependency install, which is the right order for caching, though the whole tree comes in including anything the ignore file does not exclude.

## The speed claim comes from merging adjacent regions

The most concrete performance number in the feature list is about layout, not about the model. Adjacent text regions are merged intelligently before recognition, which the documentation illustrates as eleven regions collapsing into three groups for a 2.5 times speedup and fewer calls to the model. That is the lever the design pulls before touching anything else, and it is also why the interface offers bidirectional highlighting between layout regions, since clicking an image box and clicking a text block select each other. Other output-side cleanups are named as well: converting LaTeX special constructs into their Unicode equivalents, stripping markdown code fences out of model output, and a reflow step that rejoins paragraphs broken across line wraps. The recognition figures themselves are a roughly 50 second cold start on the first request and roughly half a second per page afterwards, with PDF pages rendered at double scale to protect quality.

## The context default exists because of a specific assertion

One configuration value has an unusually documented origin. The troubleshooting section names a failure in which Ollama reports an assertion about array dimensions during image OCR, and states that from v3.3.1 onward the application sends a context window of 16384 on every chat request to avoid triggering it:

```json
{
  "options": {
    "num_ctx": 16384
  }
}
```

If the assertion still appears, the advice is to raise the context further or lower the input image resolution. A second variable caps the output length of a single recognition at 4096 tokens by default, for short documents or low-memory machines. On batch size the position is explicit that no page limit is hardcoded: pages are written in order to the database and the uploads directory, recognition runs page by page, and the real ceiling is disk space, browser page count and backend stability, with 20 to 50 pages per batch suggested until the environment is known to be steady. The README itself is written in Simplified Chinese with no English version listed alongside it.

## Conclusion

This suits digitising a personal book or exam-paper collection on a machine with a local Ollama, where sending scans to a third-party service is not acceptable. Pin the Ollama image tag before your first run, read the compose file to turn on GPU support, and keep the listen address on loopback unless you have decided how to expose a document store with a delete endpoint.

## FAQ

### What models does Folio-OCR use and what does it accept?

PP-DocLayoutV3 performs layout partitioning and GLM-OCR running through Ollama performs the visual recognition. Input is PDF, PNG, JPG, GIF and BMP, with images and PDF pages mixable in one upload, and output is Markdown, plain text, Word DOCX and EPUB.

### How do I run Folio-OCR with uvx?

Pull the model with `ollama pull glm-ocr`, then run `uvx --from git+https://github.com/vorojar/Folio-OCR folio-ocr` and open http://localhost:3000. A pipx equivalent is given using `pipx run --spec`. This route assumes Python and Ollama are already installed and avoids cloning the repository.

### How do I enable GPU acceleration in Folio-OCR?

Edit the compose file and uncomment the reserved devices block under the Ollama service. By default the layout analysis model runs on the CPU so that GPU memory is left for Ollama, and the device setting also accepts an automatic mode that uses CUDA when present or a forced CUDA mode that fails at startup if it is unavailable.

### Why does Ollama report a GGML assertion failure during image OCR?

It is a context error in the backend triggered while processing images. From v3.3.1 the application sends a context window of 16384 on every chat request to avoid it. If it still occurs, raise the context with a larger OLLAMA_NUM_CTX or lower the input image resolution.

### How many pages can Folio-OCR process in one batch?

No page limit is hardcoded. Uploaded images and PDF pages are written in order to SQLite and the uploads directory and recognised page by page, so the practical ceiling depends on disk space, how many pages the browser will hold and backend stability. Batches of 20 to 50 pages are suggested until the environment is known to be steady.

### What address does Folio-OCR listen on by default?

The configuration table gives the listen address a default of all interfaces with port 3000, while the documentation tells you to open the service at localhost:3000. The eleven API routes, including document listing and document deletion, carry no authentication, so keeping the bind address on loopback is the safer assumption.

## Sources

- [License: MIT](https://github.com/vorojar/Folio-OCR/blob/main/LICENSE)
- [Project website](https://vorojar.github.io/Folio-OCR/)
- [README](https://github.com/vorojar/Folio-OCR/blob/main/README.md)
- [Releases](https://github.com/vorojar/Folio-OCR/releases)
- [vorojar/Folio-OCR on GitHub](https://github.com/vorojar/Folio-OCR)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/vorojar-folio-ocr
