paperless-gpt: LLM OCR and automatic metadata for paperless-ngx
Use LLMs and LLM Vision (OCR) to handle paperless-ngx - Document Digitalization powered by AI
At a glance
- What is it?
- paperless-gpt is a Go service that sits next to paperless-ngx, runs LLM or cloud OCR over scans, and writes back titles, tags, correspondents and custom fields. It is useful if your scans defeat Tesseract, and unnecessary if they do not.
- Who is it for?
- Adopt paperless-gpt if your paperless-ngx archive holds scans that Tesseract mangles (faxes, handwriting, low-contrast photocopies) and you can point it at an OpenAI-compatible endpoint, Google Document AI, Azure Document Intelligence or a local Ollama model. Do not adopt it if your documents already OCR cleanly: you would be paying an LLM to re-read text you already have.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 3 days ago.
- What is it written in?
- Mainly Go, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem: paperless-ngx OCR fails exactly where you need it
paperless-ngx handles ingestion, storage and search for a document archive. Its OCR step is conventional, and conventional OCR degrades in predictable places: faxes, dot-matrix invoices, handwriting, stamps over text, skewed photocopies. When that happens you get a document in the archive with a title derived from a filename and a body of text that search cannot reach. The README states the project's position plainly, that it "stands out by supercharging OCR with LLMs".
The audience is narrow and specific. It is someone already running paperless-ngx, already comfortable with Docker Compose, and already annoyed by a subset of documents. paperless-gpt is not a replacement for paperless-ngx and does not try to be one. It is a second service that talks to the paperless-ngx API, reads documents, optionally re-OCRs them, and writes metadata back. If you are not a paperless-ngx user, the project has nothing to offer you.
How paperless-gpt works: a separate service around the paperless-ngx API
The repository layout tells most of the story. There is a Go binary (main.go, background.go, jobs.go), a Gin HTTP layer (app_http_handlers.go), an LLM client (llm_client.go, ollama_config.go, app_llm.go, app_llm_googleai.go), an OCR subsystem (ocr.go, ocr_prompt.go, ocr_runs.go, and an ocr/ directory), a SQLite persistence layer through GORM (local_db.go), and an embedded web UI (embedded_assets.go, with a web-app directory built by Vite in the Dockerfile).
The data flow is a loop. The service polls or is triggered against paperless-ngx, fetches a document, sends it to a configured OCR provider, sends the resulting text to a configured LLM with a prompt template, and writes suggestions back as title, tags, correspondent, created date and custom fields. A background worker handles the queue; the web UI is where you approve or correct suggestions, and where you edit prompts under Settings.
Two design choices are worth naming. First, prompts live in a default_prompts directory and a prompts directory, so your edits persist across container restarts rather than being overwritten by the image. Second, the OCR provider is pluggable: LLM-based OCR through OpenAI or Ollama, Google Document AI, Azure Document Intelligence, or a self-hosted Docling Server. That pluggability is the most consequential thing about the project, because it determines your cost, your privacy posture and your failure modes.
Installing paperless-gpt with Docker Compose and running a first document
The README's installation path is Docker Compose alongside paperless-ngx. The repository's docker-compose.yml defines a single app service built from the local Dockerfile, exposing port 8080 and reading a .env file.
services:
app:
build:
context: .
dockerfile: Dockerfile
ports:
- "8080:8080"
environment:
- PUID=10001
- PGID=10001
env_file:
- .envThe PUID and PGID values control the container user, and the compose file's own comment says to find yours with `id -u` and `id -g`. Getting these wrong is the usual cause of permission errors on mounted volumes. Your provider credentials and paperless-ngx connection details go in .env, which the README documents under Environment Variables; it does not reproduce the full variable list in the excerpt available here, so read that section before writing your file rather than guessing at key names.
Once the service is up, open the web UI on port 8080. The README describes a unified interface with a manual review mode and an auto-processing mode. A sensible first run is manual: let it process one or two documents, look at the suggested title, tags and correspondent, and correct them. Only after you have seen what your chosen model produces on your own scans should you switch to auto processing, where the README's framing is that you "focus only on edge cases".
If you would rather not use the container, the README also lists a Manual Setup path. The build is not trivial: the Dockerfile installs gcc, musl-dev, mupdf, mupdf-dev and sed on Alpine, pre-compiles go-sqlite3 with CGO, and builds a Vite frontend before the Go binary. A manual build reproduces all of that.
OCR providers and processing modes: where cost and privacy are decided
Four providers are documented. LLM-based OCR is the default and uses OpenAI or Ollama. Google Document AI and Azure Document Intelligence are the managed alternatives. Docling Server is the self-hosted option. The choice is not cosmetic. OpenAI sends your documents to a third party. Ollama keeps them local but makes you responsible for GPU capacity and model selection; the README suggests a reasoning model such as qwen3:8b as a balance between privacy and performance, and notes that larger models improve results if you have the hardware.
Processing modes add a second axis. Image Mode is the default and converts pages to images for the OCR provider. PDF Mode and Whole PDF Mode change what is submitted. The README includes a Provider Compatibility table, and this is the part to read carefully, because not every provider supports every mode. Picking a combination the table does not list is a configuration error you will discover at runtime.
There is also an Existing OCR Detection feature. It matters more than it sounds: if a document already carries usable text, re-OCRing it is wasted money and a chance to make the text worse. Where the README is thin is in quantifying any of this. It offers example images under demo/ (ocr-example1.jpg, ocr-example2.jpg) and a "Compare for Yourself" section, but no accuracy figures and no cost model. You are expected to evaluate on your own corpus, which is honest but means the decision is yours to make with your own samples.
Custom field write modes and the metadata copying limitation
Custom field generation is opt-in twice over: it must be enabled in settings, and you must select at least one target field. The README documents three write modes with clearly different risk profiles. Append only adds fields that do not already exist and, in the README's words, "will never overwrite an existing field, even if it's empty". Update adds new fields and overwrites existing ones where a suggestion exists, leaving untouched fields alone. Replace deletes all existing custom fields on the document and substitutes the suggestions.
Replace is the one to think about. It is a destructive operation driven by model output. If the model returns an incomplete set, you have lost data that paperless-ngx held. Append is the conservative default and the right starting point.
The README also has a section titled Metadata Copying Limitations, which is a candid admission that the enhanced-PDF path does not carry everything across. Combined with the safety features and usage recommendations sections, the picture is of a tool whose authors know the write-back step is the dangerous part. That is a better sign than silence, but it does mean you should read those sections before enabling anything automatic.
When paperless-gpt is the wrong tool
The clearest case against it is a clean archive. If your documents OCR well under paperless-ngx's default pipeline, adding an LLM OCR pass introduces cost, latency and a new failure surface for no accuracy gain. The README's own framing is about "messy or low-quality scans"; documents that are not messy do not need this.
Second, the project is a metadata writer. Every feature that generates a title, a tag, a correspondent or a custom field is a write against your archive. In auto-processing mode, that happens without a human in the loop. A model that systematically mislabels a correspondent will do so across every document in a batch before you notice. The review UI mitigates this only if you actually use it.
Third, the deployment is a second stateful service. It has its own SQLite database (local_db.go, via gorm.io/driver/sqlite), its own background job queue, and its own configuration surface of environment variables and prompt templates. That is more to run, back up and upgrade than a paperless-ngx-only setup. There is no documented rollback path for metadata that has already been written back, and the README does not describe one.
Finally, the language model is doing judgement, not transcription. Titles and tags are opinions. If you need deterministic, auditable classification, a rules-based approach in paperless-ngx itself will be more predictable than a prompt.
paperless-gpt compared with paperless-ai and paperless-ngx alone
paperless-ngx on its own is the baseline: conventional OCR, manual or rules-based metadata, no external model calls. It is the right answer for most archives.
paperless-ai is the comparison people actually search for. Both projects attach a model to paperless-ngx, and the difference in approach is where each puts its emphasis. paperless-gpt's README positions the project around OCR enhancement first, with the LLM improving the extracted text itself, and metadata generation second. The provider list reflects that: Google Document AI, Azure Document Intelligence and Docling Server are all document-text services, not chat models. It also generates searchable PDFs with a transparent text layer positioned over each word, preserving the original appearance, which is an OCR-output feature rather than a classification one. If your problem is that the text in your archive is wrong, that emphasis is the reason to pick this project.
The honest caveat is that this comparison is drawn from paperless-gpt's own documentation. The README does not contain a feature-by-feature comparison with paperless-ai, and nothing here should be read as a verdict on the other project's capabilities. Evaluate both against your own worst documents.
Maintenance, licence and what upgrades cost you
The repository is not archived, and the last push was on 2026-09-09, eight days before this writing. Releases are frequent and versioned with names: v0.27.0 ("The Clarity Expansion") on 2026-07-21, v0.26.1 on 2026-07-08, v0.26.0 ("The Resilience Expansion") on 2026-07-07. The cadence suggests a project that is still moving, and the release naming suggests the maintainer treats behavioural changes as events worth labelling.
That cadence is also the upgrade cost. A tool that writes into your archive and changes its prompt handling between minor versions is not something you update unattended. The default_prompts and prompts directory split exists so your customisations survive upgrades, but the README does not promise that the template variables those prompts use are stable across releases. Check the Template Variables section after each upgrade before trusting your prompts.
The licence is MIT, which permits commercial use, modification and redistribution provided the copyright notice and permission notice are included. That is permissive and imposes no copyleft obligation on your own code. It says nothing about the terms of the OCR and LLM providers you configure, and those terms are where the real constraints live: sending documents to a hosted OCR service or an LLM API is a data-processing decision with its own legal weight. The MIT licence on this repository does not address that at all, and nothing here is legal advice.
Editorial conclusion
Adopt paperless-gpt if your paperless-ngx archive holds scans that Tesseract mangles (faxes, handwriting, low-contrast photocopies) and you can point it at an OpenAI-compatible endpoint, Google Document AI, Azure Document Intelligence or a local Ollama model. Do not adopt it if your documents already OCR cleanly: you would be paying an LLM to re-read text you already have. Do not adopt it either if you cannot accept that a language model decides your document titles and tags, because the review UI is a checkpoint, not a gate. Before deploying, verify three things in your own environment: which OCR provider and mode you will run, whether custom field writing is left at the default Append mode, and whether the container can reach your paperless-ngx instance on the network you intend.
Frequently asked questions
What is paperless-gpt?
It is a Go service that pairs with paperless-ngx to generate AI-powered document titles and tags, and to improve OCR using LLMs or specialised OCR providers. It runs as a separate container with a web UI on port 8080.
How do I install paperless-gpt?
The README documents a Docker Compose deployment alongside paperless-ngx, using the repository's docker-compose.yml, which exposes port 8080 and reads a .env file, plus a Manual Setup path. The compose file sets PUID and PGID to control the container user.
How do I use paperless-gpt?
You open the web UI, configure prompts and an OCR provider, and let it process documents either in manual review mode, where you approve or tweak suggestions, or in auto processing mode, where the README says you focus only on edge cases. Ad-hoc analysis of a document selection with a custom prompt is also available.
Which is better, paperless-ai or paperless-gpt?
The README does not compare the two projects, so no verdict can be drawn from it. paperless-gpt's documented emphasis is OCR enhancement through LLM or specialised providers such as Google Document AI, Azure Document Intelligence and Docling Server, with metadata generation alongside it.
Why is paperless-gpt not processing my documents?
The README does not document this failure mode directly. Its troubleshooting-relevant material is the Provider Compatibility table, which shows that not every OCR provider supports every processing mode, and the note that custom field generation requires the feature to be enabled in settings and at least one field to be selected.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/icereed-paperless-gpt)