Self-hosted service
eikek/docspell avatar
eikek/docspell

Docspell: Self-Hosted Document Management with OCR and ML Tagging

Assist in organizing your piles of documents, resulting from scanners, e-mails and other sources with miminal effort.

2,336 stars185 forksElmAGPL-3.0

At a glance

What is it?
Docspell is a self-hosted document management system designed for households and small organizations that deal with scanned papers, email attachments, and other document sources. It combines OCR, fulltext search, and Stanford NLP-backed machine learning to automate metadata tagging, and it exposes everything through a REST API with an Elm-based web frontend.
Who is it for?
Docspell fits households, families, and small organizations that want a self-hosted DMS with automated OCR, NLP-based metadata suggestions, and REST API access. It is a poor fit for larger organizations that need audit trails, document versioning, or enterprise access control, or for teams that are not comfortable running a multi-service deployment with external OCR and NLP dependencies.
Can I use it commercially?
Yes, with strict conditions. AGPL-3.0 is a network copyleft licence: if people use a modified version over a network, for example as a hosted service, you must offer them its source code under the same licence.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Elm, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Docspell organizes and who runs it

Docspell describes itself as a personal document organizer, targeted at home use: families, households, and smaller groups or companies. The core workflow is upload, OCR if needed, fulltext index, and metadata tagging. Documents arrive from scanners, email integrations, the Android client app, or direct upload through the web interface.

The metadata model includes tags, correspondents (the person or organization the document relates to), and custom user-defined fields. Finding a document later means searching against those fields and against the full extracted text. The system learns from existing documents using the Stanford NLP library to suggest correspondents and tags for new documents, reducing the manual effort of classification.

Docspell is not designed for the kind of document management that enterprises need: it does not describe features for compliance workflows, document versioning, approval chains, or audit logs. The problem it solves is specifically the pile of scanned papers and email PDFs that accumulates in a household or small office with no systematic organization.

How OCR and NLP work together in Docspell

OCR runs automatically when a document is uploaded and is not already a searchable PDF. Docspell uses Tesseract as the OCR engine and relies on ocrmypdf as an additional processing tool. For document format conversion, it uses unoserver and unoconvert, which are LibreOffice-based tools for converting office formats to PDF before OCR.

Once OCR has run, the text is indexed for fulltext search. This is the foundation that makes Docspell useful as a retrieval system rather than just a filing system.

The machine learning features sit on top of the indexed text. Docspell uses the Stanford CoreNLP library to extract named entities and apply classifiers trained on existing documents. When you upload a new document, the system can compare its content against your collection and suggest the correspondent it likely belongs to or the tags that other similar documents carry. The README notes that this relies on the free GPL-licensed Stanford CoreNLP library, not a hosted AI service.

This is an on-premise ML pipeline. The quality of suggestions depends directly on how consistently you have tagged past documents. A new installation with few tagged documents will produce weaker suggestions than one with years of classified documents.

Installing Docspell via Docker Compose

The README describes Docker Compose as the quickest way to get started. The docker setup is in a separate repository at github.com/docspell/docker:

shell
git clone https://github.com/docspell/docker docspell-docker
cd docspell-docker/docker-compose
docker-compose up -d

After the services start, the web interface is at http://localhost:7880. You sign up and log in using the same name for the collective and user account on first setup.

Alternative installation methods documented in the repository include:

- Installing a .deb package on Debian-based systems from the GitHub releases page - Downloading a zip archive and running the script in bin/ - Using the Nix package manager; a NixOS module is available - Deploying with the Helm chart from the docker repository for Kubernetes

All methods require the external OCR and conversion tools (tesseract, unoserver, ocrmypdf) to be present in the execution environment. The Docker Compose setup handles this by including them in the service containers. Manual installations require installing those tools separately before Docspell can process documents that need OCR or format conversion.

Technical stack: Scala backend, Elm frontend

The backend is written in Scala in a purely functional style. The README lists the core libraries from the typelevel ecosystem: Cats, FS2, Doobie, Http4s, Circe, and Pureconfig. The build tool is sbt, visible in the build.sbt file at the repository root.

The web frontend is a single-page application written in Elm, using Tailwind CSS as the UI framework. Elm is a functional language that compiles to JavaScript; it produces no runtime exceptions at the JavaScript level, which is relevant for a file-management application where partially rendered UI states would be disruptive.

The REST/HTTP API covers all operations in the application. The Android client, maintained in a separate repository at github.com/docspell/android-client, and the CLI tool at github.com/docspell/dsc both use this API. Custom integrations or scripts can also target it directly.

The repository layout reflects this architecture: modules/ contains the Scala backend modules, website/ contains the documentation site, and nix/ contains NixOS-related configuration. The flake.nix and flake.lock files support Nix-based development environments.

Docspell versus Paperless-ngx

Paperless-ngx is a community fork of the original Paperless document management system. It is written in Python and Django for the backend, with a Vue-based frontend. Its OCR pipeline uses Tesseract and it also supports automatic tagging through its own document classifier.

The two projects target similar use cases: self-hosted personal or small-organization DMS with OCR and automatic classification. The differences are in the technology stack and the NLP approach. Docspell uses the Stanford CoreNLP library for named entity recognition as part of its classification, while Paperless-ngx uses its own classifier based on training on user-tagged documents. Docspell's backend is Scala on the typelevel functional stack; Paperless-ngx's is Python/Django, which many more developers are familiar with.

For a team already running Python services, Paperless-ngx may be easier to extend or debug. Docspell is the better choice for someone who prefers a strongly-typed functional backend or needs the Stanford NLP entity extraction specifically.

Limitations and known constraints

Docspell requires external binaries to process many common document types. Tesseract must be available for OCR on scanned images, unoserver and unoconvert must be running for converting .docx, .odt, and similar formats to PDF, and ocrmypdf is used as a wrapper for OCR-quality PDF production. In the Docker Compose setup this is handled automatically, but bare-metal or minimal container installations will fail silently on file types that need these tools if they are not present.

The machine learning suggestions are only as good as the labeled documents in your collection. A new installation produces no useful suggestions. The suggestion quality improves gradually as you consistently tag documents, and it degrades if tagging is inconsistent.

The latest stable GitHub release is v0.43.0 from March 2025. The nightly release is updated more frequently (the most recent is from September 2026). Users who need the latest features and bug fixes from the master branch can use the nightly build, but it carries less stability than a tagged release.

Docspell is licensed under AGPLv3 or later. For individuals and most self-hosted deployments, this is not a constraint. For organizations that plan to offer Docspell as a hosted service to others, the AGPL copyleft terms require that the source of any modified version be made available to users of that service.

Maintenance and license

The repository is not archived. The last push was on 2026-09-25. The project has an active nightly release and the latest stable release, v0.43.0, was published on 2025-03-15. The codebase is under active development with Scala Steward handling automated dependency updates, visible in the .scala-steward.conf configuration file.

The license is AGPLv3 or later. The project accepts donations via Liberapay and PayPal.

Editorial conclusion

Docspell fits households, families, and small organizations that want a self-hosted DMS with automated OCR, NLP-based metadata suggestions, and REST API access. It is a poor fit for larger organizations that need audit trails, document versioning, or enterprise access control, or for teams that are not comfortable running a multi-service deployment with external OCR and NLP dependencies. Before deploying, confirm that your environment supports tesseract for OCR and unoserver/unoconvert for document conversion, since those external tools are required for full file processing and must be available in the deployment environment.

Frequently asked questions

How does Docspell compare to Paperless-ng?

Both are self-hosted document management systems with OCR and automatic tagging. Docspell uses a Scala backend with Stanford CoreNLP for named entity extraction, while Paperless-ngx (the maintained fork of Paperless-ng) uses a Python/Django backend with its own document classifier. The feature scope is similar; the main difference is the technology stack and the NLP approach.

How does Docspell compare to Papermerge?

Papermerge is a Python-based open-source DMS focused on document organization and OCR. Docspell adds a machine learning layer using Stanford CoreNLP to suggest correspondents and tags from document content, which Papermerge does not provide. Docspell is written in Scala and Elm; Papermerge uses Python and Django.

What are the main alternatives to Docspell?

The most commonly compared alternatives are Paperless-ngx, Papermerge, and Mayan EDMS. All three are self-hosted open-source DMS projects with OCR support. The README does not compare them directly; differences in stack, license, and NLP approach are the main decision factors.

Official sources

  1. eikek/docspell on GitHub
  2. License: AGPL-3.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/eikek-docspell.svg)](https://hysenlabs.com/projects/eikek-docspell)