Paperless-ngx: Self-Hosted Document Management System with OCR and Full-Text Search
A community-supported supercharged document management system: scan, index and archive all your documents
At a glance
- What is it?
- Paperless-ngx is a self-hosted document management system that ingests scanned and digital documents, OCRs them, and makes them searchable in a web interface. It is the officially maintained successor to Paperless and Paperless-ng, released under GPL-3.0, and deployed primarily via Docker Compose.
- Who is it for?
- Paperless-ngx is a well-maintained self-hosted document archive for home users and small teams who can run Docker Compose on trusted hardware and who accept that documents are stored in cleartext. It is not appropriate for regulated environments handling sensitive data without additional encryption infrastructure.
- Can I use it commercially?
- Yes, with conditions. GPL-3.0 is a copyleft licence: if you distribute software that includes it, you must release that software's source code under the same licence. Running it internally without distributing it does not trigger that obligation.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What Paperless-ngx Does and Who It Is For
Paperless-ngx transforms physical documents and digital files into a searchable online archive. It ingests scanned PDFs, images, and other document types, applies OCR to make the text searchable, assigns metadata such as tags, correspondents, and document types, and presents everything through a web interface. A full feature list and screenshots are in the documentation at docs.paperless-ngx.com.
The primary audience is home users and small organizations who want to reduce physical paper and maintain a searchable digital archive on hardware they control. The project's own README warns that Paperless-ngx should never be run on an untrusted host because information is stored in cleartext without encryption. The README recommends running it on a local server in your own home with backups in place. Self-hosters running Home Assistant, Synology NAS devices, Unraid, TrueNAS, or Proxmox are common deployers, as shown by the related search queries.
Architecture: Django Backend, Celery Workers, and Angular Frontend
The backend is a Django application with a Django REST Framework API. Celery handles background processing tasks such as OCR, document consumption, and email checking. Redis is the Celery broker. The frontend is an Angular application in the `src-ui/` directory; the Dockerfile compiles it with Node.js in a separate build stage before the final image. The Python backend is in `src/` and requires Python 3.11 or later.
The dependency list in `pyproject.toml` reveals the project's functional scope. It includes `celery[redis]` for async task processing, `django-allauth` with MFA and social account support, `llama-index-core` and `llama-index-embeddings-huggingface` for AI-powered features, `langdetect` for automatic language detection, `imap-tools` for email consumption, `gotenberg-client` for document conversion, and `azure-ai-documentintelligence` for Azure OCR integration. The version is 3.2.1. The Dockerfile uses s6-overlay for process supervision and builds on `ghcr.io/astral-sh/uv` for Python dependency management.
Installing Paperless-ngx with Docker Compose
The README presents Docker Compose as the recommended deployment path. The fastest start uses the install script:
bash -c "$(curl -L https://raw.githubusercontent.com/paperless-ngx/paperless-ngx/main/install-paperless-ngx.sh)"This script prompts for configuration and writes a `docker-compose.yml` and a `paperless.conf` environment file. The compose files in the `docker/compose/` directory of the repository pull the image from the GitHub container registry. After the containers start, the web interface is accessible on the configured port (the default is 8000). A demo is available at demo.paperless-ngx.com with login `demo` / `demo`; the README notes that demo content is reset frequently.
Migrating from the older Paperless-ng project is described in the documentation: the process involves replacing the Docker image and running Django migrations. Users on Synology or other NAS platforms follow platform-specific guides that the community maintains.
Document Consumption, OCR, and AI Features
Documents enter Paperless-ngx through a watched consume directory, email ingestion via IMAP, or direct web upload. Once a document lands in the consume directory, a Celery worker picks it up, runs OCR using a configured backend (the project supports several OCR engines), detects the document language via `langdetect`, and stores the result in the database with the extracted text.
The `pyproject.toml` includes `llama-index-core` and `llama-index-embeddings-huggingface`, indicating that Paperless-ngx 3.x has added AI-powered features for document analysis and search beyond keyword matching. The `azure-ai-documentintelligence` dependency provides an alternative OCR backend using Azure's document intelligence service. The exact configuration for these AI features is in the documentation at docs.paperless-ngx.com.
Security Warning: Cleartext Storage
The README contains an explicit security warning in bold: Paperless-ngx stores information in cleartext without encryption, and no guarantees are made regarding security. The warning specifically states that document scanners are typically used to scan sensitive documents such as social insurance numbers, tax records, and invoices. The recommended configuration is a local server in a home network with backups in place.
This is a hard constraint for any deployment scenario involving documents with regulatory confidentiality requirements. Organizations bound by healthcare privacy law, financial data regulations, or legal professional privilege must not run Paperless-ngx on untrusted shared infrastructure without additional encryption at the storage or network layer, which the project itself does not provide. Self-hosters who expose the web interface to the public internet should use a reverse proxy with authentication in front of it.
Limitations and Cases Where Paperless-ngx Is the Wrong Tool
Paperless-ngx is designed for personal or small organizational use on trusted hardware. It is not designed for multi-tenant deployments where different users' documents must be isolated from each other at the storage level. The cleartext storage constraint is not a configuration option, it is a fundamental design characteristic.
The Android app situation is unclear from the README, which links to a wiki page of related projects maintained by the community rather than an official mobile app. Backup is not automated by the project itself; the README points to documentation but does not provide a built-in backup job. Users must arrange their own PostgreSQL dump and filesystem snapshots or use the consume directory export function.
The dependency set is large. The Docker image is the practical deployment path; running Paperless-ngx outside Docker requires managing a Python environment, a PostgreSQL database, Redis, and an OCR engine, as well as compiling the frontend separately.
Comparing Paperless-ngx to Docspell
Docspell is the most commonly cited alternative. It is also a self-hosted document management system with OCR and full-text search, released under GPL-3.0. Docspell is written in Scala and uses PostgreSQL or MariaDB. Its data model centers on items (groups of files) rather than individual documents, which suits workflows where multi-page correspondence is stored as a unit. Paperless-ngx uses a flat document model with tags and correspondents as the primary organizational unit.
Paperless-ngx has a more active release cadence (version 3.2.1 as of 2026-09-20) and a Docker-first deployment story. Docspell requires more configuration to get running. For users already familiar with Django administration or who want an active community-supported project, Paperless-ngx is the practical choice. Docspell suits users who prefer a Scala-based stack or the item-centric document model.
Maintenance, Licence, and Community
Paperless-ngx is the official successor to the original Paperless project and to Paperless-ng. The project is structured around team responsibility rather than a single maintainer: the GitHub organization has teams for frontend, CI/CD, and other areas. The last push was on 2026-09-27, the most recent release is v3.2.1 from 2026-09-20, and the project is under active development.
The licence is GPL-3.0. The copyleft terms apply to the software itself; documents stored in Paperless-ngx are not affected by the licence. Translations are coordinated on Crowdin. Feature requests go through GitHub Discussions. The project also maintains a Matrix room for community support.
The `pyproject.toml` confirms the project targets Python 3.11 through 3.15, which is an unusually wide range and suggests intentional forward-compatibility testing.
Editorial conclusion
Paperless-ngx is a well-maintained self-hosted document archive for home users and small teams who can run Docker Compose on trusted hardware and who accept that documents are stored in cleartext. It is not appropriate for regulated environments handling sensitive data without additional encryption infrastructure. Before deploying, run the demo at demo.paperless-ngx.com to confirm the web interface meets your workflow, verify that your hardware can run Docker Compose, and read the security warning in the README about cleartext storage before scanning any sensitive documents.
Frequently asked questions
Is Paperless-ngx any good?
Paperless-ngx is under active development with version 3.2.1 released on 2026-09-20. It provides OCR, full-text search, tag-based organization, email ingestion, and AI-powered features in version 3.x. A live demo is available at demo.paperless-ngx.com for evaluation.
What is Paperless-ngx for?
Paperless-ngx turns scanned documents and digital files into a searchable online archive. It ingests documents through a watched folder, IMAP email, or web upload, applies OCR, and organizes them with tags, correspondents, and document types in a self-hosted web interface.
How do you install Paperless-ngx?
The recommended path is Docker Compose. Run the install script with `bash -c "$(curl -L https://raw.githubusercontent.com/paperless-ngx/paperless-ngx/main/install-paperless-ngx.sh)"` to configure and start the containers. Manual setup options are in the documentation at docs.paperless-ngx.com.
Is Paperless-ngx secure?
The README explicitly states that documents are stored in cleartext without encryption, and recommends running Paperless-ngx only on a local server in your own home with backups in place. It should not be exposed on a public network without a reverse proxy with authentication.
Which is better for document management, Docspell or Paperless-ngx?
Paperless-ngx uses a flat document model with tags and correspondents, has an active release cadence, and deploys via Docker Compose. Docspell uses an item-centric model grouping files together and is written in Scala. The choice depends on the preferred document model and stack.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/paperless-ngx-paperless-ngx)