# Papra: a minimalistic self-hosted document archive you run with one Docker command

> Papra is an AGPL-3.0 TypeScript document archiving platform that installs from a single container image. It is aimed at people who want long-term storage and retrieval without the scanning-pipeline complexity of heavier tools.

**papra-hq/papra** — The minimalistic document archiving platform.

- Repository: https://github.com/papra-hq/papra
- Website: https://demo.papra.app
- Stars: 5,524 · Forks: 308
- Language: TypeScript
- License: AGPL-3.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/papra-hq-papra

## The problem Papra solves, and who it is actually for

Papra targets a specific gap: documents you already possess as files, and the slow decay of your ability to find them again. The README frames the problem in plain terms, describing a platform for "long-term document storage and management, like a digital archive for your documents", and gives the warranty for a new phone and a gift receipt as examples. That is not a scanning workflow. It is a retrieval problem.

The intended user is someone who wants a place to put PDFs and images and then forget about them until they are needed. The feature list backs this up: upload and store, full-text search with advanced filters, tags, organizations for sharing with family or colleagues, and document sharing with external users with optional expiration dates and password protection.

Organizations are the unit of separation. Custom properties are defined per-organization rather than globally, which suggests the design assumes multiple distinct collections with different metadata needs, not one flat pile. If you are a single person archiving your own paperwork, you will use one organization and ignore most of that surface. If you are a household or a small team, the organization model is the reason to pick Papra over a folder on a NAS.

What Papra is not: a scanning application, an invoice-approval workflow, or an accounting system. Nothing in the README describes OCR of a physical scanner feed, document approval chains, or financial reporting. It ingests and indexes. That boundary is worth respecting before you evaluate it.

## How ingestion and search work under the hood

The architecture is a pnpm monorepo. The top-level package.json declares the root package @papra/root at version 0.3.0, and the repository splits into apps/ and packages/. The release feed names @papra/app and @papra/webhooks as separately versioned artifacts, so the webhook delivery layer is its own package rather than something bolted into the server.

The backend stack, per the acknowledgements section, is HonoJS for the API, Drizzle as the ORM, Better Auth for authentication, and CadenceMQ as the job queue. CadenceMQ is described as "a self-hosted-friendly job queue for Node.js, made by Papra", which tells you the project deliberately avoided a Redis or cloud-queue dependency for background work. The frontend is SolidJS with Shadcn Solid components and UnoCSS.

That job queue is where the interesting ingestion paths land. The feature list includes email ingestion (forward mail to a generated address), folder ingestion (import from a watched folder), content extraction ("Automatically extract text from images or scanned documents for search"), and tagging rules that apply tags based on custom conditions. Each of those is asynchronous work, which is why a queue exists at all.

The data flow for a forwarded email is therefore roughly: mail arrives at the generated address, a job is enqueued, the attachment is stored, content extraction runs if configured, tagging rules evaluate, and the resulting document becomes searchable. The README does not document the queue's retry semantics or what happens to a job that fails permanently, so treat failure handling as something to verify in your own deployment rather than something the project promises.

## Installing Papra with Docker and uploading a first document

The README gives a single quick-start command. It publishes port 1221 and sets AUTH_SECRET. Run it as written:

```bash
docker run -d --name papra -p 1221:1221 -e AUTH_SECRET=a-dummy-secret-for-testing-purposes-only ghcr.io/papra-hq/papra:latest
```

The image is ghcr.io/papra-hq/papra:latest. The README explicitly warns that the placeholder secret is only for kicking the tires, and that a real instance should generate its own with openssl rand -hex 48 or similar. Take that seriously: AUTH_SECRET is the value your authentication depends on, and a published placeholder is a published placeholder.

```bash
openssl rand -hex 48
```

Use the output as the value of AUTH_SECRET in your run command or compose file. The README does not document a docker compose file, but the related searches suggest people look for one; the self-hosting documentation at docs.papra.app/self-hosting/using-docker is where the README points for "more information and configuration options", so that page is the place to look for the full variable set rather than guessing.

Once the container is up, open http://localhost:1221, create an account, then upload a document through the interface. The README does not spell out the first-run flow beyond authentication being a feature, so the exact onboarding sequence is something you will discover in the UI. After upload, the document should appear in the list and be reachable through full-text search once content extraction has processed it. If you want to try the interface before running anything, the README offers demo.papra.app and notes it has "no backend, client-side local storage only", so nothing you put there is stored anywhere real.

## Where Papra stops short

The sharpest limitation is stated by the project itself. The README lists document requests as "Coming soon", and the mobile app, desktop app, browser extension and AI features as "Coming maybe one day". Those are not shipped capabilities, and the phrasing is honest about it. If your workflow depends on a phone camera upload or a desktop sync folder, Papra does not have either today.

The second limitation is the absence of a scanner pipeline. Content extraction exists, and the related searches show people asking about Papra OCR, but extraction is described as a step applied to documents that are already in the system. There is no documented driver for a physical scanner, no ADF batch import, and no duplex scan handling. If your documents are paper in a filing cabinet, Papra is the wrong end of the workflow to start from. Scan to files first with something else, then ingest.

The third is operational. The README documents one required environment variable and one port. Everything else lives in the self-hosting docs. That is fine for a quick start and thin for production: backup strategy, database choice, storage backend and migration behaviour are not covered in the README at all. The repository has a .changeset directory and a renovate.json, which indicates the maintainers manage dependency updates and versioning deliberately, but that says nothing about how you upgrade a running instance without losing data. Verify that before you put anything irreplaceable in it.

Finally, the demo is explicitly frontend-only. It is useful for judging whether the interface suits you and useless for judging ingestion, search quality or extraction accuracy.

## Papra compared with Paperless-ngx

The most common comparison in the search data is papra vs paperless-ngx, and the difference is one of scope rather than quality. Paperless-ngx is built around a document consumption pipeline: you point it at a scanner or a watch folder, it consumes, OCRs, and classifies. Papra's README describes it as "minimalistic" and puts the emphasis on storage, organizations, tags, custom properties and sharing.

So the practical distinction: if your intake is paper and you want the software to own the scanning step, a consumption-pipeline tool fits better. If your intake is already digital files and you want a clean place to archive them with search, sharing and per-organization metadata, Papra's smaller surface is the point. Fewer moving parts means fewer things to configure and fewer things to break.

Papra does have ingestion paths that overlap: email forwarding to a generated address, folder ingestion, and tagging rules. Those cover a lot of what people use a watch folder for. What is missing relative to a scanning-first tool is the scanner driver and the batch paper workflow.

The other real difference is the licence. Papra is AGPL-3.0, stated in the README and confirmed by the root package.json, which declares AGPL-3.0-or-later. If you intend to offer a modified Papra as a network service to others, the AGPL's network clause is the thing your legal team will want to read. That is not a reason to avoid the project for personal or internal use, but it is a reason to check before building a product on top of it.

## Maintenance, releases and the cost of staying current

The repository is not archived, and the last push was on 2026-09-10. Releases are frequent and versioned per package: @papra/app@26.6.2 on 2026-09-01, @papra/app@26.6.1 on 2026-07-04, and @papra/webhooks@0.3.4 on 2026-07-02. The gap between 26.6.1 and 26.6.2 is roughly two months, with 26.6.2 carrying a date-based major version, which is a common convention for apps that release on a schedule.

The upgrade cost is the part the README does not address. There is a .changeset directory and a changeset script in the root package.json, so changes are tracked and versioned with intent. What is not documented in the README is the migration path for a running instance: whether database migrations run automatically on container start, whether they are reversible, or what happens if you pin to :latest and a schema change lands. The README does not document rollback.

For a personal archive, the practical approach is to pin a specific image tag rather than :latest, and to keep the underlying data volume backed up independently of the container. Because the README does not specify where documents and the database live inside the container, that is the first thing to confirm from the self-hosting documentation before you rely on a backup you have not verified.

On licence: AGPL-3.0-or-later means you can self-host and modify freely, but distributing a modified version, or running one as a network service for others, brings obligations. The README points to the LICENSE file for details. Read it rather than assuming, and get proper advice if the network clause touches your use case.

## Conclusion

Adopt Papra if you want a self-hosted archive for documents you already have as files, and you are comfortable with Docker and a single AUTH_SECRET. Skip it if you need a scanner-driven capture workflow, a mobile app, or a desktop client, none of which exist yet. Before committing, run the demo at demo.papra.app to judge the interface, then check the self-hosting documentation at docs.papra.app for the full environment variable list, since the README covers only AUTH_SECRET and port 1221.

## FAQ

### What is Papra?

Papra is a minimalistic document management and archiving platform, built in TypeScript and licensed AGPL-3.0. It stores, tags and searches documents, with organizations, sharing, email and folder ingestion, and a CLI, API and webhooks.

### How do I install Papra?

The README gives a single Docker command: docker run -d --name papra -p 1221:1221 -e AUTH_SECRET=... ghcr.io/papra-hq/papra:latest. The README warns the placeholder secret is only for testing and says to generate your own with openssl rand -hex 48, and points to docs.papra.app/self-hosting/using-docker for further configuration.

### How does Papra compare with Paperless-ngx?

The README positions Papra as minimalistic, focused on storing and retrieving documents you already have, with organizations, tags, custom properties and sharing. A consumption-pipeline tool is built around scanning and OCR intake, which Papra's README does not describe; Papra's ingestion paths are email forwarding, folder import and tagging rules.

### What are alternatives to Papra?

Any self-hosted document manager that owns the scanning and OCR pipeline is the natural alternative, since Papra's README describes ingestion through email forwarding, folder import and tagging rules rather than a scanner driver. The choice comes down to whether your documents arrive as paper or as files.

## Sources

- [License: AGPL-3.0](https://github.com/papra-hq/papra/blob/main/LICENSE)
- [papra-hq/papra on GitHub](https://github.com/papra-hq/papra)
- [Project website](https://demo.papra.app)
- [README](https://github.com/papra-hq/papra/blob/main/README.md)
- [Releases](https://github.com/papra-hq/papra/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/papra-hq-papra
