# BentoPDF: a self-hosted, client-side PDF toolkit with a dual licence

> BentoPDF runs PDF merging, splitting, conversion and OCR in the browser, with no upload step. The AGPL-3.0 build is free; a $79 commercial licence removes the copyleft obligation.

**alam00000/bentopdf** — The Privacy First PDF Toolkit. BentoPDF BentoPDF** is a powerful, privacy-first, client-side PDF toolkit that is self hostable and allows you to manipulate, edit, merge, and process PDF files directly in your browser.

- Repository: https://github.com/alam00000/bentopdf
- Website: https://bentopdf.com/
- Stars: 15,782 · Forks: 1,366
- Language: JavaScript
- License: AGPL-3.0
- Published: 2026-08-08 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/alam00000-bentopdf

## What BentoPDF solves, and who actually needs it

Most online PDF tools work the same way: you upload a file to a server, the server processes it, you download the result. That model fails for anyone handling contracts, medical records, identity documents or internal financials, because the file leaves the machine. BentoPDF's answer is to move the entire processing step into the browser. The README states that all processing happens in the browser and that files are never uploaded to a server. There is no upload limit to hit, because there is no upload.

The project targets two groups. The first is individuals and small teams who want the convenience of a web PDF tool without the data-handling question. The second is organisations that want to run the toolkit on their own infrastructure: the README documents static hosting on Netlify, Vercel and GitHub Pages, plus Docker, Podman and air-gapped deployments. The tool list covers the usual ground (merge, split, organise pages, extract, delete, rotate) and extends into conversion in both directions, PDF to text, Markdown, SVG and DOCX, plus EPUB, MOBI and XPS into PDF, along with compression, deskew, PDF/A conversion and digital signatures.

That breadth is the interesting part. A browser-only PDF merger is easy to write. Browser-only OCR, PDF/A conversion and font outlining are not, and those are the features that decide whether BentoPDF replaces a desktop application or merely complements one. The README also advertises unlimited file counts and no restrictions, which follows directly from the architecture: with no server-side queue and no per-file cost to the operator, there is nothing to meter.

## How the processing pipeline works: WASM modules loaded from a CDN

BentoPDF is a Vite application written in JavaScript, and it does not ship the heavy PDF engines in its own source tree. The README is explicit: the project does not bundle AGPL-licensed processing libraries but pre-configures CDN URLs so features work out of the box. Three components are fetched at runtime from jsDelivr: PyMuPDF for PDF to text, Markdown, SVG and DOCX conversion, image and table extraction, EPUB, MOBI and XPS conversion, compression and deskew; Ghostscript for PDF/A conversion and font to outline; and CoherentPDF for merge, split by bookmarks, table of contents, PDF to and from JSON, and attachments.

The .env.example file pins the exact URLs, including @bentopdf/pymupdf-wasm@0.11.16, @bentopdf/gs-wasm@0.1.1/assets/ and coherentpdf@2.5.5/dist/. This design keeps the repository small and the initial page load light, but it creates a hard runtime dependency on a third-party CDN for the more advanced tools. The project acknowledges this: the README points air-gapped and self-hosted users at a WASM Configuration section, and the Dockerfile exposes VITE_WASM_PYMUPDF_URL, VITE_WASM_GS_URL and VITE_WASM_CPDF_URL as build arguments. OCR assets are handled the same way through VITE_TESSERACT_WORKER_URL, VITE_TESSERACT_CORE_URL and VITE_TESSERACT_LANG_URL, with the note that all three must be set together for self-hosted or air-gapped OCR.

So the honest architecture description is: a static front end, plus WASM binaries pulled from a CDN unless you override the URLs at build time. If your network blocks jsDelivr, the basic tools still load but the PyMuPDF, Ghostscript and CoherentPDF features will not. Note which tools sit on which engine, because the split is not obvious from the UI: merge is CoherentPDF, compression is PyMuPDF, and PDF/A is Ghostscript. Losing one CDN path does not take down the whole toolkit, it takes down a specific slice of it.

## Installing BentoPDF with Docker Compose

The repository ships a docker-compose.yml that pulls a prebuilt image. The compose file lists two image families: ghcr.io/alam00000/bentopdf-simple:latest for the self-hosted build and ghcr.io/alam00000/bentopdf:latest for the commercial build, with bentopdfteam/ equivalents on Docker Hub. The default in the file is the simple image on port 8080.

```yaml
services:
  bentopdf:
    image: ghcr.io/alam00000/bentopdf-simple:latest
    container_name: bentopdf
    restart: unless-stopped
    ports:
      - '8080:8080'
```

Run docker compose up -d and open http://localhost:8080. The compose file also carries a commented-out environment block with DISABLE_IPV6=true, described as being for IPv4-only environments. If your host has no IPv6 route, uncomment it before starting the container.

For a source build, package.json defines the scripts. npm run dev starts the Vite dev server, and npm run build runs a chain that generates blog pages, static tool links, TypeScript compilation, the Vite build, SEO enhancement, i18n pages, a sitemap, security headers and an SEO audit. There is also a dedicated simple-mode script:

```bash
npm run serve:simple
```

That command sets SIMPLE_MODE=true, builds, and starts the preview server on port 3000. The README recommends Docker Compose or Podman Compose, so treat the source build as the path for contributors and for anyone who needs to change build-time variables such as VITE_DEFAULT_LANGUAGE, VITE_BRAND_NAME or the WASM URLs. The Dockerfile also accepts BASE_URL for serving under a subdirectory, which is the setting to reach for if you plan to put the app behind a reverse proxy on a path rather than a subdomain.

One thing the README does document and that is easy to miss: the digital signature feature needs a CORS proxy, configured through VITE_CORS_PROXY_URL and VITE_CORS_PROXY_SECRET. The README labels that section as required. Without it, certificate chain fetching will fail.

## The CDN dependency is the real deployment constraint

The privacy claim is about the files, and it holds: the browser does the work, so the document never travels. The claim does not extend to the tooling. On a default build, opening a conversion tool causes the browser to fetch a WASM binary from jsDelivr. That is a request to a third party, and on a locked-down corporate network it may simply fail.

The fix is documented but not trivial. You must build the image yourself with the three WASM URL arguments pointed at internal hosting, and for OCR you must set three Tesseract URLs plus, if you want non-default fonts, VITE_OCR_FONT_BASE_URL. The repository includes a bentopdf-airgap-bundle/ directory and an Air-Gapped / Offline Deployment section in the README, which suggests the maintainers have thought about this case. But the default path is CDN-first, and the README's own framing ("Zero-config by default") tells you which case is the happy path.

There is a second constraint worth naming: this is a browser application. There is no CLI, no server-side batch endpoint and no library API described in the README. If your workflow is "convert 4,000 PDFs overnight on a build server", BentoPDF is the wrong shape. It is a tool for a person sitting in front of a browser tab, or for an organisation that wants to give its staff one.

A third, quieter constraint is memory. Because the WASM engines run inside the tab, the ceiling on document size is the browser's, not a server's. The README claims the toolkit handles large PDF files with ease, but it does not publish a size threshold, and there is no documented way to raise a limit because there is no server-side limit to raise. Test with your own worst-case file before you standardise on it.

## BentoPDF compared with Stirling PDF

The most common comparison is with Stirling PDF, and the difference is architectural rather than feature-by-feature. Stirling PDF is a server application: you send the file to it, it processes the file with server-side libraries, and it returns the result. That makes it scriptable, batchable and easy to put behind an API, and it also means the document exists on the server for the duration of the job.

BentoPDF inverts that. The server, if you run one, is only serving static files. Everything else happens in the browser tab, which is why the README can claim files are never uploaded. The trade-off is that you cannot call BentoPDF from a shell script, and heavy jobs are bounded by the browser's memory rather than the server's. For an interactive tool used by a person, the browser model is fine. For an automated pipeline, it is not.

The other comparison that comes up is Adobe Acrobat. Acrobat is a desktop application with a paid licence and a much longer history; BentoPDF is a web toolkit under AGPL-3.0 or a $79 commercial licence. They overlap on editing and conversion, but Acrobat's strength is deep editing and a mature plugin ecosystem, while BentoPDF's is deployment flexibility and the absence of an upload step. Neither replaces the other cleanly.

## Licence, maintenance and upgrade cost

BentoPDF is dual-licensed. The package.json declares AGPL-3.0-only, and the README presents a table: AGPL-3.0 for open-source projects with public source code, free; commercial for proprietary or closed-source applications, $79 lifetime, described as one-time with unlimited devices and users, lifetime updates and no AGPL obligations. The repository also contains CCLA.md and ICLA.md, contributor licence agreements that support the dual-licence model.

The AGPL point that matters for adopters is the network clause. If you modify BentoPDF and let users interact with it over a network, the AGPL's source-availability obligation is triggered. That is the reason the commercial option exists, and it is the reason a company embedding BentoPDF in a closed internal product should read the licensing page rather than assume the free tier applies. This is not legal advice; the licence text and the project's licensing page are the sources.

On maintenance, the last push to the repository was on 2026-07-27, and that commit is tagged v2.8.7, described in the release list as CVE fixes. The two releases before it were v2.8.6 on 2026-06-28 and v2.8.5 on 2026-05-23, described as bug fixes. The cadence is roughly monthly, and the most recent release being a security fix is a signal worth noting: it means the project is receiving security attention, and it also means you should be tracking releases rather than pinning once and forgetting. The repository is not archived.

Upgrade cost is low if you run the container, since the compose file points at :latest and a pull gets you the new build. It is higher if you built with custom WASM URLs or custom branding, because those are build-time arguments and a rebuild is required. The repository also carries a .trivyignore file and security-headers-docs.conf, which suggests the maintainers run image scanning and care about response headers; if you serve the app through your own nginx rather than the bundled container, those headers are yours to reproduce.

## Who should adopt BentoPDF, and what to verify first

BentoPDF fits teams that handle documents they would rather not upload, and that want a browser UI their non-technical colleagues can use without training. It fits organisations with an internal Docker host and the appetite to mirror three WASM bundles if the network is restricted. It fits open-source projects, where the AGPL obligation is already satisfied.

It does not fit anyone who needs a command-line interface, a REST endpoint or unattended batch processing. It does not fit a team that cannot self-host the WASM assets and whose network blocks jsDelivr, because the advanced tools will silently be unavailable. And it does not fit a closed-source product that is unwilling to buy the commercial licence.

Three things to verify before you commit. First, open the v2.8.7 release notes and read what the CVE fixes address, then decide whether your deployment model is affected. Second, test one PyMuPDF-backed tool and one Ghostscript-backed tool from your actual network, so you find out early whether the CDN is reachable. Third, check the digital signature CORS proxy requirement against your environment, since the README marks it as required and it is not part of the default compose setup.

## Conclusion

Adopt BentoPDF if your PDF work has to stay on the machine, if you want a browser UI rather than a scripting API, and if you can live with AGPL-3.0 or pay the $79 commercial licence. Do not adopt it if you need server-side batch processing, or if you cannot host the PyMuPDF, Ghostscript and CoherentPDF WASM modules yourself. Before committing, check the release notes for the v2.8.7 CVE fixes, confirm whether your deployment can reach jsDelivr or must mirror the WASM assets, and read the licensing page against how you intend to distribute the software.

## FAQ

### Is BentoPDF safe to use?

The README states that all processing happens in the browser and that files are never uploaded to a server, so the document itself does not leave the machine. The WASM processing modules are fetched from jsDelivr by default, which is a third-party request, though the URLs can be overridden for self-hosted or air-gapped deployments.

### Is BentoPDF free?

It is dual-licensed. The AGPL-3.0 option is free for open-source projects with public source code, and a commercial licence is listed at $79 lifetime for proprietary or closed-source applications.

### What does BentoPDF do?

It is a client-side PDF toolkit covering organise and manage operations such as merge, split, extract, delete and rotate, conversion to and from PDF, and secure and optimise operations including compression, PDF/A conversion and digital signatures.

### How do I install BentoPDF?

The repository ships a docker-compose.yml that pulls ghcr.io/alam00000/bentopdf-simple:latest and maps port 8080. For a source build, package.json provides npm run dev and npm run serve:simple.

### How do I use BentoPDF offline?

The README documents an Air-Gapped / Offline Deployment section and the repository includes a bentopdf-airgap-bundle/ directory. The Dockerfile exposes VITE_WASM_PYMUPDF_URL, VITE_WASM_GS_URL and VITE_WASM_CPDF_URL so you can point the WASM modules at internal hosting instead of the default CDN.

### How does BentoPDF compare with Stirling PDF?

Stirling PDF processes files on a server, which makes it scriptable and batchable but means the document exists on that server during the job. BentoPDF runs the processing in the browser tab, so the file is never uploaded, but it has no CLI or server-side batch endpoint described in the README.

## Sources

- [Official documentation](https://bentopdf.com/)
- [Official README](https://github.com/alam00000/bentopdf#readme)
- [Project repository](https://github.com/alam00000/bentopdf)
- [Release notes](https://github.com/alam00000/bentopdf/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/alam00000-bentopdf
