Library / SDK
pymupdf/PyMuPDF avatar
pymupdf/PyMuPDF

PyMuPDF: a Python PDF engine built on MuPDF

PyMuPDF is a high performance Python library for data extraction, analysis, conversion & manipulation of PDF (and other) documents.

10,812 stars809 forksPythonAGPL-3.0

At a glance

What is it?
PyMuPDF wraps the MuPDF C engine in a Python API for text extraction, rendering, redaction, table detection and format conversion. It is the right tool when you need page-level control, and the wrong one when AGPL-3.0 does not fit your distribution model.
Who is it for?
Adopt PyMuPDF if you need page-level geometry, rendering, redaction or table detection and can live with AGPL-3.0. Do not adopt it if you ship closed-source software and cannot comply with that licence, or if you only need to assemble and split PDFs, where a permissively licensed pure-Python library is lighter.
Can I use it commercially?
Yes, with strict conditions. AGPL-3.0 is a network copyleft licence: if people use a modified version over a network, for example as a hosted service, you must offer them its source code under the same licence.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What PyMuPDF is for, and who ends up using it

PyMuPDF is a Python binding to MuPDF, a C engine. The README describes it as a library for data extraction, analysis, conversion, rendering and manipulation of PDF and other documents, and lists the design goals as fast, accurate, versatile and LLM-ready, with no mandatory external dependencies. The problem it solves is that PDF is a page description format, not a data format. Text on a page carries coordinates, fonts and sizes, and anything that needs those coordinates has to go through a renderer rather than a parser. PyMuPDF exposes that renderer directly.

The audience follows from that. Engineers building document ingestion for search or retrieval pipelines need per-span text with positions. Teams doing redaction need to remove content from the page content stream, not draw a black box over it. Anyone converting scanned pages needs rasterisation plus an OCR hook. The README also positions the project for AI work through PyMuPDF4LLM, which produces Markdown and JSON aimed at RAG pipelines. If your job is to fill in an AcroForm or stamp a watermark, the same API covers it.

How the MuPDF binding shapes the API

The architecture is a thin Python layer over a C library. setup.py states that the build downloads a hard-coded MuPDF tarball, extracts and builds it locally, and then builds PyMuPDF, so a source build always pairs the binding with the exact MuPDF release it expects. That is why wheels matter: on Windows, macOS and Linux with Python 3.10 to 3.14, pip installs a pre-built pair and no C toolchain is needed. Without a wheel, pip compiles from source and needs a C/C++ toolchain.

The data flow is document, page, then extraction mode. pymupdf.open returns a Document you can iterate. Each Page offers get_text with different output shapes: plain text by default, and a dict form whose blocks, lines and spans carry font, size and position. Rendering goes through get_pixmap, which returns a Pixmap you save to an image file. Tables are found by find_tables on a page, and each result converts to Markdown or to a Pandas DataFrame. Annotations and redactions are added to a page and then applied before saving. OCR is a separate path: get_textpage_ocr requires Tesseract installed and on PATH, and the README lists it as a separate install, not a Python dependency.

Installing PyMuPDF with pip and extracting your first page

The README gives a single install command. Wheels cover Windows, macOS and Linux on Python 3.10 to 3.14, so on those platforms nothing else is required.

bash
pip install pymupdf

If pip reports that it is building from source, you are on a platform without a wheel and will need a C/C++ toolchain. That is the main install failure mode, and it appears at build time rather than at import time.

The quick start opens a document and prints the text of every page. Run this against any PDF and you should see the extracted text on stdout, one page after another.

python
import pymupdf

doc = pymupdf.open("document.pdf")
for page in doc:
    print(page.get_text())

The second example is the one that shows why the binding exists. It reads the dict form of the first page and walks blocks, lines and spans, printing each span's text alongside its font name and size. Use it when plain text is not enough and you need to know which words are headings.

python
import pymupdf

doc = pymupdf.open("document.pdf")
page = doc[0]

blocks = page.get_text("dict")["blocks"]
for block in blocks:
    if block["type"] == 0:  # text block
        for line in block["lines"]:
            for span in line["spans"]:
                print(f"{span['text']!r}  font={span['font']}  size={span['size']:.1f}")

Rendering follows the same shape: open, index a page, call get_pixmap with a dpi value, and save. The README's example writes page_0.png at 150 dpi.

python
import pymupdf

doc = pymupdf.open("document.pdf")
page = doc[0]

pixmap = page.get_pixmap(dpi=150)
pixmap.save("page_0.png")

Optional extras install separately. pymupdf-fonts adds a font collection for output, pymupdf4llm adds Markdown and JSON extraction for LLM pipelines, and pymupdfpro adds Office document support. Tesseract is not a pip package at all; the README shows brew install tesseract on macOS and sudo apt install tesseract-ocr on Ubuntu and Debian.

Where PyMuPDF stops being the right choice

The licence is the first constraint, and it is not a footnote. PyMuPDF is AGPL-3.0. If you embed it in a network service, the AGPL's source-availability terms reach that service, and the repository does not document an alternative licensing path in what is available here. Teams that cannot accept those terms should decide before writing code, not after.

The second constraint is the dependency model. No mandatory external dependencies is a real advantage, but the build downloads a hard-coded MuPDF release and compiles it. That means your build is tied to whatever MuPDF version the PyMuPDF release pins, and you cannot swap in a system MuPDF without taking the build into your own hands; setup.py documents environment variables such as PYMUPDF_MUPDF_BUILD and PYMUPDF_MUPDF_LIB for exactly that, which tells you it is a supported but non-default path.

OCR is the third. get_textpage_ocr requires Tesseract on PATH, and the README treats it as a separate install. If your pipeline is mostly scanned documents, you are adding a second system dependency and its own language data, and the OCR quality is Tesseract's, not PyMuPDF's.

Finally, if your task is only to merge, split or stamp files, the page-geometry machinery is overhead you will not use. The README's merge example is four lines, which is a fair signal that the heavy parts of the library live elsewhere.

PyMuPDF compared with pypdf and with docling

The comparison people search for most is PyMuPDF against pypdf. The difference is the engine. PyMuPDF binds MuPDF, a C renderer, and gets coordinates, font metadata, rasterisation and table detection from it. pypdf is a pure-Python library, so it installs anywhere Python runs and has no compiled component, but it does not have a rendering engine behind it. If you need a page as pixels, or you need to know where a word sits on the page, that is a PyMuPDF job. If you need to concatenate files and rewrite metadata, pypdf is the lighter dependency.

The other comparison in the search data is docling. Docling is a document conversion pipeline aimed at structured output for downstream models, and it brings its own model and layout stack. PyMuPDF stays at the engine layer: it gives you the page and its primitives, and PyMuPDF4LLM is the optional package that turns those primitives into Markdown. The practical difference is where the intelligence sits. With PyMuPDF you decide what to do with the spans; with a pipeline tool, the layout decisions are largely made for you, which is convenient until it gets a page wrong and you have no coordinate to point at.

Release cadence, licence terms and upgrade cost

The repository is not archived, and the last push was on 2026-09-21. Recent releases are 1.28.2 on 2026-08-06, 1.28.0 on 2026-06-29 and 1.27.2.3 on 2026-04-24, so the project ships on a roughly two-month cycle. That cadence is the upgrade cost. Because a source build compiles a pinned MuPDF release, upgrading PyMuPDF can mean a new MuPDF underneath, which is a larger change than a patch version suggests. Pin the version in your requirements and test extraction output against a fixture set before moving.

The licence is AGPL-3.0, and the repository carries a COPYING file at the top level. The AGPL's network clause is the part that surprises people: offering the software's functionality over a network can trigger source disclosure obligations. Nothing here is legal advice, and the documentation is where you should check the project's own statement of terms. What is verifiable is the licence identifier and the fact that the README does not present a commercial or alternative licence option.

Editorial conclusion

Adopt PyMuPDF if you need page-level geometry, rendering, redaction or table detection and can live with AGPL-3.0. Do not adopt it if you ship closed-source software and cannot comply with that licence, or if you only need to assemble and split PDFs, where a permissively licensed pure-Python library is lighter. Before committing, install it on your target Python version and run the two snippets in this article against one of your own documents, then read the COPYING file and the licence section of the documentation to decide whether the AGPL terms fit your product.

Frequently asked questions

What is PyMuPDF used for?

It is used for data extraction, analysis, conversion, rendering and manipulation of PDF and other documents, according to the README. Concrete tasks listed there include extracting text with layout metadata, finding tables, rendering pages to images, OCR on scanned pages, redaction and merging files.

Is PyMuPDF free to use?

The repository is licensed AGPL-3.0 and carries a COPYING file. That is a free software licence, but its network clause can require source disclosure when the software is offered as a service, so the terms are not the same as a permissive licence.

Which is better, PDFPlumber or PyMuPDF?

The README does not cover PDFPlumber, so no direct comparison can be made from the project's own documentation. What can be said is that PyMuPDF binds the MuPDF C engine, which gives it rendering, font metadata and table detection through find_tables, and that its licence is AGPL-3.0.

How can I convert a PDF to an image using PyMuPDF?

Open the document, index the page, call get_pixmap with a dpi value and save the result. The README's example uses dpi=150 and saves to page_0.png.

How do I install PyMuPDF with pip?

Run pip install pymupdf. Wheels are available for Windows, macOS and Linux on Python 3.10 to 3.14; if no wheel exists for your platform, pip compiles from source and needs a C/C++ toolchain.

How do I use PyMuPDF to extract text from a PDF?

Open the file with pymupdf.open, iterate the document, and call get_text on each page. Passing "dict" instead of no argument returns blocks, lines and spans with font and size metadata.

Official sources

  1. License: AGPL-3.0
  2. Project website
  3. pymupdf/PyMuPDF on GitHub
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/pymupdf-pymupdf.svg)](https://hysenlabs.com/projects/pymupdf-pymupdf)