pdf-craft: turning scanned books into Markdown and EPUB with DeepSeek OCR
PDF craft can convert PDF files into various other formats. This project will focus on processing PDF files of scanned books.
At a glance
- What is it?
- PDF Craft is a Python library for scanned books and technical documents that extracts page content into Markdown or EPUB, with optional translation and a reusable .pcex extraction file. It is a beta project, and it needs an OCR service, local or remote.
- Who is it for?
- Adopt pdf-craft when your input is a scanned book or technical document and your output is Markdown or EPUB that someone will edit or read. Do not adopt it as a general PDF converter for digital-native files, a form-filling tool, or an offline pipeline: OCR and translation both call a service, and the project classifies itself as Beta.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The scanned-book problem pdf-craft targets
Most PDF tooling assumes the file already contains text. A scanned book does not. It contains page images, and everything a reader cares about (chapter order, footnotes, tables, formulas, the table of contents) has to be reconstructed rather than read off a text layer. That reconstruction is the problem pdf-craft addresses, and it is narrower than "convert PDF files into various other formats" suggests.
The README states the focus plainly: the project processes PDF files of scanned books. It names scanned books and academic or technical documents as the target, and it organizes body text, chapters, tables of contents, footnotes, tables, formulas, and images for further editing and reading. The intended reader is a developer or digitization worker who wants an editable Markdown file or an EPUB that opens in an ebook reader, not someone who needs a page-for-page visual copy of the original.
The project also ships an online app at pdfcraft.ai. That path is for trying the workflow in a browser, and the README says features and usage requirements there are defined by the online app itself, not by the library. Treat the two as separate products with a shared name.
How the pipeline works: OCR vendor, extraction, rendering
The architecture visible in the README has three separable stages. First, pages are rasterized: pdf2image is a declared dependency, and Poppler is listed as a requirement, which is consistent with rendering PDF pages to images before OCR. Second, an OCR vendor turns those images into structured page content. Third, a renderer writes that content out as Markdown, EPUB, or a translated PDF.
The OCR stage is pluggable by configuration rather than by code. The README lists DeepSeek OCR, DeepSeek OCR 2, and Unlimited OCR as supported, each with local and remote configurations. A remote OCR setup sends pages to the configured service and does not require local CUDA; a local setup runs the model on your own NVIDIA GPU. Translation and table-of-contents analysis are a separate concern with a separate text LLM, so OCR and translation have independent configurations.
The extraction stage is the part worth understanding before you commit. A conversion can write a .pcex extraction file, described as reusable for later rendering, translation, or processing on another machine. That is the design answer to a real cost problem: OCR is the slow and expensive step, so the project lets you pay for it once and render many times. The .pcex format has its own reference document in docs/en/PCEX_FORMAT.md.
One detail in the README is easy to skip past. For EPUB bibliographic metadata, front-page OCR metadata extraction is opt-in and uses a separate metadata LLM. The README states that every accepted value is verified against OCR evidence, and that PDF file properties are used only to fill missing fields, never to override printed book information. That is a deliberate ordering, and it is the opposite of what most converters do, which is to trust embedded metadata first.
Installing pdf-craft and converting your first book
The README's quick start requires Python 3.11 to 3.13, Poppler, and a working DeepSeek OCR-compatible service configuration. Poppler setup is covered in docs/en/INSTALLATION.md. The package installs from PyPI:
python -m pip install pdf-craftThe pyproject.toml confirms the Python constraint as >=3.11,<3.14 and lists the runtime dependencies, which include pdf2image, pypdf, openai, pydantic, and epub-generator. The standard installation supports remote OCR. Local OCR is a separate optional dependency group named local, which the README describes as requiring CUDA, sufficient VRAM, and model files.
The first conversion takes a PDF path and an output path, with the OCR vendor configured inline. Place input.pdf in the directory where you run the script:
from pdf_craft import DeepSeekOCRVendorConfig, PDFCraft, PDFOptions
craft = PDFCraft(
pdf=PDFOptions(
ocr=DeepSeekOCRVendorConfig(
base_url="https://example.com/v1",
api_key="your-api-key",
model="deepseek-ocr",
),
),
)
craft.convert_pdf_to_markdown("input.pdf", "output.md")The base_url in that snippet is a placeholder, not a working endpoint. The README says to use a compatible service that actually provides the OCR model, and points to docs/en/OCR_BACKENDS.md for other models. After the run, open output.md. Documents containing images also produce asset files, and the README is explicit that you keep those files with the Markdown when moving or sharing it.
EPUB output reuses the same configured craft instance and swaps the final call:
craft.convert_pdf_to_epub("input.pdf", "output.epub")If you expect to re-render or translate the same book, add the extraction path so the OCR work is not repeated:
craft.convert_pdf_to_markdown(
"input.pdf",
"output.md",
extraction_path="book.pcex",
)Translation is a separate configuration from OCR. The README describes supplying a chapter translator when converting to Markdown or EPUB, translating an existing EPUB directly, and choosing between replacing the original text and appending the translation for bilingual reading. Producing a translated PDF is a different workflow: it extracts content, translates it, and writes the translation back onto the original pages, and it additionally needs Ghostscript and suitable local fonts.
Where pdf-craft breaks down or is the wrong tool
The README itself sets the first limit: results depend on scan quality, page layout, and the OCR model, and it advises checking a representative document before processing a larger collection. That is not boilerplate caution. A skewed scan, a two-column academic layout, or a page with marginalia will produce different output than a clean single-column novel, and the library gives you no quality score to tell you which happened.
The translated-PDF workflow is the weakest link. Writing translated text back onto source pages is a layout problem, and the README only says to check the resulting layout against the source and translated text. There is no documented reflow guarantee, no statement about what happens when the translation is longer than the original text block, and no rollback path described. If your deliverable is a translated PDF that must look right, budget manual review per page.
The runtime split is a second constraint. Remote OCR avoids local CUDA, but it sends page images to the configured service. Local OCR keeps pages on your machine, but the README notes that using a remote LLM for translation or table-of-contents analysis still sends the corresponding content to that service. So "local" is a per-stage property, not a property of the pipeline. If your constraint is that no page content may leave the machine, you need local OCR and local models for the other stages too, and the README does not describe a fully offline configuration.
Finally, the project is not a general PDF toolkit. It has no documented support for filling forms, editing existing text layers, splitting or merging pages, or converting digital-native PDFs where a text layer already exists. The pyproject.toml classifier is Development Status :: 4 - Beta, and the README's own framing is about scanned books. Pointing it at a born-digital report and expecting a clean round trip is a misuse.
Alternatives and how the approach differs
The obvious comparison is a general-purpose OCR engine such as Tesseract, usually driven through a wrapper. The difference is what gets reconstructed. A Tesseract-based pipeline returns text per page; you then write your own logic for chapter boundaries, footnotes, tables, formulas, and a table of contents. pdf-craft treats those structures as part of the conversion, and its dependency list reflects that, with pylatexenc and mathml2latex for formulas and a dedicated epub-generator for ebook output. If your source is a clean single-column book and you only need raw text, that extra machinery is overhead you are paying for anyway.
The second comparison is a commercial document-AI service. Those typically bundle layout analysis, OCR, and export behind an API and handle the model hosting for you. pdf-craft instead makes the OCR vendor a configuration value, so you can point it at a compatible endpoint you already run or pay for, and it lets you persist the extraction as a .pcex file so later renders and translations do not re-run OCR. That is a different cost shape: more setup, but the expensive step is yours to cache and yours to re-run.
If your actual goal is translating an existing EPUB rather than converting a scan, the README describes an EPUB translation path that takes an EPUB directly. Using the full PDF pipeline for that job adds OCR and layout reconstruction you do not need.
Maintenance, upgrades, and what the MIT licence means here
The repository is not archived, and the last push was on 2026-09-18, four days before this writing. Releases v2.3.0 and v2.3.1 both landed on 2026-09-16, following v2.2.1 on 2026-09-10. That cadence means you should expect the dependency surface to move.
Upgrade cost is mostly dependency cost. The lockfile is poetry.lock, and the declared ranges are tight in places: json-repair is pinned to >=0.55.0,<0.56.0, epub-generator is pinned to exactly 0.1.7, and resource-segmentation and mathml2latex are on 0.0.x and 0.2.x ranges. A patch bump inside those ranges is unlikely to break you, but a major bump in pydantic, openai, or pypdf would, and those are the packages most likely to move. The pyside6 dependency is worth noting for a library: it is a GUI toolkit, and it pulls a large install footprint into environments that only ever call the conversion API.
The licence is MIT, declared in both the repository LICENSE file and the pyproject.toml project metadata. That is permissive: you can use, modify, and redistribute the code, including in closed products, provided the copyright notice and permission notice travel with it. Two things sit outside that grant and are not legal advice, just facts to check: the OCR models you configure are separate works with their own terms, and the online app at pdfcraft.ai is a service with its own usage requirements, which the README says are defined by that app.
Editorial conclusion
Adopt pdf-craft when your input is a scanned book or technical document and your output is Markdown or EPUB that someone will edit or read. Do not adopt it as a general PDF converter for digital-native files, a form-filling tool, or an offline pipeline: OCR and translation both call a service, and the project classifies itself as Beta. Before committing a collection, run one representative book end to end, check that the Markdown assets travel with the file, and confirm your OCR endpoint and model name are the ones your service actually serves.
Frequently asked questions
What is pdf-craft and what does it convert?
It is a Python library for scanned books and academic or technical documents that converts PDF pages to Markdown or EPUB, with optional translation and translated PDF output. The README describes it as focused on PDF files of scanned books.
How do I install pdf-craft?
Install it from PyPI with python -m pip install pdf-craft, on Python 3.11 to 3.13, with Poppler set up as described in docs/en/INSTALLATION.md. The standard installation supports remote OCR; local OCR is a separate optional dependency group.
Does pdf-craft need a GPU or an internet connection?
Remote OCR sends pages to the configured service and does not require local CUDA, while local OCR runs on your own NVIDIA GPU with CUDA, sufficient VRAM, and model files. Using a remote LLM for translation or table-of-contents analysis still sends that content to the service.
Can pdf-craft translate a book during conversion?
Yes. The README describes supplying a chapter translator when converting to Markdown or EPUB, or translating an existing EPUB directly, with translation-only and bilingual output modes. Translation uses a separate text LLM, configured independently of OCR.
What is the .pcex file for?
It is a reusable extraction file. Saving one lets you render, translate, or process the same book later, including on another machine, without repeating the OCR step. The README points to docs/en/PCEX_FORMAT.md for the format reference.
What licence does pdf-craft use?
MIT, declared in the repository LICENSE file and in the pyproject.toml project metadata. The OCR models you configure and the online app at pdfcraft.ai are separate from that grant.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/oomol-lab-pdf-craft)