MarkPDFdown: a PDF to Markdown converter that calls a vision model per page
A high-quality PDF-to-Markdown tool based on large language model visual recognition.
At a glance
- What is it?
- A small Apache-2.0 Python package that hands pages to a multimodal model through LiteLLM, converts them to Markdown, and ships a CLI with file mode, pipe mode and a Docker path.
- Who is it for?
- MarkPDFdown is thin on purpose, and the interesting decisions are all in the thin part. It does not parse PDFs, it renders them and shows the image to a model, which means quality tracks the model you point it at and cost tracks the page count.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 16 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 8, 2026, and from our analysis. They are not legal advice.
Editorial analysis
A vision model in the loop, not a PDF parser
The architecture in one sentence is that pages go to a multimodal model and the answer is Markdown. The README describes the project as using multimodal large language models to transcribe PDF files, and names LiteLLM as the layer that carries the request. That single choice determines most of the trade-offs you will care about, and it is worth being blunt about them.
Because there is no PDF text or layout extraction step, the tool does not need to reconstruct reading order, guess column boundaries or decode a font's embedded encoding. Those are the failure modes that make conventional converters produce plausible garbage on real documents. Vision models do not have those failure modes because they are reading the rendered page the way a person does.
The cost of that approach is equally direct. Every page becomes an image sent to a paid API unless you point the tool at something local. Output quality is a property of the model you configure, not of this code, and there is no offline extraction path to fall back on when a document contains something the model will not reproduce. Whether that is the right trade depends entirely on what your input looks like: for scanned or complex-layout PDFs, a vision model is often the only thing that works, and for clean born-digital PDFs it is an expensive way to extract text.
The package itself is Python, Apache-2.0 licensed, and the requirements say Python 3.9 or newer. There are five translated READMEs alongside the English one, for Chinese, Japanese, Russian, Persian and Arabic, which is a lot of localization for a tool with under a hundred open issues.
Configuring the model with environment variables
Configuration is a `.env` file, and the repository ships `.env.sample` to copy from:
# Copy the sample configuration
cp .env.sample .envThe variables fall into three groups. The model name, the API key, and three optional parameters that have sensible defaults worth knowing about: `TEMPERATURE=0.3`, `MAX_TOKENS=8192` and `RETRY_TIMES=3`. A temperature that low is the right choice for transcription, where you want the model to report what it sees rather than improve on it, and it is a reminder that this is an extraction task wearing a generation model's clothes.
Model naming is where the three providers differ, and the difference is visible in the README's own examples. OpenAI models are bare names:
MODEL_NAME=gpt-4o
MODEL_NAME=gpt-4o-mini
MODEL_NAME=gpt-4-vision-previewOpenRouter and Requesty models carry a provider prefix, `openrouter/anthropic/claude-3.5-sonnet`, `openrouter/google/gemini-pro-vision`, `requesty/openai/gpt-4o` and `requesty/anthropic/claude-sonnet-4-5`. The keys are named for the provider rather than the model, so `OPENAI_API_KEY`, `OPENROUTER_API_KEY` and `REQUESTY_API_KEY` are interchangeable depending on which prefix you use, and LiteLLM picks up whichever is set.
That model list is also where the README stops being current, which the next section comes back to. The list of supported providers and the list in the Requirements section are not the same, and the desktop app's list is a third thing again.
File mode, pipe mode and page ranges
The CLI offers two modes, and the second one is the reason to reach for this over a wrapper script.
File mode is the everyday path, with `--input` and `--output`:
# Basic conversion
markpdfdown --input document.pdf --output output.mdPage ranges are handled with `--start` and `--end`, which is the practical way to handle a long document, since you can convert the chapter you care about instead of paying for every page:
# Convert specific page range
markpdfdown --input document.pdf --output output.md --start 1 --end 10The same flags work for images, since the tool reads a rendered page either way, and `python -m markpdfdown` is available when you would rather not put the entry point on your path.
Pipe mode is the one that changes how you would deploy it:
# PDF to markdown via pipe
markpdfdown < document.pdf > output.mdStandard input to standard output means the tool composes with anything and needs no temp files, which is exactly what the Docker section is built around:
docker run -i \
-e MODEL_NAME=gpt-4o \
-e OPENAI_API_KEY=your-api-key \
markpdfdown < input.pdf > output.mdFor a directory of documents the README gives a shell loop rather than a built-in batch flag, which is honest about where the batch logic lives:
for file in *.pdf; do
markpdfdown --input "$file" --output "${file%.pdf}.md"
doneThe feature list also claims image to Markdown conversion on the same code path, which follows from the design rather than being a separate implementation.
Version 1.2.0 is an announcement for a different repository
The release history is short and shaped oddly. Version 1.1.1 and 1.1.2 were both published on 2025-09-20, a little over two hours apart, and both are single-commit releases by the same contributor. The first fixes a uv path issue in the Dockerfile, the second fixes a missing Rust compiler when building `fastuuid` in CI. Then 1.2.0 arrives on 2026-01-25 and its release notes contain no library changes at all. They are an announcement for MarkPDFdown Desktop, with a link, a screenshot and a feature list.
That is the desktop app in its own repository, `MarkPDFdown/markpdfdown-desktop`, launched with:
npx -y markpdfdownWhich produces a naming collision worth pausing on. The Python package and the npm package share the name `markpdfdown`, but `npx -y markpdfdown` launches a graphical application, not the CLI described above. If you are automating document conversion and reach for the npx invocation expecting the command line tool, you will get a window instead.
The more interesting consequence is the provider list. The core README documents OpenAI, OpenRouter and Requesty. The desktop app's release notes list OpenAI, Anthropic Claude, Google Gemini, Ollama for local models, and the OpenAI Responses API. So the GUI can talk to Anthropic and Gemini directly and can run fully local through Ollama, while the CLI documentation only shows the three LiteLLM-routed providers. Both lists are accurate for their own surface, and together they suggest the desktop app has a newer provider layer than the CLI documentation reflects. If local models matter to you, that is the version to look at.
The desktop app also advertises a side-by-side preview of PDF pages next to the generated Markdown, LaTeX rendering through KaTeX, syntax-highlighted code blocks and a multi-language interface. None of that has a CLI equivalent.
The requirements list and the feature list disagree about providers
A smaller inconsistency sits in the README's own Requirements section. The feature list states multi-provider support across OpenAI, OpenRouter and Requesty through LiteLLM, and the Configuration section shows all three API keys. The Requirements section then closes with access to supported LLM providers, naming only OpenAI or OpenRouter.
Requesty is the odd one out, and it is the omission worth resolving rather than ignoring, because the difference between the two lists is whether you can point this at Requesty at all. The Configuration section answers it: `REQUESTY_API_KEY` and the `requesty/` prefix are documented, so the feature is present and the Requirements line is short.
Two other things in the tree are worth knowing because they tell you how the project is developed rather than how it runs. `uv.lock` is committed alongside `pyproject.toml`, which is the reproducible-resolution story, and the README recommends uv as the package manager with conda and pip as alternatives. `.pre-commit-config.yaml` is present, and the documented loop is ruff for formatting and linting, `ruff format`, `ruff check`, `ruff check --fix`, then `pre-commit run --all-files`.
There is also a `Makefile`, a `CHANGELOG.md` and an `AGENTS.md`. The last of those is a convention borrowed from the agent tooling world, a file that tells a coding assistant how to work in the repository, and its presence alongside a contributor section that still reads as a generic seven-step fork-and-branch walkthrough is a fair signal of where the project's own workflow has recently been focused.
The maintenance picture is healthy. The last push was on 2026-09-23, the repository is not archived, and the combination of seven open issues with 188 forks suggests most reports arrive as pull requests rather than tickets.
Where this sits against the tools readers compare it to
The README never positions this against other converters, which is its own kind of information. Readers arriving from search mostly arrive comparing it: the related searches for this project cluster around Microsoft's MarkItDown and Marker, plus repeated variants of the same practical question, batch converting PDFs, converting on Linux, and converting with images intact.
Those variants are a fair description of what people want from this category of tool, and they line up unevenly with what MarkPDFdown documents. Batch conversion exists but as a shell loop rather than a flag. Linux is covered through the native install and the container. Images are accepted as input, and the desktop app advertises image support in the output, which is the harder direction. Preserving a document's figures as extracted image files is not something the README claims, so if your PDFs contain diagrams you need to keep, that is the gap to test first.
The architecture block is worth reading as a guide to where a limitation would live:
src/markpdfdown/
├── cli.py # Command line interface
├── main.py # Core conversion logic
├── config.py # Configuration management
└── core/
├── llm_client.py # LiteLLM integration
├── file_worker.py # File processing
└── utils.py # Utility functionsEverything that decides output quality sits behind `llm_client.py`. Everything that decides whether a PDF is split into pages, rendered and re-encoded sits in `main.py` and `file_worker.py`. That is a small enough surface that if the output is wrong, you can tell fairly quickly whether the cause is the model, the rendering or the prompt, and `core/` is the place to look.
Editorial conclusion
MarkPDFdown is thin on purpose, and the interesting decisions are all in the thin part. It does not parse PDFs, it renders them and shows the image to a model, which means quality tracks the model you point it at and cost tracks the page count. The configuration surface is a handful of environment variables and a provider-prefixed model name, and the pipe mode makes it a clean fit for a container in a larger pipeline. Two things to check before adopting it: the newest release is an announcement for the separate desktop app rather than a library change, and the Requirements section lists only OpenAI and OpenRouter while the feature list adds Requesty. Start with `markpdfdown --input document.pdf --output output.md` against a ten page sample and compare the Markdown against what your own reader produces.
Frequently asked questions
Does MarkPDFdown use a local model or require a cloud API?
The CLI documentation shows OpenAI, OpenRouter and Requesty through LiteLLM, so it expects an API key and sends page images to a remote provider. The separate desktop app's release notes add Ollama for local models along with Anthropic and Gemini, so local inference is documented for the GUI rather than for the command line tool.
What is the difference between MarkPDFdown and MarkItDown?
They are separate projects and the MarkPDFdown README never compares them. Search suggestions for this project frequently pair the two names, which suggests readers treat them as substitutes, but this repository documents only its own behavior: rendered pages sent to a multimodal model, with OpenAI, OpenRouter and Requesty named as providers.
How do I convert a batch of PDFs to Markdown?
The README gives a shell loop rather than a batch flag: `for file in *.pdf; do markpdfdown --input "$file" --output "${file%.pdf}.md"; done`. Because the tool also reads standard input and writes standard output, you can wrap that loop in the documented `docker run -i` invocation instead.
What does npx markpdfdown install?
It launches the MarkPDFdown Desktop application from the separate markpdfdown-desktop repository, not the Python command line tool. The Python package shares the name but is a different artifact, which is an easy mistake when you reach for the npx form expecting a CLI.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/markpdfdown-markpdfdown)