# MarkItDown: eleven extras, two system binaries, one silent failure

> MarkItDown converts PDF, Office, HTML, archive, audio and YouTube input into Markdown for language model pipelines, and the interesting engineering sits in the extras, the system binaries the container installs, and the places it degrades without telling you. Version 0.1.8, MIT licensed, Python 3.10 through 3.14.

**microsoft/markitdown** — Convert files and office documents to Markdown with Python.

- Repository: https://github.com/microsoft/markitdown
- Stars: 187,295 · Forks: 13,835
- Language: Python
- License: MIT
- Published: 2026-08-04 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/microsoft-markitdown

## The narrowest convert_* call is the security contract

The important note at the top of the file is the part to read twice. MarkItDown performs I/O with the privileges of the current process, and like open() or requests.get() it will access resources that the process itself can access, so inputs should be sanitized in untrusted environments and the narrowest convert_* function should be called for the case, with convert_stream() and convert_local() given as examples. Put plainly, the converter inherits your permissions. Hand it a path or a YouTube URL and it opens or fetches whatever the calling user can reach, and the note describes no privilege separation, no allowlist and no sandbox. That matters in the deployment shape this tool encourages, which is a service that takes an upload or a link from a stranger and pipes it into the converter. The instruction is the narrow one, so prefer convert_stream() or convert_local() over the general entry point and sanitize the path before it arrives. The rest of the controls are yours to build.

## Eleven extras, and the ones you skip decide what breaks

Format support is opt-in, one extra at a time, and the catalogue is long: [all], [pptx], [docx], [xlsx], [xls], [pdf], [outlook], [az-doc-intel], [az-content-understanding], [audio-transcription] and [youtube-transcription]. A narrow install is a single line, and this is the entire command for PDF, DOCX and PPTX and nothing else.

```bash
pip install 'markitdown[pdf, docx, pptx]'
```

[all] is the documented default, which is convenient locally and expensive in an image, since a build with [all] carries an Azure integration it may never call. Two consequences follow for a pipeline. A missing extra is a format that fails at conversion time rather than at install time, so your test set needs one real file per extra you intend to deploy. And the repository does not document the error text for a missing extra, so a service that only ever saw PDFs meets its first DOCX in production and you read the traceback to discover which extra was absent.

## The image installs exiftool and ffmpeg, and pip does not

The Dockerfile shows what a complete installation actually requires. It starts FROM python:3.13-slim-trixie, sets EXIFTOOL_PATH=/usr/bin/exiftool and FFMPEG_PATH=/usr/bin/ffmpeg, switches off ONNX Runtime telemetry with ORT_DISABLE_TELEMETRY=1, then installs ffmpeg and libimage-exiftool-perl through apt before copying the tree and running pip --no-cache-dir install against /app/packages/markitdown[all] and /app/packages/markitdown-sample-plugin. Two details matter if you are on a laptop. The Python extras give you libraries, not binaries, so the EXIF and audio paths expect exiftool and ffmpeg to exist on the machine, and a plain pip install on macOS or Windows does not put them there. Git is not in the base image at all: it sits behind a build argument, ARG INSTALL_GIT=false, so any workflow that needs it has to build with it switched on. The image runs as USER nobody:nogroup by default, which is the right default for a converter that should never be writing files.

## The OCR plugin loads without llm_client and quietly does nothing

Plugins are off by default. You list what is installed with `markitdown --list-plugins`, enable them with `markitdown --use-plugins path-to-file.pdf`, and find new ones by searching GitHub for the hashtag #markitdown-plugin, with the sample implementation in packages/markitdown-sample-plugin. The markitdown-ocr plugin adds OCR to the PDF, DOCX, PPTX and XLSX converters by reading text out of embedded images with LLM Vision, reusing the same llm_client and llm_model settings used for image descriptions, and it pulls in no new ML or binary dependencies. A run with it enabled looks like this.

```python
from markitdown import MarkItDown
from openai import OpenAI

md = MarkItDown(
    enable_plugins=True,
    llm_client=OpenAI(),
    llm_model="gpt-4o",
)
result = md.convert("document_with_images.pdf")
print(result.markdown)
```

Then read the sentence that follows those examples, because it is the failure mode: if no llm_client is provided the plugin still loads, but OCR is silently skipped and the standard built-in converter is used instead. No exception, no warning, exit code zero. On a scanned PDF that means Markdown with the page text missing, and a downstream index that stores the document and never learns the text was never read.

## YouTube URLs and ZIP archives make it a fetcher as well as a parser

The supported input list is longer than the name suggests: PDF, PowerPoint, Word, Excel, images with EXIF metadata and OCR, audio with EXIF metadata and speech transcription, HTML, text-based formats such as CSV, JSON and XML, ZIP files where the converter iterates over the contents, YouTube URLs, and EPub. Two of those turn a local converter into something with network reach. A YouTube URL pulls a remote resource, which is why the [youtube-transcription] extra exists and why the security note is about sanitizing inputs rather than only about file paths. ZIP handling is the quieter case: the converter iterates over the contents, so what you hand it is a container of formats, including formats whose extra you never installed. The repository documents no recursion limit, no per-entry size cap, and no behaviour for an entry inside the archive that cannot be converted, so an archive in a batch job is a case where you should choose the failure policy yourself rather than inherit one.

## Azure Content Understanding is the only video path and the only field extractor

Two extras point at Microsoft cloud services, and what they add is stated plainly. Install the content understanding extra with `pip install 'markitdown[az-content-understanding]'`. For video it is the only option, because the built-in converters have no video support at all and only basic audio transcription. For structure it is also the only option: prebuilt or custom-built analyzers extract domain specific fields such as invoice amounts, receipt dates and contract clauses, and serialize them as YAML front matter, while neither the built-in converters nor the Document Intelligence integration expose fields. It also does cloud layout analysis and OCR for scanned PDFs, complex tables and multi-page documents, which is exactly where the built-in path is weakest. The cost of that quality is locality. A pipeline needing fields from invoices or needing video has two routes, send the document to the cloud analyzer or write your own extraction over the built-in Markdown, and the repository publishes no accuracy figures for either, so the decision is about where data goes rather than about output quality.

## textract is the named alternative, and structure is the difference

The project names its comparison point in the opening paragraph: MarkItDown is most comparable to textract, and the difference is given in the same sentence, a focus on preserving important document structure and content as Markdown, including headings, lists, tables and links. The output target is narrower than that sounds. The documentation says the Markdown is meant to be consumed by text analysis tools and may not be the best option for high fidelity conversion for human consumption, while the argument for Markdown at all is that it sits close to plain text, carries minimal markup, is what models such as GPT-4o natively produce, and is token efficient. Read that as a contract: you are getting something to hand to a model or an index, not a document to publish. If the requirement is a faithful PDF for a person to read, this project tells you plainly that it is the wrong tool, and the extractor it points you to is built on a different idea.

## Version 0.1.8 on Python 3.10 through 3.14, so pin it

The release line is still pre 1.0 and it moves through betas: v0.1.8b1 on 2026-09-04, v0.1.8b2 on 2026-09-14 and v0.1.8 on 2026-09-21, with the last push on 2026-09-21 and the repository not archived. None of that promises a stable command line, so a script built around -o, --list-plugins or --use-plugins should pin the version it was written against and a container tag should name a version rather than latest. The interpreter constraint is the other half, since the project requires Python 3.10 through 3.14 and rules out both 3.9 and 3.15. A virtual environment is recommended, and the activation step is the POSIX form in all three documented cases.

```bash
python -m venv .venv
source .venv/bin/activate
```

The code is MIT licensed, which makes reuse easy; the interface is not, which is why the version belongs in your lock file. If you prefer working from a clone, the source path is three lines and it installs the same extras in editable mode from the packages/ directory.

```bash
git clone git@github.com:microsoft/markitdown.git
cd markitdown
pip install -e 'packages/markitdown[all]'
```

## Conclusion

Use MarkItDown when the destination is a model or a text index, when you can pin a version, and when you can say which formats actually need to work. Do not use it for a faithful document for human readers, and do not feed it untrusted paths or URLs from a privileged process, because it runs with the caller's access. Verify three things before you trust a pipeline: that a scanned PDF actually produced its text rather than skipping OCR, that exiftool and ffmpeg exist on the machine if you need EXIF or audio, and that the exact version you pinned is the one whose convert_stream() or convert_local() you reviewed.

## FAQ

### What is MarkItDown used for?

It is a lightweight Python utility that converts files to Markdown for use with LLMs and related text analysis pipelines, preserving structure such as headings, lists, tables and links. It handles PDF, PowerPoint, Word, Excel, images, audio, HTML, CSV, JSON, XML, ZIP files, YouTube URLs and EPub, and the documentation says the output may not be the best option for high fidelity conversion for human consumption.

### how to install markitdown

Python 3.10 through 3.14 is required and a virtual environment is recommended. Install the whole thing with `pip install 'markitdown[all]'`, or a subset such as `pip install 'markitdown[pdf, docx, pptx]'`, and use uv pip install rather than pip install when you created the environment with uv.

### how to use markitdown in python

Import MarkItDown from the markitdown package, pass any llm_client and llm_model you want for image descriptions, then call convert() and read result.markdown. In untrusted environments the documentation asks for the narrowest convert_* function, naming convert_stream() and convert_local().

### how to use markitdown

From a shell, redirect the output with `markitdown path-to-file.pdf > document.md`, pass `-o` to write the file directly, or pipe content in with `cat path-to-file.pdf | markitdown`. Third-party plugins are disabled by default and are turned on per run with the --use-plugins flag after checking what is present with --list-plugins.

### Is MarkItDown free?

The code is published under the MIT license. Two optional extras, az-doc-intel and az-content-understanding, integrate with Microsoft Azure services instead of running locally, and the content understanding extra is the only option for video and for structured field extraction.

## Sources

- [Official README](https://github.com/microsoft/markitdown#readme)
- [Project repository](https://github.com/microsoft/markitdown)
- [Release notes](https://github.com/microsoft/markitdown/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/microsoft-markitdown
