Library / SDK
microsoft/markitdown avatar
microsoft/markitdown

MarkItDown: converting office files to Markdown for LLM pipelines

Convert files and office documents to Markdown with Python.

187,295 stars13,835 forksPythonMIT

At a glance

What is it?
Microsoft's MIT-licensed Python utility turns PDFs, Office documents, images, audio and more into Markdown. It is built for text analysis pipelines, not for high-fidelity human-facing conversion, and the README says so.
Who is it for?
Adopt MarkItDown if you need a permissively licensed Python converter that emits Markdown for a text pipeline and you can control the input files. Do not adopt it if you need high-fidelity output for human readers, or if you must process untrusted uploads without sandboxing, since the README states conversion runs with the privileges of the current process.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 8 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.

DEEP OPEN-SOURCE ANALYSIS

The problem MarkItDown addresses in an LLM pipeline

Feeding a PDF or a spreadsheet to a language model means turning a binary container into text first. Most converters in that space target human readers: they render a facsimile of the page. MarkItDown targets the opposite consumer. The README frames it as "a lightweight Python utility for converting various files to Markdown for use with LLMs and related text analysis pipelines", and explicitly positions it against textract while claiming a focus on preserving document structure such as headings, lists, tables and links.

The intended user is a developer building a retrieval or summarisation pipeline who wants one call that accepts a file and returns Markdown. The README is direct about the trade-off: the output is "often reasonably presentable and human-friendly" but "may not be the best option for high-fidelity document conversions for human consumption." If your goal is a pixel-faithful rendering of a scanned contract for a legal reviewer, this is the wrong tool by the author's own description.

How conversion works: one entry point, format-specific converters, optional extras

The architecture is a dispatcher over per-format converters. You construct a MarkItDown object and call convert() on a path, or you call a narrower function such as convert_stream() or convert_local(), which the README recommends for untrusted input. The object inspects the input and routes it to the converter that handles that format.

The formats listed in the README are PDF, PowerPoint, Word, Excel, images (EXIF metadata and OCR), audio (EXIF metadata and speech transcription), HTML, text-based formats including CSV, JSON and XML, ZIP files (iterated over their contents), YouTube URLs and EPubs, with the list left open-ended. Dependencies for those formats are optional extras rather than a single monolithic install, so the dependency footprint is a choice you make at install time.

Plugins are a second extension path. They are disabled by default, listed with markitdown --list-plugins, and enabled per invocation with --use-plugins. The repository ships a sample plugin under packages/markitdown-sample-plugin, and the README points developers to the hashtag #markitdown-plugin on GitHub to find published ones. The markitdown-ocr plugin is documented in the README as adding OCR to the PDF, DOCX, PPTX and XLSX converters by reusing the llm_client and llm_model pattern already used for image descriptions, with no new ML libraries or binary dependencies. Notably, the README states that if no llm_client is provided the plugin still loads but OCR is silently skipped and the standard built-in converter is used. Silent fallback is a design decision worth knowing about: a misconfigured pipeline will produce output that looks fine and simply lacks the extracted text.

Installing MarkItDown and converting your first PDF

MarkItDown requires Python 3.10 or higher, and the README recommends a virtual environment. Any of the three environment recipes in the README works; the standard one is:

bash
python -m venv .venv
source .venv/bin/activate

With the environment active, install the package. The README gives pip install 'markitdown[all]' for every optional dependency at once:

bash
pip install 'markitdown[all]'

If you only care about a subset of formats, the README shows the extras being combined, for example pip install 'markitdown[pdf, docx, pptx]' to pull in only the PDF, DOCX and PPTX dependencies. The full list of extras in the README is [all], [pptx], [docx], [xlsx], [xls], [pdf], [outlook], [az-doc-intel], [az-content-understanding], [audio-transcription] and [youtube-transcription].

From the command line, the README's first example redirects output to a file:

bash
markitdown path-to-file.pdf > document.md

You should see a Markdown file containing the document's extracted text and structure. The README also documents -o for writing directly to a path, and piping content in:

bash
markitdown path-to-file.pdf -o document.md
cat path-to-file.pdf | markitdown

In Python, the pattern is a MarkItDown instance plus convert(). The README's plugin example shows the shape, including the llm_client and llm_model arguments used for image descriptions:

python
from markitdown import MarkItDown
from openai import OpenAI

md = MarkItDown(
    enable_plugins=True,
    llm_client=OpenAI(),
    llm_model="gpt-4o",
)
result = md.convert("document_with_images.pdf")
print(result.text_content)

The converted document is exposed as result.text_content. If you are not using plugins or LLM-backed image descriptions, you can drop enable_plugins, llm_client and llm_model and call convert() on the file directly.

The security model is the caller's responsibility

The most important paragraph in the README is the warning at the top. MarkItDown "performs I/O with the privileges of the current process", and the README compares this to open() or requests.get(): it will reach whatever the process itself can reach. The stated mitigation is to sanitize inputs in untrusted environments and to call the narrowest convert_* function needed, naming convert_stream() and convert_local() as examples, with a pointer to a Security Considerations section of the documentation.

That is an honest disclosure rather than a solved problem. If you are building an upload endpoint that accepts arbitrary files from the public, the converter is not a sandbox and does not claim to be. The README's own advice is to narrow the call and sanitize the input, which in practice means the isolation has to come from your process boundary, not from the library. Teams that treat "it parses the file" as equivalent to "it is safe to parse the file" will get this wrong.

Where MarkItDown is the wrong choice

Two limits stand out. The first is output fidelity. The README states the output is meant to be consumed by text analysis tools and may not suit high-fidelity conversions for human consumption. If your downstream step is a person reading a converted document, a layout-preserving renderer will serve you better, and the README's own comparison to textract is about structure preservation for machines, not visual accuracy.

The second is that the project leans on optional dependencies and, for the harder formats, on cloud services. The README lists [az-doc-intel] and [az-content-understanding] as extras, and describes Content Understanding as providing higher-quality conversion with structured field extraction, multi-modal support and configurable analyzers. It states that Content Understanding is the only option for video, that built-in converters have no video support and only basic audio transcription, and that neither the built-in converters nor the Document Intelligence integration exposes extracted fields. So for scanned PDFs, complex tables and structured field extraction, the README itself points you off the local path and toward an Azure service. If you need a fully local pipeline with structured field extraction, MarkItDown's built-in converters do not offer it.

The truncated README also leaves some operational questions open. It does not document rollback behavior, version pinning strategy, or what happens when a converter encounters a malformed file. Those are gaps you would need to probe in code rather than in the documentation.

MarkItDown compared with Docling and textract

The README names textract as the closest comparison and states the difference plainly: MarkItDown focuses on preserving important document structure and content as Markdown, including headings, lists, tables and links. That is a narrower, more opinionated output target than textract's, which is a general extraction library.

Docling is the other name that comes up in searches about this project, and the README does not mention it, so any comparison here would be speculation. What can be said from the repository is structural: MarkItDown is a Python package installable from PyPI with per-format extras, a CLI entry point named markitdown, a plugin system that is off by default, and an optional Azure path for the formats its local converters handle least well. That combination, a small local core plus a documented cloud escape hatch, is the shape to weigh against any alternative you are considering. If you need a single self-contained local binary with no cloud option, the extra surface here is weight you are not using. If you need structured field extraction from invoices or receipts, the README's answer is Content Understanding, not the built-in converters.

Maintenance, licence and what a version bump costs

The repository is not archived. Its last push was on 2026-07-29, which is the same date as the v0.1.7 release, following v0.1.6 on 2026-05-26 and v0.1.5 on 2026-02-20. The version numbers are still in the 0.1.x range, which is worth factoring into how you pin it: a minor bump in a pre-1.0 package can carry behavior changes, and the README does not document a deprecation policy.

The licence is MIT, which permits commercial and closed-source use with the usual requirement to retain the copyright notice and permission notice. That is a permissive choice and consistent with the project being positioned as a general-purpose utility rather than a product. This is not legal advice; read the LICENSE file in the repository for the operative text.

Upgrade cost is dominated by the optional dependencies rather than the core package. Because formats are gated behind extras such as [pdf], [docx] and [az-content-understanding], a version bump can pull in new transitive dependencies for the formats you have enabled. Pinning the extras you actually use, rather than [all], keeps that surface smaller. The Dockerfile in the repository installs the packages from the local source tree with pip --no-cache-dir install /app/packages/markitdown[all] and runs as an unprivileged USER, which is a reasonable starting point if you want the converter behind a process boundary rather than in your application process.

Editorial conclusion

Adopt MarkItDown if you need a permissively licensed Python converter that emits Markdown for a text pipeline and you can control the input files. Do not adopt it if you need high-fidelity output for human readers, or if you must process untrusted uploads without sandboxing, since the README states conversion runs with the privileges of the current process. Before committing, verify three things yourself: that the optional extras for your formats install on your platform, that the output structure of your specific documents survives conversion, and which convert_* entry point your code path actually calls.

Frequently asked questions

What is MarkItDown used for?

It converts files such as PDF, PowerPoint, Word, Excel, images, audio, HTML, CSV, JSON, XML, ZIP archives, YouTube URLs and EPubs into Markdown. The README describes it as a lightweight Python utility for use with LLMs and related text analysis pipelines.

How does MarkItDown work?

You create a MarkItDown object and call convert() on a path, or call a narrower function such as convert_stream() or convert_local(). The object dispatches to a per-format converter, and format-specific dependencies are installed through optional extras.

Is MarkItDown free?

The repository is licensed under MIT, which permits commercial and closed-source use provided the copyright and permission notices are retained. Some optional features point at paid cloud services, such as the az-doc-intel and az-content-understanding extras, but the package itself is MIT-licensed.

How do I install MarkItDown?

It requires Python 3.10 or higher, and the README recommends a virtual environment. Install with pip install 'markitdown[all]' for every optional dependency, or name only the formats you need, for example pip install 'markitdown[pdf, docx, pptx]'.

How do I use MarkItDown in Python?

Construct a MarkItDown instance and call convert() on the file path; the converted text is available as result.text_content. The README's example also shows enable_plugins, llm_client and llm_model arguments, which are used for plugin support and LLM-backed image descriptions.

Official sources

  1. Official README
  2. Project repository
  3. Release notes
For maintainers

Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/microsoft-markitdown.svg)](https://hysenlabs.com/projects/microsoft-markitdown)
Community notes

Community notes