# markdownify-mcp: a Markdown conversion layer for MCP clients

> The zcaceres/markdownify-mcp server wraps markitdown and repomix behind MCP tools so an assistant can turn PDFs, Office files, audio and web pages into Markdown. It is convenient, but the Docker image ships a reduced feature set and the README is thin on failure handling.

**zcaceres/markdownify-mcp** — A Model Context Protocol server for converting almost anything to Markdown.

- Repository: https://github.com/zcaceres/markdownify-mcp
- Stars: 2,997 · Forks: 255
- Language: TypeScript
- License: MIT
- Published: 2026-08-04 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/zcaceres-markdownify-mcp

## What markdownify-mcp solves, and for whom

MCP clients can read files and call functions, but they generally cannot open a PDF, a slide deck or an audio recording and return usable text. markdownify-mcp fills that gap. It is a Model Context Protocol server that, in the README's words, "converts various file types and web content to Markdown format", exposing each conversion as a separate tool: `pdf-to-markdown`, `docx-to-markdown`, `xlsx-to-markdown`, `pptx-to-markdown`, `image-to-markdown`, `audio-to-markdown`, `youtube-to-markdown`, `bing-search-to-markdown`, `webpage-to-markdown` and `get-markdown-file`.

The audience is narrow and specific. You need an MCP-capable desktop client, Node or Bun on the machine, and a reason to want Markdown rather than plain text. A developer who drops a vendor PDF into a chat window and asks for a summary is the intended user. Someone who wants a general-purpose document conversion service is not: there is no HTTP endpoint here, no queue, and no job history. The server is a local process that the client launches over stdio.

## How the TypeScript server delegates to markitdown and repomix

The architecture is deliberately thin. `src/server.ts` holds the MCP wiring and `src/tools.ts` holds the tool definitions; the package depends on `@modelcontextprotocol/sdk`, `zod`, `private-ip` and `repomix`. Conversion itself is not implemented in TypeScript. The `preinstall` script creates a Python virtual environment at `.venv` and installs `markitdown[all]`, and the tools shell out to that executable. `git-repo-to-markdown` uses `repomix` from `node_modules/.bin`.

That delegation explains the environment variables. `MARKITDOWN_PATH` defaults to `<project>/.venv/bin/markitdown` and then to `markitdown` on `PATH`; `REPOMIX_PATH` follows the same pattern for the repomix binary. Both exist because the server assumes a layout that a system-wide install breaks. If you installed markitdown with `pipx install "markitdown[pdf]"`, the README says to point `MARKITDOWN_PATH` at that executable instead.

`private-ip` is the other interesting dependency. A server that fetches arbitrary URLs on request is an obvious request-forgery risk, and its presence suggests the web-fetching tools check resolved addresses before connecting. The README does not document that behaviour, so treat it as an inference from the dependency list rather than a guarantee.

## Installing markdownify-mcp and running a first conversion

The README gives a four-step local install. Clone the repository, then run the dependency install, which triggers the `preinstall` hook and builds the Python environment:

```bash
bun install
```

Next build the TypeScript and start the server. `bun run build` runs `tsc` and then marks the output executable; `bun start` runs `dist/index.js`. You should end up with a `dist/` directory containing `index.js`.

```bash
bun run build
bun start
```

For a desktop client, the configuration is a command and an absolute path. The README uses `node` here even though the start script uses Bun, so the client must be able to find `node` on its own PATH.

```js
{
  "mcpServers": {
    "markdownify": {
      "command": "node",
      "args": ["{ABSOLUTE PATH TO FILE HERE}/dist/index.js"]
    }
  }
}
```

After restarting the client, its tool list should include the conversions above. The first real use is to hand `pdf-to-markdown` a path to a local PDF and ask for a summary. For a bounded setup, restrict what the server can read:

```bash
MD_ALLOWED_PATHS=/data/in:/data/out bun start
```

With that set, every file-input tool rejects paths outside those directories. The README notes `MD_SHARE_DIR` as a deprecated single-directory alias that is still honored, so new setups should prefer `MD_ALLOWED_PATHS`.

## The Docker image is slimmer than the local install

The repository ships a multi-stage Dockerfile. The base stage installs `python3`, `python3-venv`, `bash` and `git` on top of `oven/bun:debian`, deletes `.python-version`, then creates `.venv` and installs `"markitdown[pdf]>=0.1.5"`. A builder stage runs `bun install` and `bun run build`; the runner stage installs production dependencies and copies `dist` from the builder. The README states the split saves about 100MB.

The consequence matters more than the size. The published image installs the `[pdf]` extras only, so `audio-to-markdown` and `image-to-markdown` fail inside it. Audio transcription and image OCR need `markitdown[all]`, which the local `bun install` provides. If those tools are part of your workflow, the container is the wrong deployment.

The README also flags a path trap for the Docker MCP catalog entry `mcp/markdownify`: mount host directories into the container and pass container paths to the tools, so `/data/foo.pdf` rather than `/Users/you/Documents/foo.pdf`. Set `MD_ALLOWED_PATHS` to the colon-separated list of mounted directories so the read boundary matches the bind mount.

```sh
docker build -t markdownify-mcp .
docker run --rm -i \
  -v "$HOME/Documents:/data:ro" \
  -e MD_ALLOWED_PATHS=/data \
  markdownify-mcp
```

## Where markdownify-mcp is the wrong tool

The server has no documented retry, timeout or partial-failure behaviour. If markitdown cannot parse a malformed PDF, the README does not say what the tool returns: an error, an empty document, or a best-effort extraction. For interactive use that is tolerable, because a human sees the result and tries again. For an automated pipeline that ingests hundreds of documents, it is a real gap, and the README does not document rollback or a dry-run mode either.

Output quality is inherited, not controlled. Everything routes through markitdown, so table fidelity in XLSX files, layout in multi-column PDFs and slide ordering in PPTX files are markitdown's problem, not this project's. The repository adds tool names and a transport, not a converter.

`get-markdown-file` is narrower than its name suggests: the README says the file extension must end with `*.md` or `*.markdown`. It reads existing Markdown, it does not convert anything, and it will reject a `.txt` file that already contains Markdown.

Finally, the security model is opt-in. `MD_ALLOWED_PATHS` is unset by default, which the README describes as unrestricted. A client with a prompt-injection vector and an unrestricted file-reading tool is a poor combination. The default is convenient for a laptop and wrong for a shared machine.

## markdownify-mcp compared with calling markitdown directly

The obvious alternative is markitdown itself, the Python package this project wraps. Calling it directly means one command per file and no MCP layer, and it exposes the full option set of the underlying library rather than the subset the MCP tools surface. What it does not do is let an assistant decide, mid-conversation, that a file needs converting. With markdownify-mcp the model chooses the tool and the path; with markitdown the human decides in advance.

Repomix is the second reference point, and it is not a competitor so much as a component: the package depends on `repomix` and the README describes `git-repo-to-markdown` as using the `repomix` executable, with `REPOMIX_PATH` to locate it. If your only need is packing a repository into a single Markdown file, install repomix on its own and skip the MCP server entirely. The search phrase "Markdownify vs markitdown" points at the same distinction: this project is the integration, markitdown is the engine.

## Maintenance, licensing and upgrade cost

The repository is not archived, and the last push was on 2026-05-01, which is also the date of release v1.1.0. Before that, v1.0.4 landed on 2026-04-17 and v1.0.3 on 2026-04-02, so the three most recent releases span about a month. The project has moved recently, but the cadence visible in those three releases is not a promise about the next six months.

Upgrading carries a real cost because the dependency graph is split across ecosystems. `package.json` pins `@modelcontextprotocol/sdk` at `^1.27.1`, `repomix` at `^1.12.0`, `zod` at `^4.3.6` and `private-ip` at `^3.0.2`, while `pyproject.toml` requires Python 3.11 or later and `markitdown[all]>=0.1.5`. A change in the MCP SDK or in markitdown can require a rebuild of the virtual environment, not just a `bun install`. The Dockerfile deletes `.python-version` before creating the venv, which is a workaround worth remembering if you build your own image.

Licensing is straightforward: the project is MIT, and the LICENSE file is in the repository root. That covers this code. It does not cover markitdown, repomix or any content you convert, and MIT says nothing about the rights attached to the documents you feed through the tools. Check those separately.

## Conclusion

Adopt markdownify-mcp if you already run an MCP-capable client and want file conversion exposed as tools rather than as a separate CLI step. Skip it if you need a stable HTTP API, batch conversion at scale, or audio and OCR inside a container, because the published image installs markitdown[pdf] only and those tools fail there. Before wiring it in, confirm the absolute path to dist/index.js, set MD_ALLOWED_PATHS to a directory you are willing to expose, and check whether your install needs markitdown[all] rather than the slim extras.

## FAQ

### What does an MCP actually do?

In this project MCP is the protocol that lets a desktop client launch the server and call its tools, such as pdf-to-markdown or webpage-to-markdown. The README lists the available tools and shows an mcpServers entry that starts dist/index.js.

### Does ChatGPT understand Markdown?

The README does not discuss ChatGPT or any specific client. It describes the server as a Model Context Protocol server that converts files and web content to Markdown, and gives a desktop app configuration example.

### What is Markdown mainly used for?

The README frames Markdown as the readable, shareable output format for converted content, which is why the tools turn PDFs, images, audio, DOCX, XLSX, PPTX, YouTube transcripts, Bing results and web pages into it. get-markdown-file also retrieves existing .md or .markdown files.

### Can I convert MD to PDF?

No. The README describes conversion into Markdown only, and get-markdown-file retrieves existing files whose extension ends with *.md or *.markdown. There is no PDF-writing tool in the list.

## Sources

- [Official README](https://github.com/zcaceres/markdownify-mcp#readme)
- [Project repository](https://github.com/zcaceres/markdownify-mcp)
- [Release notes](https://github.com/zcaceres/markdownify-mcp/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/zcaceres-markdownify-mcp
