mrmps/pdf2md: a browser-side PDF to Markdown converter built on @opendocsg/pdf2md
Browser based tool to convert PDFs to Markdown
At a glance
- What is it?
- mrmps/pdf2md wraps the @opendocsg/pdf2md library in a Next.js interface so conversion runs in the browser tab. It suits people who cannot upload documents to a server, and it is the wrong tool for scanned PDFs and image-heavy layouts.
- Who is it for?
- Adopt mrmps/pdf2md if your PDFs are text-based, under roughly 10MB, and you cannot send them to a server; the README states processing stays in the browser and the code is MIT licensed. Do not adopt it for scanned pages or image-heavy layouts, because the README says images and complex layouts may not be fully preserved.
- Can I use it commercially?
- Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
- Is it still maintained?
- Yes. The repository last received commits 88 days ago.
- What is it written in?
- Mainly TypeScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Who mrmps/pdf2md is for, and the problem it removes
Most PDF to Markdown converters are services. You upload a file, a server parses it, and you download the result. That is fine for a published paper and unacceptable for a contract, a medical record, or an internal specification that cannot leave the machine. mrmps/pdf2md takes the opposite approach: the conversion happens inside the browser tab, and the README states plainly that no files are uploaded or stored and that nothing is saved after you close the tab.
The project is a Next.js application, not a library. You run it yourself, or you use the hosted instance at pdftomarkdown.co. The README lists developers, writers, and anyone who works with Markdown as the intended audience, and the repository topics include rag, which points at the other common reason to want this: turning a pile of PDFs into text you can feed to a retrieval pipeline without shipping the originals to a third party.
The trade-off is honest and visible. Because the work happens in the tab, the ceiling is your browser and your machine. The README suggests keeping files under 10MB and warns that larger files may slow down the browser. A server-side converter has no such ceiling. This one buys privacy with a size limit.
How the conversion actually runs: Next.js shell, @opendocsg/pdf2md inside
The dependency list makes the architecture clear. The app depends on @opendocsg/pdf2md, listed as "latest" in package.json, and the README describes the flow in four steps: upload the PDF, convert with that library, display the Markdown, then download or copy it. There is no API route, no queue, no worker service in the described flow. The parsing library is bundled into the client.
The rest of the dependency tree is a standard shadcn-style Next.js interface: a long list of @radix-ui primitives, tailwindcss with @tailwindcss/typography, lucide-react for icons, and @vercel/analytics. The repository layout confirms it, with app/, components/, hooks/, lib/, styles/ and a docker/ directory.
Two details are worth flagging. First, pinning @opendocsg/pdf2md to "latest" means a fresh install can pull a different parser than the one the interface was written against; that is a reproducibility risk in a project whose whole output quality depends on that one library. Second, the README claims offline capability once the app is loaded, which follows from client-side parsing but also means the first load has to fetch the bundle. The README does not describe a service worker, so treat the offline claim as applying to an already-open tab.
Installing mrmps/pdf2md and converting your first PDF
The README gives clone, install, dev server, open localhost. The package manager is your choice; the examples cover pnpm, npm and yarn. Node and a package manager are the only prerequisites the README names.
git clone https://github.com/mrmps/pdf2md
cd pdf2md
pnpm installThen start the development server. The build script in package.json is next build and the production start script is next start, so the same project can be deployed rather than only run locally.
pnpm devOpen http://localhost:3000. You should see the converter interface. Drag a text-based PDF onto it, or select one with the file picker. The README says the result appears instantly, with buttons to download the Markdown file or copy it to the clipboard.
npm install
npm run devThe npm path is identical; the README lists it as an alternative, not a separate setup. There is a docker/ directory in the repository, but the README documents no Docker command, so do not assume a published image exists. If you want a container, read that directory before planning a deployment around it.
What the Markdown output preserves, and where it breaks
The README is specific about the happy path: headings, paragraphs, lists, tables, bold and italic text, and links when possible. It is equally specific about the failure case, stating that complex layouts, images, and advanced formatting may not be fully preserved. That sentence is the most important one in the document.
A two-column academic paper is a complex layout. A scanned page is not text at all; it is an image of text, and nothing in the described pipeline performs OCR. A datasheet whose meaning lives in a figure will come out as Markdown with the figure missing. If your source material looks like any of those, this tool will produce a plausible-looking file with quietly missing content, which is worse than an obvious failure.
The 10MB guidance compounds this. The README frames it as a performance recommendation rather than a hard limit, so a large file may still convert, just slowly and with the tab unresponsive while it works. There is also no documented batch mode. The described interface takes one file at a time, so converting a directory of PDFs means repeating the drag-and-drop for each one. For a RAG ingestion job with hundreds of documents, that is the point at which you want the underlying library directly rather than this interface.
The alternative: calling @opendocsg/pdf2md yourself
The most meaningful alternative is not a different product. It is the dependency this project is built on. @opendocsg/pdf2md is the parser doing the work; mrmps/pdf2md is the interface, the styling, the download button and the privacy story around it. If you need to convert a folder of PDFs on a schedule, or inside a Node process, the interface is the part you do not want.
That difference matters in both directions. Using the library directly gives you scripting, batching and control over the version you run. It also means you handle the file reading, the output writing and the environment. Using this app gives you a working interface in one command and a guarantee that the document never leaves the browser, at the cost of doing everything by hand, one file at a time, in a tab.
For a single confidential document, the app is the shorter path. For a pipeline, the library is. The README treats the library as an implementation detail, which is fair for its audience but worth knowing before you adopt the app for bulk work.
Maintenance, licence and what an upgrade costs you
The repository is not archived, and the last push was on 2026-07-04. The README states the project is MIT licensed and points at a LICENSE file, which permits commercial and private use with the usual attribution requirement; that is a plain statement about the licence text, not legal advice, and you should read the file yourself if the distinction matters to your organisation.
The maintenance picture has one real cost centre: @opendocsg/pdf2md is pinned to "latest". Any install can resolve to a newer parser than the one the interface was tested against, and any change in that library's output quality lands in your build without a version bump on this side. If you self-host, pin the dependency in your own lockfile and treat parser upgrades as something to check against a set of sample PDFs.
The interface layer is low-risk to upgrade. It is a Next.js app with Radix components and Tailwind, and the README documents no database, no migrations and no server state, so there is nothing to migrate when you rebuild. The upgrade cost sits almost entirely in the parser and in the layout of the pages you feed it.
Editorial conclusion
Adopt mrmps/pdf2md if your PDFs are text-based, under roughly 10MB, and you cannot send them to a server; the README states processing stays in the browser and the code is MIT licensed. Do not adopt it for scanned pages or image-heavy layouts, because the README says images and complex layouts may not be fully preserved. Before relying on it, clone the repo, run pnpm dev, and convert one of your own documents to see whether headings, lists and tables survive.
Frequently asked questions
Is it safe to use PDF converters like mrmps/pdf2md?
With mrmps/pdf2md the README states that all processing happens in the browser, that no files are uploaded or stored, and that nothing is saved after you close the tab. That claim applies to the app itself; if you use a different hosted converter, its data handling is a separate question.
Is there a file size limit in mrmps/pdf2md?
The README recommends keeping files under 10MB for best performance and notes that larger files may slow down your browser. It presents this as guidance rather than a hard cap.
Does mrmps/pdf2md preserve images and tables from my PDF?
Tables are listed among the supported formatting, but the README states that complex layouts, images, and advanced formatting may not be fully preserved. Test one of your own documents before trusting the output.
Which browsers does mrmps/pdf2md support?
The README lists Chrome, Firefox, Edge, Safari and Opera as supported browsers.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/mrmps-pdf2md)