Model or dataset
ibrahimqureshae/mdflux avatar
ibrahimqureshae/mdflux

MDFlux review: a local-first desktop app that turns scanned PDFs and office files into Markdown

Turn any document into clean, AI-ready Markdown. Local-first desktop app: reads scanned PDFs, batches folders, runs offline, and uses far fewer tokens than vision models.

541 stars39 forksPythonMIT

At a glance

What is it?
MDFlux wraps Microsoft's MarkItDown in a Tauri desktop shell, adds OCR for image-only PDFs and batch folder conversion, and claims up to 6x fewer tokens than sending pages to a vision model. Here is what the repository actually documents, and where it stops.
Who is it for?
Adopt MDFlux if you are converting documents on Windows or Linux x64 and want the OCR and batch work to stay on your own machine, and start with the Full build so nothing is fetched on first launch. Do not adopt it if you need macOS, a headless server pipeline, or a library you can import into your own Python code, because the project ships as a desktop application and its README lists only Windows and Linux (x64) platforms.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 23 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The gap MDFlux fills: scanned pages that extract as nothing

Plain text extraction fails silently on image-only PDFs. The README makes this the centre of its argument: point a plain extractor at a scanned document and, in its own words, you get zero usable text, because the characters are pixels rather than an embedded text layer. The alternative people reach for is rendering each page as an image and sending it to a vision model, which works but means the document leaves your machine and you pay per page in image tokens.

MDFlux sits between those two. It is a desktop application, built with Tauri and Svelte according to the repository topics, with a Python core around Microsoft's MarkItDown. The audience is narrow and identifiable: engineers and analysts who need their own documents in Markdown for a RAG index or a prompt, who cannot or will not upload those documents, and who have a folder of scans rather than a single clean PDF. If your documents already have a text layer and you only need one file converted once, the OCR machinery is dead weight.

How MDFlux converts a file: MarkItDown plus an OCR pass

The pipeline is documented at a high level rather than in detail. A file or folder goes in, MarkItDown handles the formats it already understands, and an OCR stage recovers text from pages where extraction returns nothing. A cleanup pass is optional and, per the README, exists to strip broken layout, repeated whitespace and junk characters from messy extractions. Output is Markdown.

The format list is broad: PDF, DOCX, PPTX, XLSX, EPUB, HTML, CSV, JSON, XML, TXT, MD, images and audio. Because MarkItDown is the engine, the non-PDF formats are largely inherited behaviour rather than something MDFlux wrote. The repository does not publish an architecture document, so the boundary between MarkItDown, the OCR component and the cleanup pass is not visible from the README. That matters if you plan to debug a bad conversion: you will not know from the documentation which stage dropped your table.

The token argument is the project's main claim. The README states that ordinary documents come out around 4 times lighter than sending pages as images, and up to 5.7 times lighter on scanned pages, and it gives a worked comparison for one scanned, image-only PDF: 0 usable tokens from a plain extractor, 10,731 from a vision model, 1,893 from MDFlux OCR to Markdown. Those are the project's own numbers, not an independent measurement, and the README does not describe the corpus or the method behind them. Treat the direction as plausible and the exact multiplier as unverified.

Installing MDFlux and converting your first folder

MDFlux is distributed as a release download, not as a package on PyPI or a container image. The README points at the latest release page and offers two editions per platform. Full bundles everything needed for conversion and OCR, so nothing is downloaded when you first open the app. Lite is a small first download (the README gives 4.9 MB for Windows and 5.8 MB for Linux) that fetches the required components on first launch.

The repository lists Windows and Linux x64 builds. The Linux build is the glibc variant, and the release file names carry the version. The README's own download links are the only distribution channel it documents; the release page is where you pick the archive for your platform. After downloading, the README says to verify the file against the instructions in docs/cross-platform/releases-and-verification.md before running it. That step is worth taking: the Full archive is the large one, and the project publishes a NOTICE and THIRD-PARTY-LICENSES.md at the repository root, which tells you the bundle contains third-party components whose terms you inherit.

There is no documented CLI in the README. The workflow it describes is graphical: drop in a file or a folder, get Markdown back. Batch processing is a property of the app, not a flag you pass. A command-line invocation of the kind below is not something the README or the repository files show, so do not expect it to exist:

bash
# Not documented: the README shows no CLI entry point for MDFlux.

For a first real use, pick a directory of scanned PDFs rather than a single clean file. That exercises the OCR path, which is the reason to choose this project over a plain extractor, and it shows you the per-file output quality before you point it at a large batch.

Where MDFlux is the wrong tool

Platform coverage is the first limit. The README lists Windows and Linux x64 only. There is no macOS build in the download section, so Mac users are out unless they build from source, and the README does not describe that path. ARM Linux is likewise unaddressed.

The second limit is headless use. MDFlux is a desktop app with a graphical workflow. There is no documented command line, no server mode and no HTTP API in the README. If your conversion step lives inside a CI job, a cron task or a container, this project does not fit that shape, and wrapping a GUI binary in a pipeline is not something the documentation supports.

The third is reproducibility of the numbers. The 4x and 5.7x token figures come from the project itself, with no stated methodology. If you are building a cost model for a RAG pipeline, measure your own documents rather than adopting those multipliers. And if your PDFs already carry a clean text layer with intact tables, MarkItDown alone will get you most of the way, and the OCR stage will not earn its download size.

MDFlux compared with MarkItDown and with the vision-model route

MDFlux is explicit that it is built on Microsoft's MarkItDown, and the README frames the project as the parts that make that engine usable day to day: OCR, batch folders, a cleanup pass and a desktop interface. The practical difference is that MarkItDown is a Python library you call from your own code, while MDFlux is an application you run. If you need to embed conversion in a service, call MarkItDown directly and add your own OCR; you give up the packaged OCR and the cleanup pass, and you gain control over every stage.

The other comparison is the vision-model route, and here the difference is architectural rather than cosmetic. Sending page images to a vision model keeps the document in a format the model reads as pixels, which the README argues is the expensive way to spend tokens, and it requires uploading the document. MDFlux converts locally and hands the model text. The trade-off is that OCR output is not perfect: a vision model can use layout and visual context that a text-layer OCR pass may not reconstruct, so for documents where table geometry or figure captions carry meaning, the cheaper Markdown can be the weaker input. The README does not compare output accuracy against a vision model, only token counts.

Maintenance, releases and what the MIT licence means here

The last push to the default branch was on 2026-09-07, and the repository is not archived. Releases are recent and versioned: v0.1.0 on 2026-06-22, v0.2.0 on 2026-08-19, and v0.3.0 on 2026-08-23. That cadence over roughly two months suggests the project is moving, though it says nothing about how long that will continue, and the release history is short enough that no long-term support pattern exists yet.

The upgrade cost is the download itself. Because MDFlux is distributed as a full or lite archive per platform rather than through a package manager, upgrading means fetching a new archive and replacing the old one. There is no documented auto-update mechanism in the README, and no migration notes, so a version bump is a manual operation. The CHANGELOG.md at the repository root is where the project records what changed between versions; read it before replacing a working build.

The project is MIT-licensed. That covers MDFlux's own code, but the repository ships NOTICE and THIRD-PARTY-LICENSES.md, which means the Full bundle includes components under other terms. The OCR engine and any bundled models are the likely candidates. If you redistribute the Full archive inside your organisation or in a product, read those two files rather than assuming MIT covers the whole download. This is not legal advice; the files are the source.

Editorial conclusion

Adopt MDFlux if you are converting documents on Windows or Linux x64 and want the OCR and batch work to stay on your own machine, and start with the Full build so nothing is fetched on first launch. Do not adopt it if you need macOS, a headless server pipeline, or a library you can import into your own Python code, because the project ships as a desktop application and its README lists only Windows and Linux (x64) platforms. Before you commit, verify the release checksum against docs/cross-platform/releases-and-verification.md and confirm the bundled OCR model covers the languages in your documents, since the README does not state which languages ship with the Full download.

Frequently asked questions

How can I convert anything to Markdown with MDFlux?

Download the Full build for Windows or Linux x64 from the releases page, open the app, and drop in a file or a folder. The README states that MDFlux handles PDF, DOCX, PPTX, XLSX, EPUB, HTML, CSV, JSON, XML, TXT, MD, images and audio, with OCR for scanned pages and optional cleanup of messy extraction.

Does MDFlux work offline?

Yes. The README describes it as local-first and says it runs entirely on your machine, and the Full download is presented as containing everything needed for document conversion and OCR so there are no additional setup downloads on first open. The Lite edition is the exception: it downloads required components when you first launch it.

Does MDFlux run on macOS?

The README's download section lists Windows and Linux (x64) builds only, and the Linux build is the glibc variant. No macOS build is documented, and the README does not describe a build-from-source path for it.

Is there a command line or API for MDFlux?

The README documents a desktop workflow: drop in a file or folder and get Markdown back. It does not document a CLI, a server mode or an HTTP API, so headless and pipeline use is not covered by the project's own instructions.

Official sources

  1. ibrahimqureshae/mdflux on GitHub
  2. License: MIT
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/ibrahimqureshae-mdflux.svg)](https://hysenlabs.com/projects/ibrahimqureshae-mdflux)