Pix2Text: A Python Image-to-Markdown OCR Toolkit and Free Mathpix Alternative
An Open-Source Python3 tool with SMALL models for recognizing layouts, tables, math formulas (LaTeX), and text in images, converting them into Markdown format. A free alternative to Mathpix, empowering seamless conversion of visual content into text-based representations. 80+ languages are supported.
At a glance
- What is it?
- Pix2Text chains layout analysis, table recognition, formula detection and OCR into a single Markdown pipeline for images and PDFs. It is a good fit for Python teams that need LaTeX output and can absorb the model download and GPU cost.
- Who is it for?
- Adopt Pix2Text if you are comfortable in Python and your documents mix prose, tables and formulas that must come out as Markdown with LaTeX inline. Skip it if you need a supported desktop product, a hosted API with an uptime commitment, or OCR for a script outside the 80+ language set, since the other-language path runs through EasyOCR rather than the CnOCR models used for English and Simplified Chinese.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 39 days ago.
- What is it written in?
- Mainly Jupyter Notebook, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The gap Pix2Text fills between OCR libraries and Mathpix
Most OCR libraries return text in reading order and stop there. A scanned page from a textbook or a paper is not a flat string: it has a column layout, a table whose cells must stay aligned, and inline formulas that are useless as plain characters. Pix2Text is built for that specific case. The README describes it as a free and open-source Python alternative to Mathpix that can already accomplish Mathpix's core functionality, and the output format it targets is Markdown, with formulas written as LaTeX.
The audience is narrow and specific. It is a Python3 toolkit, and the README says so plainly, adding that it may not be very user-friendly for those who are not familiar with Python. If you want a GUI, the project points to a free-to-use online web service instead. The repository is not archived, and the last push was on 2026-08-23, which is recent enough that the codebase is still moving.
Who benefits: anyone building a corpus out of PDFs, lecture notes, or photographed problem sets, where the alternative is hand-copying equations into LaTeX. Who does not: teams that need a stable, versioned API with a support contract, or that need to parse a language the underlying OCR engines do not cover well.
How the Pix2Text pipeline chains five models into Markdown
Pix2Text is not one model. The README lists the components it integrates, and the architecture diagram in docs/figs/arch-flow.jpg shows them as a flow rather than a single pass.
Layout analysis runs first, using breezedeus/pix2text-layout; the release notes for V1.1.2 say a DocLayout-YOLO model was integrated to improve layout analysis accuracy. Table regions are handled by breezedeus/pix2text-table-rec. Formula regions are found by a Mathematical Formula Detection model, breezedeus/pix2text-mfd-1.5, which the release notes for V1.1.4 describe as the new default, implemented on top of CnSTD. Those regions are then recognized by the Mathematical Formula Recognition model breezedeus/pix2text-mfr-1.5, which emits LaTeX. Everything else goes to a text recognition engine: CnOCR for English and Simplified Chinese, EasyOCR for the other languages in the 80+ set.
That decomposition is the design decision worth understanding. Each stage can be swapped or configured independently, which is why the CLI exposes separate flags for the analyzer, the text OCR config and the formula OCR config. It also means errors compound: a formula the MFD model never detects is a formula the MFR model never sees, and it will appear in your Markdown as garbled text rather than as a missing equation. The release notes for V1.1.1 frame the improved MFD models as a fix for formula detection accuracy, which suggests the project treats that stage as the weak link.
Installing Pix2Text and running a first prediction
The package is on PyPI under the name pix2text, and setup.py declares the console entry point p2t = pix2text.cli:cli, so the CLI is available as soon as the install finishes. The base install pulls in cnocr[ort-cpu], cnstd, torch, torchvision, transformers, PyMuPDF and doclayout-yolo, among others, so expect a large dependency tree.
pip install pix2textMultilingual support and the VLM backends are optional extras rather than part of the base install. The README gives the install command for the VLM path as pip install pix2text[vlm], and setup.py defines a multilingual extra that adds easyocr.
pip install pix2text[vlm]The Makefile contains the shortest working example in the repository. It runs prediction on docs/examples/mixed.jpg with English and Simplified Chinese, the MFD analyzer, a yolov7_tiny detector, and writes an annotated image to tmp-output.jpg.
p2t predict -l en,ch_sim -a mfd -t yolov7_tiny -i docs/examples/mixed.jpg --save-analysis-res tmp-output.jpgAfter it runs you should get Markdown on stdout and the annotated analysis image on disk. The Makefile also carries commented variants for Vietnamese (vi) and Traditional Chinese (ch_tra) that select explicit model file paths, which is a useful hint that per-language runs may need their own model configuration rather than just a different -l value.
Where Pix2Text breaks down or is the wrong tool
The first constraint is that the open-source models are not the best models the project has. The README states directly that the online web version uses the latest models, resulting in better performance compared to the open-source models. If you need the highest accuracy, the local install is not the path to it, and that is an unusual thing for a project to say about itself.
The second is the language split. English and Simplified Chinese go through CnOCR; everything else in the 80+ set goes through EasyOCR. Those are different engines with different error profiles, so a Vietnamese page and an English page are not processed by the same recognizer, and tuning that helps one will not necessarily help the other.
The third is the dependency surface. requirements.txt is generated by pip-compile and pins an index-url of mirrors.aliyun.com with extra indexes at pypi.tuna.tsinghua.edu.cn and pypi.org. That is fine for the project's own build, but it tells you the maintainers are working from a China-based network, and it is worth checking that your environment resolves the same versions before you assume a bug is in Pix2Text rather than in a transitive pin.
Finally, there is no documented rollback or downgrade procedure in the README. If a model upgrade changes your output, the README does not tell you how to pin back to the previous model version. You would be working from the release notes and the model repository names alone.
Pix2Text compared with Mathpix and with plain OCR libraries
The comparison the project invites is with Mathpix, and the README makes the claim explicitly: free and open-source, and able to accomplish Mathpix's core functionality. The difference in approach is not just price. Mathpix is a hosted service, so you send the image and get a result; Pix2Text runs the models locally, which means your documents do not leave your machine, but you own the hardware, the model downloads and the version pinning. The README's own admission that the web version outperforms the open-source models is the honest version of this trade-off: you are buying privacy and control with accuracy.
The second comparison is against a plain OCR library. A general OCR engine returns text boxes and strings. Pix2Text adds the layout, table and formula stages on top, and the value is entirely in those stages. If your documents are single-column prose with no math, the extra models are overhead you are paying for and not using. If they are dense with equations, the formula stages are the whole point.
A third option the repository itself points to is the VLM route. The release notes for V1.1.3 say Pix2Text supports VlmTableOCR and VlmTextFormulaOCR models based on the LiteLLM interface, allowing the use of closed-source VLM models, with examples in tests/test_vlm.py and tests/test_pix2text.py. That is a hybrid: local layout, remote recognition. It is worth knowing this exists before you assume the choice is binary.
Maintenance, licence and the upgrade cost you should budget for
The licence is MIT, declared in setup.py and shown by the license badge in the README. MIT is permissive: you can use Pix2Text in commercial products, modify it and redistribute it, provided the copyright notice and permission notice are preserved. That is a statement about the licence text, not legal advice, and if you are shipping it inside a product you should read the LICENSE file in the repository yourself.
The maintenance picture is active but not high-cadence. The last push was on 2026-08-23, and the release history shows v1.1.7 on 2026-08-23, v1.1.6 on 2026-02-07 and v1.1.5 on 2026-02-07. Two releases on the same day in February, then a six-month gap, then one more. That is a project that ships in bursts. Plan for long stretches without upstream fixes.
The upgrade cost is real because the models are versioned separately from the package. The release notes for V1.1.4 say the MFD and MFR models were upgraded to version 1.5 and that all default configurations, documentation and examples now use mfd-1.5 and mfr-1.5. A model bump like that can change your output on the same input, and since the README does not document a rollback path, you should snapshot both the package version and the downloaded model files if you need reproducible output. The Makefile's commented examples, which pass explicit --analyzer-model-fp and --latex-ocr-model-fp paths, show the mechanism for pinning a specific model file if you need to.
Editorial conclusion
Adopt Pix2Text if you are comfortable in Python and your documents mix prose, tables and formulas that must come out as Markdown with LaTeX inline. Skip it if you need a supported desktop product, a hosted API with an uptime commitment, or OCR for a script outside the 80+ language set, since the other-language path runs through EasyOCR rather than the CnOCR models used for English and Simplified Chinese. Before committing, install the package and run the predict command from the Makefile against one page of your own scanned PDF, then compare the Markdown by hand against the source. Check the model weights download cleanly from Hugging Face or the hf-mirror mirror, since the release notes for v1.1.5 describe a refactor of model downloading to use HuggingFaceDownloader.
Frequently asked questions
Is there a free version of Mathpix?
Pix2Text is positioned as a free and open-source Python alternative to Mathpix under the MIT licence, and the README states it can already accomplish Mathpix's core functionality. The project also runs a free-to-use online web service at p2t.breezedeus.com for people who do not want to install the Python package.
How can I convert a math image to text with Pix2Text?
Install the package with pip install pix2text, then run the p2t predict command from the Makefile against your image with the mfd analyzer selected. The formula regions are detected by the MFD model and recognized into LaTeX by the MFR model, and the result is assembled as Markdown.
How to use Pix2Text?
Install it from PyPI as pix2text, which also provides the p2t console command. The Makefile shows a working invocation using p2t predict with -l en,ch_sim, -a mfd, -t yolov7_tiny and an input image, and the README points to the online documentation for fuller examples.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/breezedeus-pix2text)