Model or dataset
NanoNets/docext avatar
NanoNets/docext

docext: on-premises PDF to markdown and document extraction without OCR

An on-premises, OCR-free unstructured data extraction, markdown conversion and benchmarking toolkit. (https://idp-leaderboard.org/)

2,086 stars157 forksPythonApache-2.0

At a glance

What is it?
NanoNets/docext is an Apache-2.0 Python toolkit that turns PDFs and images into markdown and structured fields using vision-language models instead of a classic OCR pipeline. It installs from PyPI, ships a Gradio UI and a REST API, and its benchmark lives at idp-leaderboard.org.
Who is it for?
Adopt docext if you need document extraction to stay on your own Linux or macOS hardware and you can supply the GPU memory a vision-language model needs, or if you want to reproduce the idp-leaderboard.org numbers yourself. Do not adopt it if you need a maintained release cadence, Windows support, or an OCR-free path for handwriting-heavy scans, since the README lists no handwriting capability for the extraction flow.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Activity is slowing. The repository last received commits 6 months ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What docext actually replaces in a document pipeline

A conventional extraction stack chains three moving parts: an OCR engine that produces text, a layout model that reconstructs reading order, and a parser that pulls fields out. Each stage loses something, and tables and equations are usually where it shows. docext takes the other route. It sends page images straight to a vision-language model and asks for markdown or for structured fields, so the intermediate text layer never exists. The README describes this as "OCR-free extraction of structured information (fields, tables, etc.) from documents such as invoices, passports, and other document types, with confidence scoring."

The audience is narrow and specific. You are building document ingestion for an application that cannot send scans to a hosted API, you have Linux or macOS hardware with a GPU, and you are willing to run a model yourself. The README lists on-premises deployment as a feature and names Linux and MacOS as the supported platforms. If your constraint is accuracy on a benchmark rather than data residency, the same repository also publishes the Intelligent Document Processing Leaderboard, which evaluates models across key information extraction, visual question answering, OCR, document classification, long document processing, table extraction, and confidence score calibration. That leaderboard is useful even if you never run the extraction code.

How the VLM pipeline is wired: markdown tags, templates and a REST API

The markdown converter is where the design is most visible. Rather than emitting plain text, it wraps recognised elements in semantic tags. Signatures land inside `<signature></signature>`, watermarks inside `<watermark></watermark>`, page numbers inside `<page_number></page_number>`, and image descriptions inside `<img></img>`. Checkboxes and radio buttons become standard Unicode symbols for unchecked, checked and crossed boxes. Tables come out as HTML tables, and inline and block LaTeX equations are converted to markdown. This tagging scheme is the actual interface: downstream code can strip or route regions by tag instead of guessing from text position.

The extraction side follows a template model. The README names two pre-built templates, invoices and passports, and states that fields or columns can be added and deleted for other templates. Output includes confidence scores per extracted item, which matters because a VLM can produce a plausible value for a field that is not present on the page. The repository also exposes a REST API for programmatic access, and the Dockerfile shows the runtime shape: it is built on `vllm/vllm-openai:v0.8.2`, installs Python 3.11 into a virtual environment, and starts the app with `--no-share --ui_port 7860`. Dependencies are pinned tightly, including `transformers>=4.51.1,<4.53.0`, `vllm==0.8.3` and `xgrammar==0.1.17`. Those pins are the honest signal about how coupled this is to a specific serving stack.

Installing docext and converting your first PDF

The package is on PyPI and requires Python 3.11 or newer according to setup.py. The README points to EXT_README.md for installation and usage details beyond the summary, so treat the commands below as the documented entry points rather than a complete walkthrough.

bash
pip install docext

After installation, the console script registered in setup.py is `docext`, which maps to `docext.__main__:main`. The repository also ships `docext.ipynb` and `pdf2markdown.ipynb` notebooks, and the README links a Colab notebook, which is the fastest way to see the markdown output before you commit hardware.

For a containerised run, the Dockerfile builds on the vLLM OpenAI image and exposes the Gradio UI on port 7860:

dockerfile
ENV GRADIO_SERVER_PORT=7860
ENV GRADIO_SERVER_NAME="0.0.0.0"
EXPOSE 7860
ENTRYPOINT ["/app/.venv/bin/python", "-m", "docext.app.app", "--no-share", "--ui_port", "7860"]

One caveat worth stating plainly: the Dockerfile installs `flash-attn` with `--no-build-isolation` after the application install, which is a build step that frequently needs a matching CUDA toolchain. The README does not document a CPU-only path.

Where docext is the wrong tool

The dependency list is the first limitation. `vllm==0.8.3` and `flash-attn` mean a GPU is the expected runtime, and the README's platform list stops at Linux and macOS. Windows is not mentioned. A team that wants to run extraction on a CPU-only CI runner or a Windows workstation will spend more time on the environment than on documents.

The second limitation is release cadence. The most recent release in the repository is v0.1.14 from 2025-06-30, and the last push to the default branch was on 2026-03-17. The version number is still 0.1.x. If your adoption decision requires a project with a frequent tagged release history, that history is not here. The change log entries that are present cluster around leaderboard additions and support for multi-document and PDF input in extraction, which suggests the extraction feature set is younger than the benchmarking side.

The third case is handwriting and degraded scans. The README lists LaTeX equations, image descriptions, signatures, watermarks, page numbers, checkboxes, radio buttons and tables as recognised elements. Handwriting is not on that list for the extraction flow, and an OCR-free approach gives you no separate text layer to fall back on when the model misreads a region. If you need a searchable text layer for compliance or full-text search, docext's markdown output is not the same artifact as an OCR text dump.

docext against a classic OCR plus layout stack

The obvious alternative is the traditional combination of an OCR engine such as Tesseract with a layout or table model, assembled in something like Unstructured or a custom pipeline. The difference is not a matter of quality, it is a matter of where errors appear. A classic stack produces a text layer you can inspect, diff and index. When the OCR engine misreads a character, you can see it in the text and correct it. docext produces markdown or JSON with no equivalent intermediate artifact, so a misread field arrives as a confident-looking value with a confidence score attached. That score is the only signal you get, and the leaderboard includes confidence score calibration as one of its seven evaluation tasks precisely because that signal is not automatically trustworthy.

The other difference is operational. A Tesseract pipeline runs on CPU and is cheap to scale horizontally. docext wants a GPU and a served vision-language model, which changes your cost model from per-page CPU time to GPU occupancy. The trade is that equations, signatures and watermarks come out tagged rather than as noise in a text stream. If your documents are clean, typed invoices, the classic stack is likely cheaper and easier to debug. If your documents are visually rich and you need the structure preserved, the VLM route is the one docext implements.

Licence, maintenance and what an upgrade costs

The repository is Apache-2.0, and the README carries the Apache-2.0 badge. There is a discrepancy worth flagging: setup.py declares `"License :: OSI Approved :: MIT License"` in its classifiers while the LICENSE file and README both point to Apache-2.0. If you are running a licence scan, that classifier will report MIT. Resolve it against the LICENSE file rather than the metadata. Apache-2.0 includes an explicit patent grant and requires attribution and notice retention; this is a description of the licence terms, not legal advice.

Upgrade cost is dominated by the pinned serving stack rather than by docext's own API. `vllm==0.8.3`, `xgrammar==0.1.17`, `transformers>=4.51.1,<4.53.0` and `python-levenshtein==0.27.1` are exact or narrow pins. Moving to a newer transformers release means moving vLLM, which means moving the CUDA and flash-attn versions underneath it. The `dependency_links` entry in setup.py points at a specific transformers tarball commit, which is another sign that the project tracks a particular build. Budget for a coordinated upgrade across the model server and the client library, not a `pip install --upgrade docext`.

Editorial conclusion

Adopt docext if you need document extraction to stay on your own Linux or macOS hardware and you can supply the GPU memory a vision-language model needs, or if you want to reproduce the idp-leaderboard.org numbers yourself. Do not adopt it if you need a maintained release cadence, Windows support, or an OCR-free path for handwriting-heavy scans, since the README lists no handwriting capability for the extraction flow. Before committing, verify three things: that pip install docext resolves against transformers>=4.51.1,<4.53.0 on your Python 3.11 environment, that your chosen VLM fits the GPU you have, and that the fields you need are covered by the invoice and passport templates or can be added through the custom field mechanism described in EXT_README.md.

Frequently asked questions

Is Nanonets OCR free?

The docext repository is published under Apache-2.0, so the toolkit itself is free to use under those terms. The README also links Nanonets' commercial document parsing and extraction service, which is a separate product from the open-source package.

How do I install docext with pip?

Run pip install docext. setup.py requires Python 3.11 or newer, and the README points to EXT_README.md for the full installation and usage guide.

What document types does docext have pre-built templates for?

The README names two pre-built templates, invoices and passports, and states that fields or columns can be added and deleted to build templates for other document types.

Does docext run on-premises?

Yes. The README lists on-premises deployment as a feature and names Linux and MacOS as the supported platforms, with a Dockerfile that serves the app on port 7860.

Official sources

  1. License: Apache-2.0
  2. NanoNets/docext on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/nanonets-docext.svg)](https://hysenlabs.com/projects/nanonets-docext)