DeepSeek-OCR: Running a Vision-Token OCR Model Locally
Project brief: Contexts Optical Compression. Model Download | Paper Link | Arxiv Paper Link | DeepSeek-OCR: Contexts Optical Compression Explore the boundaries of visual-text compression.
At a glance
- What is it?
- DeepSeek-OCR compresses document images into a small number of vision tokens and decodes them with a language model. Here is how the install works, what the prompts do, and where the approach breaks down.
- Who is it for?
- Adopt DeepSeek-OCR if you have a CUDA 11.8 or newer GPU, need layout-aware markdown from documents, and are willing to run a 3B-class model yourself; the repository ships both a vLLM path and a Transformers path, so pick one before you install anything. Do not adopt it if you need a hosted API key, CPU-only inference, or a supported Ollama build, because the README documents none of those.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Activity is slowing. The repository last received commits 8 months ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
DEEP OPEN-SOURCE ANALYSIS
What DeepSeek-OCR actually does with a page
Most OCR pipelines hand a page to a detector, then a recognizer, then a layout model, and glue the three outputs together. DeepSeek-OCR collapses that into one model. The README describes the project as an investigation into "the role of vision encoders from an LLM-centric viewpoint," and the subtitle on the repository page is "Contexts Optical Compression." The claim embedded in that framing is that a document does not need to be represented as thousands of text tokens when it can be represented as a few hundred vision tokens and decoded back into text.
The supported modes make the compression explicit. Native resolution runs at Tiny (512x512, 64 vision tokens), Small (640x640, 100 vision tokens), Base (1024x1024, 256 vision tokens) and Large (1280x1280, 400 vision tokens). A dynamic mode called Gundam tiles the page as n tiles of 640x640 plus one 1024x1024 tile. Those token counts are the whole point: a dense page of text might cost several thousand text tokens, while Base spends 256 vision tokens before the decoder writes anything.
Who is this for? Teams that already run local GPU inference and want document conversion they control, plus researchers working on vision-text compression. It is not aimed at someone who wants to paste a URL into a web form.
Install: conda, pinned wheels, and the flash-attn build
The README states the reference environment as cuda11.8+torch2.6.0, and the install sequence is deliberately pinned. Start by cloning the repository:
git clone https://github.com/deepseek-ai/DeepSeek-OCR.gitThen create the Python environment. The README uses Python 3.12.9:
conda create -n deepseek-ocr python=3.12.9 -y
conda activate deepseek-ocrPackages come in four steps, and the order matters because the vLLM wheel is built against CUDA 11.8. The README points at the vllm-0.8.5 release page for the wheel file:
pip install torch==2.6.0 torchvision==0.21.0 torchaudio==2.6.0 --index-url https://download.pytorch.org/whl/cu118
pip install vllm-0.8.5+cu118-cp38-abi3-manylinux1_x86_64.whl
pip install -r requirements.txt
pip install flash-attn==2.7.3 --no-build-isolationThe requirements.txt in the repository root pins transformers==4.46.3 and tokenizers==0.20.3 alongside PyMuPDF, img2pdf, einops, easydict, addict, Pillow and numpy. Note that the README itself warns about a version conflict: vllm 0.8.5+cu118 wants transformers>=4.51.1, while requirements.txt pins 4.46.3. The README says you do not need to worry about that error if you want both vLLM and Transformers code in one environment, but it does not explain which version wins. If you only need one inference path, install only that path's dependencies.
First run: image, PDF, and the prompt that sets the output format
The vLLM scripts live under DeepSeek-OCR-master/DeepSeek-OCR-vllm, and the README tells you to edit INPUT_PATH, OUTPUT_PATH and other settings in config.py before running anything:
cd DeepSeek-OCR-master/DeepSeek-OCR-vllm
python run_dpsk_ocr_image.pyFor PDFs there is a second script. The README annotates it with a throughput note of roughly 2500 tokens/s on an A100-40G, which is the project's own figure, not an independent measurement:
python run_dpsk_ocr_pdf.pyA third script, run_dpsk_ocr_eval_batch.py, is for benchmark evaluation. The prompt is what selects the task. The README lists these examples, and the distinction between them matters more than it looks:
# document: <image>\n<|grounding|>Convert the document to markdown.
# other image: <image>\n<|grounding|>OCR this image.
# without layouts: <image>\nFree OCR.
# figures in document: <image>\nParse the figure.
# general: <image>\nDescribe this image in detail.
# rec: <image>\nLocate <|ref|>xxxx<|/ref|> in the image.Use the grounding prompt when you want reading order and layout preserved as markdown. Use Free OCR when you only want the characters and do not care where they sat on the page. The Transformers path is separate; the README points at DeepSeek-OCR-master/DeepSeek-OCR-hf and run_dpsk_ocr.py, and shows loading the model with trust_remote_code=True and use_safetensors=True, then casting to bfloat16 with flash_attention_2.
Through upstream vLLM, with one logits processor you must not forget
Since 2025-10-23 the model is supported in upstream vLLM, and the README documents that recipe separately from the pinned local install. The nightly index is required until v0.11.1:
uv venv
source .venv/bin/activate
uv pip install -U vllm --pre --extra-index-url https://wheels.vllm.ai/nightlyThe Python entry point is where a subtle requirement appears. The example imports NGramPerReqLogitsProcessor from vllm.model_executor.models.deepseek_ocr and passes it in logits_processors, with enable_prefix_caching=False and mm_processor_cache_gb=0:
from vllm import LLM, SamplingParams
from vllm.model_executor.models.deepseek_ocr import NGramPerReqLogitsProcessor
llm = LLM(
model="deepseek-ai/DeepSeek-OCR",
enable_prefix_caching=False,
mm_processor_cache_gb=0,
logits_processors=[NGramPerReqLogitsProcessor]
)Both caching switches are off. That is a throughput cost, and the README does not explain why they are disabled, only that they are. If you are tuning for latency on repeated pages, that is the first thing you would want to understand before assuming the defaults are conservative. The README's code sample is also truncated mid-dictionary in the model_input list, so you will be reconstructing the batch format from the full source rather than from the README alone.
Where this model is the wrong tool
The install is the first real constraint. Everything documented assumes a CUDA 11.8 environment, torch 2.6.0, and a flash-attn 2.7.3 build. The README gives no CPU path, no Apple Silicon path, and no Docker image. If your inference host is CPU-only, this project does not meet you there.
The second constraint is the resolution trade-off. Vision token counts are fixed per mode, so a Tiny run at 64 tokens is cheap but is compressing a page aggressively. Small print, dense tables and footnotes are exactly the content that suffers when the encoder is given 512x512. The README lists the modes but does not publish accuracy per mode, so you cannot pick a mode from the documentation alone; you have to run your own pages through each one.
Third, there is no hosted API documented here. The repository is a model release with local inference scripts. If you need a managed endpoint, this is not that product, and nothing in the README describes authentication, rate limits or a service endpoint.
Finally, the README does not document rollback, version pinning policy, or how to switch between the pinned vLLM 0.8.5 build and the upstream nightly recipe once the pinned version falls behind. That is an operational gap you inherit.
How it differs from GOT-OCR2.0 and MinerU
The acknowledgement section names the projects this one builds on: Vary, GOT-OCR2.0, MinerU, PaddleOCR, OneChart and Slow Perception. Two of those are the most natural comparisons.
GOT-OCR2.0 is a unified OCR model in the same family of ideas, and DeepSeek-OCR's authors credit it. The difference the README makes visible is the emphasis on optical compression as the research question itself, with an explicit ladder of vision-token budgets from 64 up to 400 plus the Gundam tiling mode. GOT-OCR2.0 is presented as a capable end-to-end OCR model; DeepSeek-OCR is presented as an experiment in how far you can compress before the text degrades.
MinerU is a document parsing pipeline, credited in the same acknowledgement list, and it is the better comparison for someone whose goal is PDF-to-markdown at scale rather than model research. A pipeline approach tends to give you more knobs for table structure, reading order and post-processing rules. DeepSeek-OCR gives you one model and a prompt, and the layout behaviour is whatever the model learned. If your PDFs have unusual structures and you need deterministic rules for them, a pipeline is easier to reason about. If you want a single model that also handles figures and free-form images, the prompt list here covers Parse the figure and Describe this image in detail in the same interface.
Licence, upgrades, and the OCR2 question
The repository is MIT licensed, which is permissive and places few obligations on how you redistribute or embed the code. The licence file covers the repository; model weights are distributed through the Hugging Face page linked at the top of the README, and the README does not restate weight licensing terms in the text. Check the model card on that page rather than assuming the MIT text extends to the weights.
The upgrade picture is unusual. The release list shows DeepSeek-OCR2 announced on 2026-01-27 at a separate repository, deepseek-ai/DeepSeek-OCR-2. So the project you are reading about is superseded by a successor with its own repository, and the README here does not describe a migration path, shared checkpoints, or API compatibility between the two. Anyone starting today should decide which of the two they are actually adopting before they write install scripts.
On maintenance cost: the pinned stack (torch 2.6.0, vLLM 0.8.5+cu118, flash-attn 2.7.3, transformers 4.46.3) will drift out of step with current CUDA and PyTorch releases. The upstream vLLM recipe is the escape hatch, but it requires the nightly index until v0.11.1, which is not a stable target for production.
Editorial conclusion
Adopt DeepSeek-OCR if you have a CUDA 11.8 or newer GPU, need layout-aware markdown from documents, and are willing to run a 3B-class model yourself; the repository ships both a vLLM path and a Transformers path, so pick one before you install anything. Do not adopt it if you need a hosted API key, CPU-only inference, or a supported Ollama build, because the README documents none of those. Before committing, verify three things: that your driver can run the pinned torch 2.6.0 cu118 wheels, that the flash-attn 2.7.3 build succeeds on your machine, and that the resolution mode you pick (Tiny at 512x512 through Gundam at n x 640x640 plus 1024x1024) still holds the small text in your own scans.
Frequently asked questions
Does DeepSeek have OCR?
Yes. DeepSeek-OCR was released on 2025-10-20 as a model for investigating vision encoders from an LLM-centric viewpoint, and it is MIT licensed. A successor, DeepSeek-OCR2, was announced on 2026-01-27 in a separate repository.
What API does DeepSeek OCR use?
The README does not document a hosted API. It documents two local inference paths: a pinned vLLM 0.8.5 install with scripts under DeepSeek-OCR-master/DeepSeek-OCR-vllm, and a Transformers path under DeepSeek-OCR-master/DeepSeek-OCR-hf. Upstream vLLM also supports the model through its Python LLM class.
How to install DeepSeek OCR?
Clone the repository, create a conda environment with Python 3.12.9, then install torch 2.6.0 with the cu118 index, the vllm-0.8.5+cu118 wheel, requirements.txt, and flash-attn 2.7.3 with --no-build-isolation. The README states the reference environment is cuda11.8+torch2.6.0.
How to use DeepSeek OCR locally?
After installing, edit INPUT_PATH and OUTPUT_PATH in DeepSeek-OCR-master/DeepSeek-OCR-vllm/config.py, then run run_dpsk_ocr_image.py for a single image or run_dpsk_ocr_pdf.py for PDFs. The prompt you pass selects the task, for example <image>\n<|grounding|>Convert the document to markdown.
How to use DeepSeek OCR with Ollama?
The README does not document Ollama support, and no Ollama model name or command appears in the repository files. The documented paths are the pinned vLLM install, the Transformers scripts, and upstream vLLM.
How to use DeepSeek OCR and Docling for PDF parsing?
The README does not mention Docling. For PDFs it documents run_dpsk_ocr_pdf.py under DeepSeek-OCR-master/DeepSeek-OCR-vllm, with INPUT_PATH and OUTPUT_PATH set in config.py, and a throughput note of roughly 2500 tokens/s on an A100-40G.
Community notes