# Dolphin by ByteDance: Document Image Parsing with Two-Stage Architecture

> Dolphin is a document image parsing model from ByteDance that classifies a page as digital or photographed, performs layout analysis, then parses text, tables, formulas, and code blocks in parallel, outputting structured JSON and Markdown.

**bytedance/Dolphin** — The official repo for “Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting”, ACL, 2025.

- Repository: https://github.com/bytedance/Dolphin
- Stars: 9,058 · Forks: 777
- Language: Python
- License: NOASSERTION
- Published: 2026-09-21 · Updated: 2026-09-21 · Language: en
- Canonical page: https://hysenlabs.com/projects/bytedance-dolphin

## Document Parsing as a Two-Stage Classification Problem

Dolphin solves a problem that other document extraction tools handle inconsistently: mixed document collections that contain both digitally created PDFs (where text is stored as vector data) and photographed or scanned pages (where text exists only as pixels). Most parsers tune their approach for one type and produce degraded output on the other.

Dolphin-v2's first stage classifies each page as digital or photographed, then performs layout analysis to identify regions and reading order. The second stage applies different parsing strategies based on that classification: holistic parsing for photographed pages, and parallel element-wise parsing for digital pages. The parallel approach for digital pages is cited in the README as the source of the model's efficiency claim, since multiple elements can be processed simultaneously rather than sequentially.

The model handles 21 element types in Dolphin-v2, including text paragraphs, figures, formulas, tables, and code blocks. The changelog notes this expanded element count as a key feature of the v2 release compared to earlier versions.

## Model Versions and Architecture Changes

The repository tracks three model generations. The original Dolphin (0.3B parameters) was released in May 2025 and scored 74.67 on the OmniDocBench overall metric. Dolphin-1.5 (still 0.3B parameters) improved that score to 85.06 with significant parsing improvements. Dolphin-v2 (3B parameters) reaches 89.78 overall and is the current recommended version.

The parameter count increase from 0.3B to 3B is substantial. Running Dolphin-v2 requires more GPU memory than the earlier versions. The requirements file specifies torch 2.6.0, transformers 4.51.0, and accelerate 1.4.0 as core dependencies. The model weights are hosted on Hugging Face under the identifier `ByteDance/Dolphin-v2`.

The repository also references a Fox-Page Benchmark, a manually refined subset of the Fox dataset, available on Baidu Yun and Google Drive. This benchmark can be used to evaluate parsing accuracy on a standardized document set.

## Installing Dolphin and Running the Demo Scripts

To get started, clone the repository and install dependencies:

```bash
git clone https://github.com/ByteDance/Dolphin.git
cd Dolphin
```

```bash
pip install -r requirements.txt
```

Download the Dolphin-v2 model weights from Hugging Face:

```bash
git lfs install
git clone https://huggingface.co/ByteDance/Dolphin-v2 ./hf_model
```

To parse a single document image:

```bash
python demo_page.py --model_path ./hf_model --save_dir ./results \
    --input_path ./demo/page_imgs/page_1.png
```

To parse a PDF:

```bash
python demo_page.py --model_path ./hf_model --save_dir ./results \
    --input_path ./demo/page_imgs/page_6.pdf
```

To parse individual document elements (specifying the element type):

```bash
python demo_element.py --model_path ./hf_model --save_dir ./results \
    --input_path  \
    --element_type [table|formula|text|code]
```

The results are saved to the directory specified by `--save_dir` in JSON and Markdown format.

## Accelerated Inference with vLLM and TensorRT-LLM

The repository includes separate deployment guides for two inference acceleration frameworks. vLLM support was added in June 2025, and TensorRT-LLM support in the same month. These allow Dolphin to run with higher throughput than the default Hugging Face Transformers inference path.

For production deployments that process high document volumes, vLLM or TensorRT-LLM reduce per-document latency by batching requests and optimizing the compute graph. The guides live in the `deployment/` subdirectory of the repository. Neither the vLLM nor TensorRT-LLM guide is reproduced in the main README; users must navigate to the deployment subdirectory for those instructions.

Multi-page PDF parsing was added in June 2025 as well, extending the model's scope from single-page images to full PDF documents. The batch size for parallel element decoding in page-level parsing is configurable with the `--batch_size` flag on `demo_page.py`.

## Limitations and Maintenance Status

The last push to the Dolphin repository was on 2026-03-25. At the time of this review, that puts the project over six months without a commit. The repository has no GitHub releases. The changelog entries in the README all refer to 2025 dates, with the most recent change being the Dolphin-v2 model release in December 2025.

The license field in the repository metadata reads `NOASSERTION`. The README references the MIT license in a badge link, but the metadata discrepancy means the actual license terms should be confirmed in the LICENSE file before any distribution or integration into commercial products.

Dolphin requires substantial GPU resources to run Dolphin-v2 at reasonable speed. The requirements file specifies PyMuPDF, torch, torchvision, and transformers as dependencies, along with deepspeed, triton, and accelerate. The full requirements.txt lists 20 packages with pinned versions, which means compatibility with newer PyTorch or Transformers releases may require adjustment.

## Dolphin Against Marker

Marker is another open-source document parsing tool that converts PDFs and images to Markdown. It uses a pipeline of traditional OCR tools combined with small language models for layout classification and text extraction. Marker does not use a single unified model architecture for the full page; it chains specialized models for each element type.

Dolphin's two-stage architecture is designed as an integrated model trained end-to-end for the document parsing task, which means the classification and parsing stages share representations rather than operating as entirely separate systems. The practical implication is that Dolphin may generalize better across document types that were represented in its training data, while Marker's pipeline may be easier to replace or tune component-by-component for specific document types.

Neither tool provides a managed API. Both require local deployment with GPU hardware for acceptable performance on large document batches.

## Conclusion

Dolphin is worth evaluating for pipelines that need to extract structured content from mixed document collections containing both scanned and digital PDFs. The last commit to the repository was on 2026-03-25, which puts it more than six months without a push. Teams building production systems on it should plan for the possibility that the model will not receive further updates and should verify whether the NOASSERTION license field in the repository metadata is clarified in the LICENSE file before redistribution.

## FAQ

### What is Dolphin from ByteDance?

Dolphin is a document image parsing model from ByteDance that uses a two-stage architecture to classify document pages as digital or photographed, then extracts text, tables, formulas, and code blocks in structured JSON and Markdown format. The current version is Dolphin-v2 with 3B parameters.

### What document types can Dolphin parse?

Dolphin handles both digital PDFs and photographed or scanned document images. It detects 21 element types including text paragraphs, figures, tables, mathematical formulas, and code blocks. Multi-page PDF parsing is supported via demo_page.py.

### How do you download and run the Dolphin-v2 model?

After cloning the repository and installing requirements with `pip install -r requirements.txt`, download the model with `git clone https://huggingface.co/ByteDance/Dolphin-v2 ./hf_model`. Then run `python demo_page.py --model_path ./hf_model --save_dir ./results --input_path <your file>` to parse a document.

## Sources

- [bytedance/Dolphin on GitHub](https://github.com/bytedance/Dolphin)
- [Issues](https://github.com/bytedance/Dolphin/issues)
- [README](https://github.com/bytedance/Dolphin/blob/master/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/bytedance-dolphin
