# OpenBMB/VisRAG: Parsing-Free RAG That Embeds Documents as Images

> VisRAG replaces the PDF-to-text step in retrieval-augmented generation with a vision-language embedding model, and EVisRAG (VisRAG 2.0) adds evidence-guided multi-image reasoning on top. Here is what the repository actually ships, how to install it, and where the approach breaks down.

**OpenBMB/VisRAG** — Parsing-free RAG supported by VLMs

- Repository: https://github.com/OpenBMB/VisRAG
- Stars: 981 · Forks: 80
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/openbmb-visrag

## The parsing step VisRAG removes

Classic retrieval-augmented generation over documents has a lossy front end. A PDF is parsed into text, the text is chunked, embedded, and indexed, and the generator answers from those chunks. Everything that does not survive parsing is gone: table structure, chart values, figure captions placed far from the text that references them, and the reading order of multi-column pages. The README frames the project's motivation exactly this way, stating that VisRAG maximizes retention of information in the original documents and eliminates the loss introduced during parsing.

The intended user is not a general chatbot builder. It is a team whose corpus is visually rich enough that text extraction is the bottleneck: financial filings with dense tables, scanned reports, slide decks, scientific papers where a result lives in a figure. The repository ships two distinct things for that audience. VisRAG-Ret is a document embedding model built on MiniCPM-V 2.0, which itself pairs SigLIP as the vision encoder with MiniCPM-2B as the language model. EVisRAG, branded VisRAG 2.0, is a separate end-to-end framework trained on Qwen2.5-VL-7B-Instruct and Qwen2.5-VL-3B-Instruct for multi-image question answering. One retrieves, the other reasons over what was retrieved.

## How VisRAG-Ret and EVisRAG split the work

The data flow has two stages and they are trained independently. In the retrieval stage, each document page is rendered as an image and passed through VisRAG-Ret, which produces a single embedding per page. No OCR, no layout parser, no chunker sits between the file and the vector. Those page-level vectors go into whatever index you build, and a query returns a ranked set of page images rather than ranked text passages.

In the generation stage, the retrieved images are handed to a vision-language model. EVisRAG changes what happens there. Per the README, it first observes the retrieved images linguistically to collect per-image evidence, then reasons over those cues to produce the answer. Training uses Reward-Scoped GRPO, which the README describes as binding fine-grained, token-level rewards to scope-specific tokens so that visual perception and reasoning are optimized together rather than in sequence. The repository also notes that VisRAG-Gen in the original paper used MiniCPM-V 2.0, MiniCPM-V 2.6, and GPT-4o, and that any VLM can serve as the generator, so the generator is a swappable component while VisRAG-Ret is the piece that defines the retrieval semantics.

## Installing VisRAG and running the pipeline

The repository carries two requirement sets because the two halves have different dependency trees. EVisRAG installs from EVisRAG_requirements.txt on Python 3.10. VisRAG installs from requirements.txt on Python 3.10.8, pins torch 2.1.2 with torchvision 0.16.2 and transformers 4.40.2, and additionally requires a CUDA 11.8 toolkit plus an editable install of the bundled timm_modified package. That last package is an enhanced timm that adds gradient checkpointing, which the README says is used during training to reduce memory usage. If you only intend to run inference, you still install it, because the setup instructions do not separate the two paths.

```bash
git clone https://github.com/OpenBMB/VisRAG.git
conda create --name VisRAG python==3.10.8
conda activate VisRAG
conda install nvidia/label/cuda-11.8.0::cuda-toolkit
cd VisRAG
pip install -r requirements.txt
pip install -e .
cd timm_modified
pip install -e .
cd ..
```

After that, the repository provides two entry points for turning files into something retrievable and for running the full question-answering loop. File conversion lives under visrag_scripts/file2img, and the end-to-end demo lives under visrag_scripts/demo/visrag_pipeline, which the release notes say was extended to support visual understanding across multiple PDF documents. A hosted version of that pipeline exists as a Hugging Face Space and a Colab notebook linked from the README, which is the fastest way to see the behaviour before installing anything locally.

```bash
cd visrag_scripts/file2img
python convert.py --help
```

The command above is the directory the README names for file-to-image conversion; the README does not print a specific invocation, so read the script's own arguments before running it against a corpus. Expect one image per page, and expect that image to be the unit of retrieval and the unit of context handed to the generator.

## Training cost and the two-stage dependency chain

EVisRAG training is not a single script. Stage 1 is supervised fine-tuning built on LLaMA-Factory, and stage 2 is RS-GRPO built on Easy-R, which the README names as the second framework dependency. The SFT stage clones LLaMA-Factory separately and runs a shell script from the repository.

```bash
git clone https://github.com/hiyouga/LLaMA-Factory.git
bash evisrag_scripts/full_sft.sh
```

This is the part most teams underestimate. Reproducing EVisRAG means maintaining three codebases (VisRAG, LLaMA-Factory, Easy-R) with their own version constraints, plus the timm_modified fork, plus a CUDA 11.8 toolchain. The pinned transformers 4.40.2 in requirements.txt is a hard constraint for the VisRAG half, and it will conflict with anything in your stack that expects a newer transformers release. If your plan is to fine-tune on in-domain data rather than use the released checkpoints, budget for that dependency surface before you start, not after the first failed install.

## Where the image-as-document approach is the wrong tool

The design trades compute for fidelity, and that trade does not always pay. A page image embedding carries far more tokens of information than a text chunk, so a VisRAG-Ret index is heavier to build and heavier to query than a text index over the same corpus. For a corpus of plain prose, contracts, or Markdown, a text embedding model is cheaper, faster, and already good enough; VisRAG's advantage only materializes when parsing destroys something you need.

The second limitation is granularity. Retrieval operates at the page level, and the README gives no chunking or sub-page selection mechanism. If an answer sits in one paragraph of a forty-page annual report, the whole page image enters the generator's context. That is a context-budget problem, and EVisRAG's multi-image reasoning does not remove it, it just makes the generator better at handling several such images at once. Third, the project is research code. There are no released versioned packages, the setup.py in the repository is still named openmatch-thunlp with version 0.0.1 and an MIT classifier despite the repository carrying Apache-2.0, and the README documents no rollback path or index migration story. Treat upgrades as manual work. The last push to the default branch was on 2026-08-09.

## VisRAG against VDocRAG and text-first RAG

VDocRAG is the closest comparison point and it appears in the same search space around visually rich documents. The difference is where the vision model sits. In a VDocRAG-style design, the system still extracts structure from the document and feeds structured text to the generator; the vision model helps with extraction quality. VisRAG skips extraction entirely and lets the vision-language model consume the page image at both retrieval and generation time. The practical consequence is that VisRAG has no OCR or layout-parse failure mode, but it also has no intermediate representation you can inspect, edit, or correct by hand. When VisRAG-Ret ranks the wrong page, you cannot read the parsed text and see why; you can only look at the image and the score.

Against ordinary text-first RAG the trade is starker. Text RAG gives you cheap indexes, transparent chunks, and the ability to patch a bad chunk. VisRAG gives you fidelity to the original page and a much larger compute bill. The honest framing is that these are not competitors on the same axis: VisRAG is what you reach for after text extraction has demonstrably failed on your corpus, not as a default upgrade.

## License and what to verify before adopting

The repository is licensed Apache-2.0. That covers the code in this repository. It does not automatically cover the model weights: VisRAG-Ret, EVisRAG-3B, and EVisRAG-7B are distributed separately on Hugging Face, and the base models they build on (MiniCPM-V 2.0, Qwen2.5-VL-7B-Instruct, Qwen2.5-VL-3B-Instruct) carry their own terms. If you plan to ship a product on top of these weights, check the license attached to each checkpoint you download, not just the LICENSE file in the repository. The setup.py metadata is also inconsistent with the stated license, so do not treat it as authoritative.

On maintenance: the default branch was last pushed on 2026-08-09 and the repository is not archived. There are no tagged releases, so pinning to a commit hash is the only reproducible option. Before adopting, verify that your GPU memory fits the checkpoint you choose, that your CUDA version matches the 11.8 toolkit the VisRAG setup installs, and that the file formats in your corpus are handled by visrag_scripts/file2img. Those three checks will tell you more than any benchmark table.

## Conclusion

Adopt VisRAG when your corpus is visually rich (scanned pages, charts, multi-column layouts) and you can afford GPU inference for both retrieval and generation; skip it if your documents are clean text and a standard text embedding model already retrieves well, since you would be paying vision-model cost for no gain. Before committing, verify which model checkpoints you intend to serve (VisRAG-Ret for embedding, EVisRAG-3B or EVisRAG-7B for generation), confirm the CUDA and Python versions your environment provides against the two separate requirements files, and check whether the visrag_scripts/file2img or visrag_scripts/demo/visrag_pipeline path covers your file formats, because the README documents no rollback or fallback path if image conversion fails.

## FAQ

### What is OpenBMB/VisRAG and how is it different from text-based RAG?

VisRAG is a vision-language model based RAG pipeline in which documents are embedded as images rather than parsed into text first. The README states this maximizes retention of information in the original documents and eliminates the loss introduced during parsing.

### What is the difference between VisRAG-Ret and EVisRAG?

VisRAG-Ret is the document embedding model built on MiniCPM-V 2.0 with SigLIP as the vision encoder and MiniCPM-2B as the language model. EVisRAG, also called VisRAG 2.0, is an end-to-end framework trained on Qwen2.5-VL-7B-Instruct and Qwen2.5-VL-3B-Instruct that collects per-image evidence before reasoning over retrieved images.

### How do I install VisRAG?

The README gives a conda-based setup on Python 3.10.8 that installs the CUDA 11.8 toolkit, then pip install -r requirements.txt, pip install -e ., and an editable install of the bundled timm_modified package. EVisRAG has a separate environment on Python 3.10 using EVisRAG_requirements.txt.

### What is timm_modified in the VisRAG repository?

The README describes timm_modified as an enhanced version of the timm library that supports gradient checkpointing, which the project uses during training to reduce memory usage. The setup instructions install it as a separate editable package after the main requirements.

### Can I try VisRAG without installing it locally?

Yes. The README links a VisRAG Pipeline demo hosted on Hugging Face Spaces and a Google Colab notebook, both released in late 2024. The README says the pipeline supports visual understanding across multiple PDF documents.

## Sources

- [Issues](https://github.com/OpenBMB/VisRAG/issues)
- [License: Apache-2.0](https://github.com/OpenBMB/VisRAG/blob/master/LICENSE)
- [OpenBMB/VisRAG on GitHub](https://github.com/OpenBMB/VisRAG)
- [README](https://github.com/OpenBMB/VisRAG/blob/master/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/openbmb-visrag
