RAG-Anything: a multimodal RAG framework built on LightRAG and MinerU
"RAG-Anything: All-in-One RAG Framework"
At a glance
- What is it?
- HKUDS/RAG-Anything wraps LightRAG and MinerU into one Python pipeline that parses text, images, tables, equations, audio and video into a single queryable index. It is a library, not a service, and its weight is its main cost.
- Who is it for?
- Adopt RAG-Anything when your corpus is genuinely mixed-media and you already run an LLM endpoint you control, since the package is a library that assumes you supply the model. Skip it when you want a hosted UI, a managed ingestion service, or a text-only pipeline where LightRAG or a plain vector store is less machinery.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 15 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
DEEP OPEN-SOURCE ANALYSIS
What RAG-Anything is for, and who ends up using it
The README frames the problem directly: documents now carry text, images, tables, equations, charts and multimedia, and text-only retrieval pipelines drop most of that. RAG-Anything answers with one pipeline that parses every modality and writes the result into a LightRAG index, so a single query can reach a diagram caption, a table row and a paragraph of prose.
The intended audience is developers building retrieval over academic papers, technical documentation, financial reports and enterprise knowledge bases, which is the list the README names. The repository layout supports that reading: the package is published as raganything on PyPI, the project metadata lists Development Status 4 - Beta, and the examples directory is full of integration scripts rather than application code. There is no server binary and no web UI in the top-level entries, so the person who benefits is the one writing Python, not the one looking for a hosted product.
That matters for evaluation. If you want a tool you point at a folder and query through a browser, this is the wrong shape of project. If you want a library that a larger application imports, the shape is right.
The pipeline: MinerU parses, LightRAG indexes, a VLM reads the pictures
Two upstream projects carry the weight. MinerU, pinned as mineru[core]>=3.4.1 in requirements.txt, handles document parsing, and the comment in that file is unusually candid about why the floor exists: earlier versions iterate pdftext's PageChars return value directly and crash with "TypeError: 'PageChars' object is not iterable" against pdftext>=0.7. LightRAG, pinned as lightrag-hku<1.5, provides the graph and vector retrieval layer underneath.
On top of that base, RAG-Anything adds modality-specific processors. The optional dependency groups in pyproject.toml map to distinct content types: image covers Pillow for BMP, TIFF, GIF and WebP conversion, text covers reportlab for TXT and MD to PDF, audio covers faster-whisper, and video pulls in scenedetect, moviepy, faster-whisper and opencv-python together. A separate paddleocr group adds paddleocr and pypdfium2, which the requirements comment recommends for scanned PDFs. Office formats are handled through an empty extras group whose comment states it requires LibreOffice as an external program, not a Python package.
The query side has a mode the changelog calls VLM-Enhanced Query, released in 2025.08, which sends document images into a vision language model alongside the retrieved text. The repository also ships per-provider example scripts, including examples/ollama_integration_example.py, examples/vllm_integration_example.py, examples/lmstudio_integration_example.py and examples/minimax_integration_example.py, so model hosting is a configuration choice rather than something the package bundles.
Installing RAG-Anything and running a first multimodal query
The base install is small. The project requires Python 3.10 or later according to pyproject.toml, and the core dependencies are huggingface_hub, lightrag-hku, mineru[core] and tqdm. Installing the package name alone gets you the text and PDF path.
pip install raganythingThe requirements file documents the extras explicitly, and notes that for best PDF handling you should add paddleocr. If your documents are scanned rather than born-digital, take that advice seriously, because the OCR path is a separate parser selection rather than something the default parser switches to on its own.
pip install raganything[paddleocr]For a broader install that covers images, text conversion, markdown rendering and OCR in one step, the extras group named all collects them. Note that it does not include the office group, which needs LibreOffice installed on the machine.
pip install raganything[all]The repository provides examples/raganything_example.py as the reference starting point, and the integration examples show how to point the pipeline at a locally hosted model. If you run Ollama, examples/ollama_integration_example.py is the file to read first, and examples/vllm_integration_example.py covers the vLLM route. Batch ingestion has its own examples in examples/batch_processing_example.py, with examples/batch_dry_run_example.py for validating a run before committing to it. The env.example file at the repository root is the place to look for the environment variables the examples expect; the README excerpt does not enumerate them, so read that file rather than guessing key names.
Where the all-in-one claim costs you
The dependency list is the honest limitation. A video pipeline pulls scenedetect, moviepy, faster-whisper and opencv-python; audio pulls faster-whisper; OCR pulls paddleocr and pypdfium2. Installing raganything[all] on a clean machine is a large operation, and the extras exist precisely because most users do not need all of them. Choose the narrow group that matches your corpus.
Office documents are the sharpest edge. The extras group is declared empty with a comment that it requires LibreOffice as an external program. That means pip cannot satisfy it, container images need LibreOffice baked in separately, and a deployment that works on a developer laptop can fail in a slim container for a reason no Python traceback will explain.
Quality is bounded by the parser, not by the framework. MinerU's output determines what the index ever sees. A two-column academic PDF, a scanned form, or a table with merged cells can come out mangled, and no amount of graph retrieval recovers text that was never extracted correctly. The requirements comment about the pdftext crash is a reminder that this layer has had real breakages.
The VLM-Enhanced Query mode is only as good as the vision model you attach. The package does not ship a model. Point it at a small local VLM and image-heavy queries will underperform; point it at a hosted frontier model and your per-query cost and data-egress profile change. Neither outcome is a bug in RAG-Anything, but both are decisions the framework pushes onto you.
RAG-Anything versus LightRAG alone, and versus RAGFlow
The comparison that matters most is with LightRAG, because RAG-Anything is built on it. LightRAG handles graph-based retrieval over text and, per the 2026.06 news entry, now enables multimodal RAG through native integration of RAG-Anything. So the boundary is moving. If your documents are text-only, LightRAG alone is less machinery and one fewer dependency chain. If you need images, tables and equations parsed and indexed, RAG-Anything is the layer that does the parsing and hands structured content to LightRAG. The pin lightrag-hku<1.5 also means the two projects version independently, and an upgrade on one side can constrain the other.
RAGFlow is the other name that comes up in search. It is a different kind of artifact: a document-understanding platform with its own ingestion and interface, rather than a Python library you import. The practical difference is where the work happens. With RAG-Anything you write the ingestion script, choose the model endpoint, and own the deployment; the package gives you processors and a retrieval layer. With a platform-style tool you get a running system and accept its choices about parsers, models and storage. Neither is strictly better, but the migration cost between them is high, so pick based on whether you want to own the pipeline or operate someone else's.
Maintenance, licensing and upgrade cost
The repository is not archived, and the last push was on 2026-09-15, six days before this writing. Release v1.4.1 landed on 2026-09-02, following v1.4.0 the same day and v1.3.1 on 2026-05-21. That cadence suggests the project is being worked on, though the metadata still classifies it as Beta, which is a fair description of a framework whose upstream parser has had version-specific crash fixes.
The licence is MIT, declared both in pyproject.toml and in the repository's LICENSE file. MIT is permissive: you can use, modify and redistribute the code, including commercially, provided the copyright notice and licence text travel with it. That covers RAG-Anything itself. It does not automatically cover MinerU, LightRAG, PaddleOCR, faster-whisper, moviepy or any model weights you attach, each of which carries its own terms. Check those separately before shipping a product; this is a description of the licence file, not legal advice.
Upgrade cost concentrates in two pins. The lightrag-hku<1.5 ceiling means a LightRAG 1.5 release will not be adopted until RAG-Anything widens the range, and the mineru[core]>=3.4.1 floor exists to avoid a known parser crash. When you upgrade, read the requirements.txt comments, because they carry the reasoning that the README does not.
Editorial conclusion
Adopt RAG-Anything when your corpus is genuinely mixed-media and you already run an LLM endpoint you control, since the package is a library that assumes you supply the model. Skip it when you want a hosted UI, a managed ingestion service, or a text-only pipeline where LightRAG or a plain vector store is less machinery. Before committing, verify three things against your own documents: whether MinerU parses your PDFs cleanly, whether your chosen VLM handles the image content you actually have, and whether the optional extras you need (paddleocr, office, video) install without pulling in LibreOffice or heavy vision packages you did not plan for.
Frequently asked questions
How do I use RAG-Anything with Ollama?
The repository ships examples/ollama_integration_example.py, which is the reference for pointing the pipeline at a locally hosted Ollama model. Install the package first, then read that example alongside env.example for the environment variables it expects.
What is RAG-Anything?
It is a Python library described in its README as an all-in-one multimodal document processing RAG system built on LightRAG. It parses text, images, tables, equations and multimedia into one index so a single query can reach all of them.
Is RAG-Anything any good?
The framework's quality depends on two things it does not control: how cleanly MinerU parses your documents, and which vision language model you attach for VLM-Enhanced Query mode. The project itself is still classified as Beta and its requirements file documents a past parser crash tied to a specific pdftext version.
RAG-Anything versus RAGFlow: what is the difference?
RAG-Anything is a library you import and wire into your own application, supplying the model endpoint and deployment yourself. RAGFlow is a platform-style document understanding system with its own interface and ingestion choices. The trade is ownership of the pipeline against operating an existing one.
What are alternatives to RAG-Anything?
LightRAG is the closest, since RAG-Anything is built on it and LightRAG now integrates it natively for multimodal work. If your documents are text-only, LightRAG alone removes a dependency chain. Platform-style tools occupy a different position because they are not imported as libraries.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/hkuds-rag-anything)
Community notes