NeMo Retriever Library: what the extraction pipeline actually does
NeMo Retriever Library is a scalable, performance-oriented document content and metadata extraction microservice. NeMo Retriever Library uses specialized NVIDIA NIM microservices to find, contextualize, and extract text, tables, charts and images that you can use in downstream generative applications.
At a glance
- What is it?
- NeMo Retriever Library splits documents into pages, classifies text, tables, charts and infographics, extracts them through NIM microservices or HuggingFace models, and writes embeddings to LanceDB. It is built for teams with GPU capacity and a Kubernetes path, not for a laptop experiment.
- Who is it for?
- Adopt NeMo Retriever Library if you already run NVIDIA NIM microservices or local GPUs and need text, tables, charts and infographics pulled into one JSON schema for a retrieval pipeline. Do not adopt it if your corpus is a few hundred PDFs of plain text, or if you have no GPU and no Kubernetes path, because the quickstart path is explicitly described as in development and the supported production deployment is Helm on Kubernetes.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 6 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem NeMo Retriever Library addresses
Most PDF ingestion code handles text and then gives up. A page with a cost chart, a table with merged cells and a callout box with a diagram becomes either an empty string or a wall of numbers with no labels. NeMo Retriever Library takes the position that each of those artifacts is a separate extraction problem and should be classified and handled separately before anything reaches a vector store.
The README describes the target as finding, contextualizing and extracting text, tables, charts and infographics for downstream generative and retrieval-augmented applications. The audience is therefore teams building RAG over document sets where visual content carries meaning: financial filings, product datasheets, technical manuals, scientific PDFs. If your documents are plain prose, this is a lot of machinery for a problem you do not have.
Page splitting, artifact classification and the JSON schema
The pipeline parallelizes at the page level. According to the README, documents are split into pages, artifacts on each page are classified (text, tables, charts, infographics), extracted, and then contextualized through OCR into a well defined JSON schema. Only after that does the library compute embeddings for the extracted content and store them in LanceDB.
The API reflects that ordering. Ingestion tasks are chainable and defined lazily, so nothing runs until ingest() is called. The extract step takes explicit booleans for extract_text, extract_charts, extract_tables and extract_infographics, which means you can skip chart extraction on a corpus where charts are decorative. The return value in batch and inprocess mode is a pandas.DataFrame, one row per extracted artifact, with the content under a text column. Helper functions to_markdown and to_markdown_by_page turn those rows back into markdown, with the per-page variant keyed by page number.
Two extraction backends are supported: NVIDIA NIM microservices, and a wide range of models including HuggingFace models on local GPUs. That choice is the main architectural decision you make, and it determines your deployment shape.
Installing NeMo Retriever Library and running a first ingest
The README points to the nemo_retriever library subfolder for the quickstart installation steps and notes that this setup is in development. It works with HuggingFace models on local GPUs or with NIMs hosted on build.nvidia.com, and the README scopes it to small-scale workloads of fewer than 100 PDFs.
Start from the repository root and follow the quickstart in nemo_retriever. Once the environment is ready, the README's example constructs an ingestor in batch mode, chains the tasks, and executes:
from nemo_retriever import create_ingestor
from nemo_retriever.common.io import to_markdown, to_markdown_by_page
from pathlib import Path
documents = [str(Path("data/multimodal_test.pdf"))]
ingestor = create_ingestor(run_mode="batch")
ingestor = (
ingestor.files(documents)
.extract(
extract_text=True,
extract_charts=True,
extract_tables=True,
extract_infographics=True
)
.embed()
.vdb_upload()
)
chunks = ingestor.ingest()The README states that ingest() is what actually executes the pipeline, and that the result is a pandas.DataFrame in batch and inprocess mode. Inspect it before wiring anything downstream:
chunks.iloc[0]["text"]
chunks.iloc[1]["text"]
to_markdown_by_page(chunks).keys()With the repository's sample document, the README shows the first row returning raw page text, the second returning a markdown table of animals and activities, and the third returning chart content including the axis labels and values. The per-page markdown keys come back as dict_keys([1, 2, 3]). If you see empty strings where a table should be, the extract flags or the model backend are the first things to check.
For anything beyond small scale, the README directs you to Kubernetes with Helm, starting from the NeMo Retriever Helm chart in nemo_retriever/helm. There is also a Dockerfile at the repository root, with a build command given in its header comment:
docker build -f Dockerfile -t nemo-retriever .
docker run nemo-retriever
docker run -v /host/docs:/data nemo-retriever /dataThe header also documents a service target, docker build -f Dockerfile --target service --build-arg DOWNLOAD_DEFAULT_TOKENIZER=True -t nemo-retriever-service, and a service-gpu target that adds in-pod HuggingFace support. The image installs LibreOffice headless for docx and pptx to PDF conversion, which is why the base image is a full Ubuntu rather than a slim one.
Where the pipeline breaks down or is the wrong choice
The README is explicit that the default branch tracks active development and may be ahead of the latest supported release. That is a real operational constraint: an install from main is not the same artifact as the published 26.08 line, and the README notes that legacy ingestion APIs are being phased out and dependencies simplified. Code written against the older API surface will need rework.
The second constraint is the backend dependency. Chart and infographic extraction is the reason to use this library, and it is also the part that needs a capable model behind it. Without NIM microservices or local GPUs, you are not running the pipeline the README describes. A CPU-only environment is not a supported configuration in the documentation provided.
The third is scale. The quickstart path is scoped to fewer than 100 PDFs. The README's answer for production is Helm on Kubernetes, which means cluster access, GPU scheduling and chart configuration before you get to your first document. If your corpus is small and text-only, a simpler extractor will get you to a working index faster, and you can revisit this when you hit tables and charts that matter.
How this differs from a general-purpose document loader
A general-purpose loader, such as the document readers shipped with LangChain or LlamaIndex, typically produces a flat stream of text and metadata per document. Tables survive as text runs if the parser is lucky, and charts are lost entirely. The library's approach is different in kind: it classifies artifacts per page, routes each class through a specialized model, and emits structured rows that can be turned back into markdown.
That difference shows up in the repository's examples. The examples directory contains both langchain_multimodal_rag.ipynb and llama_index_multimodal_rag.ipynb, which suggests the intended pattern is to keep your existing RAG framework and use this library as the ingestion stage in front of it. There is also an example for querying with metadata filters, nemo_retriever_retriever_query_metadata_filter.ipynb, and one for evaluation with RAGAS, nrl_ragas.ipynb. If you only need text, the heavier pipeline buys you nothing except more moving parts.
Licence, release cadence and what an upgrade costs
The repository is Apache-2.0, and source files carry SPDX headers naming NVIDIA Corporation and Affiliates. The Dockerfile complicates the picture slightly: it installs LibreOffice headless for docx and pptx conversion and carries a GPL_LIBS argument listing packages such as libltdl7, libhunspell-1.7-0, libhyphen0 and libdbus-1-3, with a comment referring to GPL source handling. If you redistribute the service image rather than run it internally, that is worth checking with your own counsel. Nothing here is legal advice.
On cadence, the most recent release listed is 26.08.1 from 2026-08-27, following 26.5.0 in May 2026 and 26.3.0 in March 2026. The last push to the repository was on 2026-09-23. The README tells you to use the 26.08 branch for the latest supported release, with 26.03 as the previous stable line, and states that the published artifacts are PyPI and Helm chart version 26.8.1. Note the version string mismatch between the release tag 26.08.1 and the chart version 26.8.1; treat them as the same line and confirm against the chart README before pinning. An upgrade means re-reading the branch note, since main and the stable line diverge.
Editorial conclusion
Adopt NeMo Retriever Library if you already run NVIDIA NIM microservices or local GPUs and need text, tables, charts and infographics pulled into one JSON schema for a retrieval pipeline. Do not adopt it if your corpus is a few hundred PDFs of plain text, or if you have no GPU and no Kubernetes path, because the quickstart path is explicitly described as in development and the supported production deployment is Helm on Kubernetes. Verify first which branch you are installing from: main tracks active development and may be ahead of the supported 26.08 release line, which is the one published as PyPI and Helm chart version 26.8.1.
Frequently asked questions
What is NeMo Retriever Library?
It is a scalable, performance-oriented framework for document content and metadata extraction, supporting NVIDIA NIM microservices and a wide range of models to find, contextualize and extract text, tables, charts and infographics. The README notes it is also referred to as NVIDIA Ingest in some NVIDIA product materials.
Is NVIDIA NeMo Retriever free?
The repository is licensed under Apache-2.0, so the library source is freely available. Running the pipeline still requires either NVIDIA NIM microservices or local GPUs for the HuggingFace path, and the Dockerfile installs GPL-licensed LibreOffice components for docx and pptx conversion, which matters if you redistribute the image.
What is NVIDIA NeMo Retriever?
It is the same project: the README describes NeMo Retriever Library as a framework that extracts text, tables, charts and infographics from documents, computes embeddings for the extracted content and stores them in LanceDB. It can run on NVIDIA NIM microservices or on local GPU models.
What is NVIDIA NeMo used for?
In this repository, the NeMo Retriever name covers document extraction and metadata generation for retrieval-augmented applications. The README scopes the library to finding, contextualizing and extracting text, tables, charts and infographics that feed downstream generative and retrieval pipelines.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/nvidia-nemo-retriever)