# NexusRAG: A Hybrid RAG System Combining Vector Search, Knowledge Graph, and Cross-Encoder Reranking

> NexusRAG is a Python/React document question-answering system that combines vector search, a LightRAG knowledge graph, and cross-encoder reranking in a three-way parallel retrieval pipeline. It supports Docling or Marker document parsing, vision-LLM image and table captioning, and inline 4-character citations with page-level navigation.

**LeDat98/NexusRAG** — Hybrid RAG system combining vector search, knowledge graph (LightRAG), and cross-encoder reranking — with Docling document parsing, visual intelligence (image/table captioning), agentic streaming chat, and inline citations. Powered by Gemini or local Ollama models.

- Repository: https://github.com/LeDat98/NexusRAG
- Stars: 530 · Forks: 113
- Language: Python
- License: not declared
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/ledat98-nexusrag

## What NexusRAG Does and Who It Targets

Most RAG systems follow a three-step loop: split text into chunks, embed them, retrieve by vector similarity, and send the top results to a language model. NexusRAG adds stages at every step of that loop and is aimed at developers building internal document-search applications who need more precision than a plain vector search provides.

The three departures from a standard RAG pipeline are: a knowledge graph that tracks entities and relationships across documents, a cross-encoder reranker that scores retrieved candidates jointly with the query rather than by cosine distance alone, and an image-and-table captioning pipeline that embeds visual content into the searchable vector space through LLM-generated text descriptions.

The backend is FastAPI, the frontend is Next.js, and asynchronous jobs such as PDF parsing, Zotero synchronisation, and audio generation run in a Celery worker. The full stack is packaged with Docker Compose.

## The Retrieval Pipeline in Detail

NexusRAG runs retrieval in three parallel streams that are then merged and reranked:

1. Vector search over ChromaDB using BAAI/bge-m3 embeddings (1024 dimensions, multilingual). The system over-fetches 20 candidates.
2. Knowledge-graph entity lookup using LightRAG, which extracts entity and relationship data from documents at ingestion time and supports multi-hop traversal at query time.
3. Cross-encoder reranking using BAAI/bge-reranker-v2-m3, which encodes query and candidate chunk together rather than as separate embeddings. The README notes this is far more precise than cosine similarity alone.

After reranking, NexusRAG keeps the top 8 results above a relevance threshold of 0.15, falling back to top 3 if all candidates fall below the threshold. Images and tables from the same pages as retrieved chunks are surfaced alongside the text results.

Two embedding models are used because vector search and knowledge-graph extraction have different requirements. Vector search benefits from a fast local model. Knowledge-graph extraction benefits from a semantically richer embedding, which is why the README offers Gemini Embedding (3072 dimensions) as an option for the KG embedding provider while keeping bge-m3 for vector search.

## Document Parsing: Docling and Marker

NexusRAG supports two document parsers, switchable via the NEXUSRAG_DOCUMENT_PARSER environment variable:

```bash
NEXUSRAG_DOCUMENT_PARSER=marker
```

Docling is the default. It uses hybrid semantic-and-structural chunking that respects headings, tables, and paragraph boundaries. Its formula-enrichment pipeline requires 18 to 20 GB of VRAM according to the README. Marker uses Surya for LaTeX and runs at 2 to 4 GB of VRAM. Both support PDF, DOCX, and PPTX.

Both parsers share the same output contract (a ParsedDocument type), so switching between them requires only the environment variable change. Both extract images and tables and pass them to the LLM captioning pipeline.

## Visual Document Intelligence: Images and Tables

Images extracted from documents are captioned by a vision LLM (Gemini Vision or an Ollama multimodal model). The caption is appended to the text chunks from the same page before embedding, so a query for "revenue chart" retrieves chunks containing the LLM-generated description of that chart. This avoids maintaining a separate image search index.

The README describes the pipeline as: extract image, generate caption with specific numbers and labels, append the caption text to the page's chunks, and embed. During retrieval, images found on matched pages are surfaced as references using 4-character IDs with page number.

Table extraction follows a parallel path: the parser exports tables as structured Markdown, a text LLM summarises each table with its purpose, key columns, and notable values (up to 500 characters), and the summary is appended to the page chunks. Table summaries are also injected back into the document Markdown for the document viewer.

## Installing and Starting NexusRAG With Docker Compose

The repository includes a docker-compose.yml for a full production stack and a docker-compose.services.yml for supporting services. The .env.example file contains all required configuration.

A minimal Gemini-backed startup:

```bash
cp .env.example .env
```

Edit .env to add a GOOGLE_AI_API_KEY and configure the DATABASE_URL. The compose file starts:

- postgres:15-alpine on host port 5433
- chromadb/chroma:latest on host port 8002
- The FastAPI backend on port 8080
- A Celery worker for async jobs
- An nginx reverse proxy

The README also lists an Ollama provider path for fully local LLM usage. Switch by commenting out the Gemini block in .env and uncommenting the Ollama block. The KG embedding provider is configured separately and can be set to a different provider than the main LLM.

## Limitations and Deployment Considerations

The repository does not list a license. Before deploying or modifying NexusRAG in a production or commercial context, this must be resolved; the absence of a license file means the default copyright terms apply, which restricts redistribution.

The system has significant infrastructure requirements. A local deployment needs PostgreSQL, ChromaDB, a Gemini API key or a local Ollama instance, and enough VRAM to run the embedding and reranking models. The Docling parser alone can consume 18 to 20 GB of VRAM for formula-heavy documents.

The last push to the repository was on 2026-04-20. There are no GitHub releases.

For teams that want a simpler RAG baseline without knowledge-graph enrichment, LlamaIndex provides a well-documented retrieval pipeline with pluggable vector stores and a large community. The difference from NexusRAG is that LlamaIndex is a framework designed for composition, while NexusRAG is a complete application with fixed opinions about its stack. Swapping out the vector store or replacing the reranker in NexusRAG requires changes to the application code rather than a configuration switch.

## Conclusion

NexusRAG is a good fit for developers who want a working hybrid RAG system with knowledge-graph enrichment and citation tracking, and who are comfortable assembling a Docker stack with PostgreSQL, ChromaDB, a Gemini API key, or a local Ollama instance. It is not a drop-in library; it is a full application. The repository has no license file listed and the last push was on 2026-04-20. Before deploying, confirm the license situation and review the .env.example carefully: the Docling parser requires up to 20 GB of VRAM for formula enrichment, while Marker runs at 2 to 4 GB.

## FAQ

### What LLM providers does NexusRAG support?

NexusRAG supports Gemini (via GOOGLE_AI_API_KEY) and Ollama for local models. The .env.example shows how to switch between them by commenting out one provider block and uncommenting the other. The KG embedding provider can be configured independently of the main LLM provider.

### How does NexusRAG handle images and tables in documents?

Images are captioned by a vision LLM and the caption is appended to the text chunks from the same page before embedding, making them searchable through text queries. Tables are summarised by a text LLM and the summaries are appended to page chunks and injected into the document view.

### How does the cross-encoder reranking improve retrieval over standard vector search?

The cross-encoder (BAAI/bge-reranker-v2-m3) encodes the query and each candidate chunk together in a single forward pass, producing a joint relevance score. Standard vector search scores query and chunk independently and compares their embeddings by cosine similarity, which the README notes is less precise.

## Sources

- [Issues](https://github.com/LeDat98/NexusRAG/issues)
- [LeDat98/NexusRAG on GitHub](https://github.com/LeDat98/NexusRAG)
- [README](https://github.com/LeDat98/NexusRAG/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/ledat98-nexusrag
