RAGLight: Modular Python RAG Framework with Hybrid Search and MCP Integration
RAGLight is a modular framework for Retrieval-Augmented Generation (RAG). It makes it easy to plug in different LLMs, embeddings, and vector stores, and now includes seamless MCP integration to connect external tools and data sources.
At a glance
- What is it?
- RAGLight is a Python library for building Retrieval-Augmented Generation pipelines that swap LLMs, embeddings, and vector stores through a common interface. Version 3.4.7 adds MCP integration, hybrid BM25 plus semantic search with Reciprocal Rank Fusion, query reformulation, and streaming output across all supported providers.
- Who is it for?
- RAGLight is a practical choice for Python developers who want a RAG pipeline without writing the retrieval, reranking, and multi-provider adapter layers from scratch, and who want to avoid vendor lock-in to a single LLM. Python 3.11 or newer is required.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 29 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What RAGLight Solves and Who It Is For
Retrieval-Augmented Generation pipelines have a predictable set of components: a document ingestion step, an embedding model to vectorize chunks, a vector store for retrieval, a reranker to reorder candidates, and an LLM to generate the final answer. The problem is that each component has competing implementations, and swapping one (say, switching from Ollama to AWS Bedrock as the LLM) typically requires rewriting the integration code.
RAGLight addresses this by providing a common interface over those components. The library is described in its README as lightweight and modular: LLMs, embeddings, and vector stores are pluggable. It is aimed at Python developers who want to build context-aware AI applications on top of their own documents, whether that is an internal knowledge base, a codebase, or a collection of PDFs, without being committed to one vendor.
The library is available on PyPI as `raglight` (version 3.4.7 as of the most recent release) and requires Python 3.11 or newer.
Installing RAGLight and Its Optional Extras
The base library installs with:
pip install raglightVector store backends are optional extras to avoid pulling in large C++ build dependencies by default. The README documents three extras:
pip install "raglight[qdrant]" # Qdrant only (Windows-friendly)
pip install "raglight[chroma]" # ChromaDB only
pip install "raglight[chroma,qdrant]" # both
pip install "raglight[qdrant,langfuse]" # Qdrant + observabilityThe README notes that ChromaDB requires a C++ compiler on Windows, while the Qdrant client is pure Python and works on Windows without one. Langfuse observability tracing is a third extra.
Once installed, the `raglight` CLI is available. The fastest way to start is the interactive chat wizard:
raglight chatThe wizard prompts for a document folder, which folders to exclude (virtual environments, node_modules, __pycache__, build artifacts, and IDE directories are excluded by default), a vector database name, an embedding model, and an LLM. After setup it indexes the documents and opens a chat session. An agentic variant of the wizard starts with `raglight agentic-chat`.
For serving over HTTP, `raglight serve` starts a FastAPI server with a Streamlit chat UI and REST endpoints for ingestion and queries.
Supported LLM Providers and Vector Stores
RAGLight supports seven LLM providers: Ollama, Google Gemini, LMStudio, vLLM, OpenAI API, Mistral API, and AWS Bedrock. The .env.example documents the configuration pattern: RAGLIGHT_LLM_PROVIDER sets the provider name (e.g., OpenAI), RAGLIGHT_LLM_MODEL sets the model name (e.g., gpt-4.1-mini), and RAGLIGHT_LLM_API_BASE optionally overrides the endpoint URL.
For embeddings, the same provider pattern applies: RAGLIGHT_EMBEDDINGS_PROVIDER and RAGLIGHT_EMBEDDINGS_MODEL. The README cites HuggingFace's all-MiniLM-L6-v2 as an example embedding model for compact vector representations.
Two vector stores are supported: ChromaDB and Qdrant. The .env.example shows RAGLIGHT_VECTOR_STORE_PROVIDER and RAGLIGHT_VECTOR_STORE_NAME as the configuration keys. The Makefile's install target installs both with `uv pip install -e ".[chroma,qdrant]"`.
AWS Bedrock deserves specific mention. The README states that Bedrock credentials load through the standard AWS credential chain (environment variables, ~/.aws/credentials, or IAM role) with no extra install step. This makes it straightforward to use in environments that already have AWS auth configured.
Hybrid Search, Query Reformulation, and Streaming
RAGLight's retrieval layer goes beyond plain vector search. The hybrid search feature combines BM25 keyword retrieval with dense semantic search and applies Reciprocal Rank Fusion to merge the result lists. The README labels this 'best-of-both-worlds results.' BM25 handles exact keyword matches that semantic search can miss; semantic search handles paraphrase and concept matches that keyword search misses.
Query reformulation rewrites follow-up questions into standalone queries using conversation history. The README explains the motivation: in a multi-turn conversation, a question like 'What does it say about rollback?' depends on context from earlier turns. Query reformulation makes the retrieval query self-contained so the vector store can answer it correctly without context.
Streaming output is available on all providers through a `generate_streaming()` method. The README states it is a drop-in alongside `generate()` with no extra configuration, meaning switching from batch to streaming requires only changing the method call.
Conversation history is supported across all providers with an optional `max_history` cap to bound token usage. Langfuse observability (version 3 and later) traces every RAG call end-to-end through retrieve, rerank, and generate stages in a Langfuse dashboard.
MCP Integration and Agentic RAG
RAGLight version 3.x adds MCP (Model Context Protocol) integration through the langchain-mcp-adapters package. MCP servers expose tools (code execution, database access, external APIs) that an agentic RAG pipeline can call during query resolution. The README marks this with a plug icon in the feature list.
The example file `examples/agentic_rag_with_mcp.py` demonstrates the pattern. An agentic RAG pipeline is one where the LLM can decide to call retrieval, call an MCP tool, or generate directly based on the query, rather than always passing through the vector store first.
The `raglight agentic-chat` CLI command launches an agentic variant of the interactive wizard. The distinction from standard RAG is that the agent reasons about whether to retrieve before answering, which reduces unnecessary retrieval on questions the LLM can answer from its own knowledge.
Limitations and What RAGLight Does Not Cover
RAGLight requires Python 3.11 or newer, as stated in pyproject.toml. Older Python environments are not supported.
The library is a local-first tool: it runs the embedding model and vector store on the same machine as the application. Distributed vector store deployments (Qdrant cluster, hosted Chroma) are possible through the provider configuration but are not documented in the examples.
The chunking strategy is described as structure-aware but the README does not specify the chunking algorithm or the chunk size defaults. Workloads with long documents that need precise chunk control will require digging into the source or opening the question in the project issues.
Langfuse observability is an optional extra rather than built-in telemetry. Teams that need observability from the start must include `raglight[langfuse]` at install time. The README does not describe what happens if you add the extra to an existing deployment with already-indexed documents.
The most direct alternative is LangChain, which covers the same RAG component categories with a larger ecosystem and more adapters. RAGLight's advantage over LangChain is a smaller default dependency footprint and a CLI that handles the full setup wizard without writing any Python.
Maintenance and License
The last push to the repository was on 2026-09-02, and version 3.4.7 is the current PyPI release. The project has published multiple minor releases in 2026 (3.1.1 in March, 3.2.0 in March, 3.4.7 in March). The MIT license permits commercial use, modification, and redistribution. The pyproject.toml shows the author as Bessouat40 with a contact email at orange.fr. The repository includes a CONTRIBUTING.md and a Makefile with format, test, and install targets.
Editorial conclusion
RAGLight is a practical choice for Python developers who want a RAG pipeline without writing the retrieval, reranking, and multi-provider adapter layers from scratch, and who want to avoid vendor lock-in to a single LLM. Python 3.11 or newer is required. Before starting, verify that your chosen LLM provider is running: Ollama must be active locally, LMStudio must have the target model loaded, and AWS Bedrock requires configured credentials. Teams that need production observability from day one should include the langfuse extra at install time, since adding it later requires reindexing to get traces on historical queries.
Frequently asked questions
What is RAGLight Python?
RAGLight is a Python library for building Retrieval-Augmented Generation pipelines. It provides a modular interface for swapping LLMs (Ollama, OpenAI, Mistral, Gemini, AWS Bedrock), embedding models, and vector stores (ChromaDB, Qdrant) without rewriting integration code. It installs from PyPI with `pip install raglight`.
Does RAGLight support streaming output?
Yes. All LLM providers in RAGLight support token-by-token streaming through a generate_streaming() method. The README states it is a drop-in alongside generate() with no extra configuration required.
Can RAGLight connect to external tools through MCP?
Yes. RAGLight version 3.x includes MCP integration via langchain-mcp-adapters. An agentic RAG pipeline can use MCP servers to call external tools such as code execution or database access during query resolution. The raglight agentic-chat CLI command provides an interactive wizard for this mode.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/bessouat40-raglight)