# RAGLite: a permissive Python RAG toolkit on DuckDB or PostgreSQL

> RAGLite is a Python toolkit that wires document conversion, hybrid search, reranking and LLM calls into one pipeline, backed by DuckDB or PostgreSQL. It targets teams that want a small dependency footprint and no PyTorch or LangChain in the tree.

**superlinear-ai/raglite** — 🥤 RAGLite is a Python toolkit for Retrieval-Augmented Generation (RAG) with DuckDB or PostgreSQL

- Repository: https://github.com/superlinear-ai/raglite
- Stars: 1,208 · Forks: 109
- Language: Python
- License: MPL-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/superlinear-ai-raglite

## The problem RAGLite targets: RAG without a framework tax

Most RAG stacks accumulate weight. A vector store, an orchestration library, a document loader, a reranker, and suddenly a deployment carries PyTorch and a dependency graph nobody can audit. RAGLite takes the opposite position. Its README states the project depends on "only lightweight and permissive open source dependencies (e.g., no PyTorch or LangChain)", and the pyproject.toml backs that up: the runtime list is duckdb, pg8000, sqlmodel-slim, litellm, rerankers, pdftext, markdown-it-py, numpy, scipy, wtpsplit-lite and a handful of utilities. That is the whole pitch. You get document conversion, chunking, hybrid retrieval, reranking, generation and evaluation in one library, and you can read the dependency list in under a minute.

The intended user is a Python engineer who already has a corpus and an LLM key, and who wants retrieval working this week without adopting an orchestration framework. The storage choice is the clearest signal of that audience: DuckDB for a local file, PostgreSQL for a shared instance, with the same code path. There is no managed control plane and no dashboard, which is a deliberate trade rather than an oversight.

## How the retrieval pipeline is assembled

The architecture is a linear pipeline with a database at the centre. Documents enter as PDFs (or any format Pandoc handles, if the pandoc extra is installed), get converted to Markdown, and are split into sentences and chunks. RAGLite does the splitting by solving a binary integer programming problem, both for sentence splitting via wtpsplit-lite and for semantic chunking, rather than using fixed-size windows. Chunks are then embedded as multi-vectors using late chunking and contextual chunk headings.

Retrieval is hybrid: the README describes using the database's native keyword and vector search, which means DuckDB's FTS and VSS extensions or PostgreSQL's tsvector and pgvector. Results are fused and passed to a reranker, with FlashRank as the default and other rerankers available through the rerankers library. Generation goes through LiteLLM, so any provider LiteLLM supports is available, plus local llama.cpp models through llama-cpp-python.

Two details are worth separating from the feature list. First, adaptive retrieval: the README says the LLM decides whether to retrieve and what to retrieve based on the query, which means an extra model call before search on some queries. Second, the closed-form linear query adapter in src/raglite/_query_adapter.py, computed by solving an orthogonal Procrustes problem. That is a trained transformation between query and document embedding spaces, and it only helps if you compute it from your own data.

## Installing RAGLite and running a first insert and query

The base install is a single pip command. The README gives it directly:

```bash
pip install raglite
```

Optional capabilities are extras: raglite[chainlit] for the ChatGPT-like frontend, raglite[pandoc] for non-PDF input formats, and raglite[ragas] for evaluation. Mistral OCR support is not an extra; the README says to install mistralai separately.

For local models, the README recommends an accelerated llama-cpp-python precompiled wheel and shows the environment variables that select it. Not every combination of version, accelerator and platform is published, so the project warns about that in the snippet itself:

```bash
LLAMA_CPP_PYTHON_VERSION=0.3.9
PYTHON_VERSION=310|311|312
ACCELERATOR=metal|cu121|cu122|cu123|cu124
PLATFORM=macosx_11_0_arm64|linux_x86_64|win_amd64
pip install "https://github.com/abetlen/llama-cpp-python/releases/download/v$LLAMA_CPP_PYTHON_VERSION-$ACCELERATOR/llama_cpp_python-$LLAMA_CPP_PYTHON_VERSION-cp$PYTHON_VERSION-cp$PYTHON_VERSION-$PLATFORM.whl"
```

Configuration starts with RAGLiteConfig, which takes the database URL and the LiteLLM model identifier. The README's configuration section shows a Python snippet beginning with `from raglite import RAGLiteConfig` and describes a "remote" config using a PostgreSQL database plus a hosted model. A llama.cpp model is selected with the identifier form `llama-cpp-python/<hugging_face_repo_id>/<filename>@<n_ctx>`, where n_ctx is optional.

For a local PostgreSQL, the repository ships a docker-compose.yml that runs pgvector/pgvector:pg17 with the user raglite_user and password raglite_password, and mounts the data directory as tmpfs, so the database is wiped when the container stops. That last part matters: it is a development setup, not a deployment recipe. The README also notes you can create a PostgreSQL database in a few clicks at neon.tech if you would rather not run one.

Once configured, the workflow order in the README is: insert documents, run RAG, optionally compute a query adapter, optionally evaluate, and optionally expose the result over MCP or Chainlit. The package also installs a CLI entry point, `raglite = "raglite._cli:cli"`, so a terminal command exists alongside the Python API.

## Where RAGLite stops being the right tool

The dependency constraint is also the ceiling. RAGLite does not ship a document loader ecosystem. PDF is the first-class format; everything else routes through Pandoc, and Mistral OCR is a separate install for PDFs, images, DOCX and PPTX. If your corpus is a pile of HTML, email, or proprietary formats, you are writing the conversion layer yourself.

Storage is the second boundary. DuckDB gives you a file, which is excellent for a laptop and awkward for concurrent writers. PostgreSQL removes that limit, but the repository's own compose file uses tmpfs for the data directory, so anyone copying it into production without changing that line will lose their index on restart. The README does not document a migration path between the two backends, so the choice is worth making early.

The heavier mechanisms are also the ones with the least documentation. Late chunking requires an embedding model that supports it; the README describes the technique but does not list which models qualify. The query adapter is presented as an optimization to compute from your data, and the README's own step list places it after basic RAG works, which is the right ordering but also means an unadapted system is the default. Adaptive retrieval adds a model call to decide whether to search at all, and that call is not free in either latency or tokens. Finally, the project has no hosted offering and no web UI beyond the optional Chainlit frontend, so operational concerns such as index rebuilds and backup are yours.

## RAGLite against LangChain and LlamaIndex

The honest comparison is not feature-by-feature, because RAGLite is deliberately smaller. LangChain and LlamaIndex both provide large integrations ecosystems: hundreds of loaders, retrievers, memory modules and agent abstractions, plus hosted tracing services. RAGLite provides one opinionated pipeline and refuses the framework role. The README names LangChain explicitly as something it avoids, which tells you the maintainers see the dependency footprint itself as the differentiator.

If your problem is "connect eleven SaaS sources and let an agent choose tools", the larger frameworks are the better fit and RAGLite will feel like a straitjacket. If your problem is "index these PDFs, retrieve well, and answer with citations", the extra surface area in a general framework is mostly maintenance you did not ask for. The same logic applies against a hosted RAG service: those remove the database and the pipeline from your plate entirely, at the cost of sending your corpus to someone else's infrastructure. RAGLite keeps both the corpus and the index local, which is the point of choosing DuckDB in the first place.

## Licence, release cadence and upgrade cost

RAGLite is licensed under MPL-2.0, a file-level copyleft licence. You can use it in a closed-source product; the obligation attaches to modifications of MPL-covered files, not to your application as a whole. That is a summary of the licence's structure, not legal advice, and anything beyond that belongs with your counsel.

The version in pyproject.toml is 1.1.1, matching the v1.1.1 release dated 2026-05-18. The previous releases were v1.0.0 on 2025-06-11 and v0.7.0 on 2025-03-17, so the project moves in occasional larger steps rather than continuous small ones. The last push to the default branch was on 2026-08-17, which is recent enough that the repository is not dormant.

Upgrade cost concentrates in the pinned dependencies. The pyproject.toml pins numpy below 2.0.0, excludes specific scipy versions, excludes duckdb 1.4.0, and caps onnxruntime below 1.24.0 for Python 3.10. Those pins exist for real reasons, but they mean RAGLite can hold back a transitive dependency in a larger environment. The supported Python range is >=3.10,<3.14. If you run a monorepo where another package needs numpy 2.x, that conflict surfaces at install time, not at runtime.

## Conclusion

Adopt RAGLite if you want a RAG pipeline whose storage is a single DuckDB file or an existing PostgreSQL instance, and whose dependency list avoids PyTorch and LangChain. Skip it if you need a hosted service, a visual pipeline builder, or a framework with a large plugin ecosystem, because RAGLite is a library plus a CLI and expects you to write the Python. Before committing, verify that your Python version falls inside the declared >=3.10,<3.14 range, and check whether your embedding model supports late chunking, since the multi-vector embedding path depends on it.

## FAQ

### Is RAGLite a RAG model?

No. RAGLite is a Python toolkit that implements retrieval-augmented generation: it converts documents, indexes them in DuckDB or PostgreSQL, retrieves and reranks chunks, and calls an LLM through LiteLLM. The model itself comes from whichever provider or local llama.cpp file you configure.

### What is RAG and why does RAGLite use it?

Retrieval-augmented generation means fetching relevant passages from your own documents and passing them to an LLM alongside the question. RAGLite exists to make that pipeline concrete: hybrid keyword and vector search in the database, reranking on top, and generation through any LiteLLM-supported provider.

### Is Perplexity built on RAG?

The material for RAGLite does not cover Perplexity, so this cannot be answered from the project's documentation. RAGLite's own README describes hybrid search, reranking and adaptive retrieval rather than any comparison with Perplexity.

### How is RAGLite different from an LLM on its own?

An LLM answers from its training data. RAGLite adds a retrieval stage before generation: documents are chunked and embedded, hybrid search and reranking select the passages, and only those passages reach the model. The README also describes adaptive retrieval, where the LLM decides whether to retrieve at all.

## Sources

- [Issues](https://github.com/superlinear-ai/raglite/issues)
- [License: MPL-2.0](https://github.com/superlinear-ai/raglite/blob/main/LICENSE)
- [README](https://github.com/superlinear-ai/raglite/blob/main/README.md)
- [Releases](https://github.com/superlinear-ai/raglite/releases)
- [superlinear-ai/raglite on GitHub](https://github.com/superlinear-ai/raglite)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/superlinear-ai-raglite
