RAG-Retrieval: one training and inference stack for embedding, ColBERT and reranker models
Unify Efficient Fine-tuning of RAG Retrieval, including Embedding, ColBERT, ReRanker.
At a glance
- What is it?
- RAG-Retrieval bundles fine-tuning code for embedding, late-interaction and reranker models with a small pip-installable library for calling rerankers behind one interface. It is MIT licensed and the last push to the repository was on 2026-08-28.
- Who is it for?
- Adopt RAG-Retrieval if you already know which base checkpoint you want to fine-tune and you need one repository that covers embedding, ColBERT and reranker training with DeepSpeed or FSDP. Do not adopt it if you want a managed retrieval service, a hosted index, or a one-command pipeline from raw documents to a running RAG application; the project stops at the model layer.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 33 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What RAG-Retrieval solves and who it is aimed at
Retrieval quality in a RAG system is usually decided by three separate model families: the embedding model that turns passages into vectors, a late-interaction model such as ColBERT that scores token-level matches, and a reranker that rescores a short candidate list. Each family has its own training scripts, its own loss conventions and its own inference wrapper. RAG-Retrieval's stated goal is to put all three under one repository so that a team fine-tuning a bge or bce checkpoint does not have to learn three codebases.
The README frames the project as end-to-end code for training, inference and distillation. The audience is fairly narrow: engineers who already have labelled query-passage pairs and a GPU budget, and who want to fine-tune an open-source retrieval model rather than call a hosted API. The project explicitly lists bge (bge-embedding, bge-m3, bge-reranker), bce (bce-embedding, bce-reranker) and gte (gte-embedding, gte-multilingual-reranker-base) as compatible starting points. If you are building a first RAG prototype and have not yet decided whether retrieval quality is your bottleneck, this is more machinery than you need.
How the training, distillation and inference pieces fit together
The repository splits into a training tree under rag_retrieval/train/ and an inference package under rag_retrieval/infer/reranker_models. Training code is organised by model type rather than by dataset, so embedding, colbert and reranker work lives in separate subdirectories with their own README files and shell entry points. The pyproject.toml packages only rag_retrieval and rag_retrieval.infer.reranker_models for distribution, which tells you the training scripts are meant to be used from a checkout, not from the installed wheel.
On the inference side the project ships a lightweight library whose job is to normalise how different rankers are called. The README describes two supported families, Cross Encoder Reranker and Decoder-Only LLM Reranker, and two strategies for long documents: truncation to a maximum length, or splitting the document and taking the maximum score across chunks. Extension is by inheritance: a new ranking model is added by subclassing BaseReranker and implementing the rank and compute_score functions.
Distillation is the third leg. The README states that larger LLM-based reranker or embedding models can be distilled into smaller ones, naming a 0.5B-parameter LLM or BERT-base as targets, and the examples directory contains distill_llm_to_bert_reranker and stella_embedding_distill as worked cases. For embedding models specifically, the project implements the MRL algorithm for reducing output vector dimensionality and the Stella distillation method described in arXiv:2412.19048. Multi-GPU training is handled through deepspeed and fsdp.
Installing RAG-Retrieval and running a first training job
The README gives two separate installation paths, and the split matters. The training path installs from a checkout of the repository and pulls in requirements.txt, which lists accelerate, transformers, torch>=2.1.0, tqdm and sentence_transformers. The prediction path installs the published package from PyPI. In both cases the README advises installing a compatible torch build manually first, because the automatically resolved version may not match your local CUDA.
Start with the training environment:
conda create -n rag-retrieval python=3.8 && conda activate rag-retrieval
# install a torch build matching your local CUDA before the next step
pip install -r requirements.txtAfter that, training runs from the subdirectory for the model type you care about. The README uses embedding as the example, and points to the README inside each subdirectory for the full procedure:
cd ./rag_retrieval/train/embedding
bash train_embedding.shFor reranker inference only, the published package is a single install:
pip install rag-retrievalWhat you should see after the training command is the shell script driving the Python training entry point in that directory; the README does not print expected loss curves or step counts, so read the subdirectory README before assuming defaults. The package version in pyproject.toml is 0.2.2 and requires Python 3.8 or newer, with classifiers up to 3.12.
Where RAG-Retrieval is the wrong choice
The project is a model toolkit, not a retrieval system. There is no index, no vector store integration, no document loader and no serving layer in what the README describes. If your problem is that you cannot get documents into a searchable store, or that your query latency is too high, RAG-Retrieval does not address either. You would pair it with something else for storage and serving.
A second boundary is the experimental record. The README's reranker results table is incomplete: the rag-retrieval-reranker row lists T2Reranking, MMarcoReranking, CMedQAv1 and CMedQAv2 values but no final average, and the model-size column shows 0.41 GB against 1.11 GB for bge-reranker-base and bce-reranker-base_v1. The table also does not state which checkpoint produced those numbers or how it was trained, so it is not enough on its own to justify swapping an existing reranker. Treat it as a starting point for your own evaluation.
Finally, the training path assumes you bring data. The examples directory includes synthetic_data_embedding, but the README does not describe a data preparation pipeline, a labelling workflow or a quality filter. If you do not already have query-passage pairs, the fine-tuning half of this repository is not usable yet.
How it differs from Sentence Transformers and FlagEmbedding
Sentence Transformers is the obvious alternative for the embedding half, and it is a direct dependency here (sentence_transformers appears in requirements.txt). The difference is scope. Sentence Transformers gives you a training loop and a model zoo centred on embedding models; RAG-Retrieval wraps embedding, ColBERT-style late interaction and reranker training in one tree, and adds distillation from a larger LLM-based model down to a 0.5B LLM or BERT-base. If you only ever train bi-encoders, Sentence Transformers is the smaller dependency and the more documented path.
BAAI's FlagEmbedding is the closer comparison, since the README names bge-embedding, bge-m3 and bge-reranker as supported starting points. FlagEmbedding is organised around BAAI's own model family and its training recipes. RAG-Retrieval is organised around model type instead, and its published artifact is a reranker inference library with a BaseReranker subclassing contract, which is a different thing to adopt: you are buying a uniform call surface for rankers more than a specific checkpoint. If your stack is already all-BGE, FlagEmbedding keeps you on one vendor's recipes; if you mix gte, bce and bge checkpoints and want one training layout, RAG-Retrieval's structure is the argument for it.
Maintenance, licence and upgrade cost
The repository is not archived and the last push was on 2026-08-28, so the code is being touched. That is not the same as a support commitment: the README does not describe a release cadence, a deprecation policy or a compatibility guarantee between the training tree and the published package. The only listed release, rag_retrieval_only_train (RAG-Retrieval v0.1), dates from 2024-05-04, while pyproject.toml declares version 0.2.2, so the release feed and the package metadata are not in step. Plan to pin the version you install.
Upgrade cost is concentrated in the training tree. Because training scripts are run from a checkout and are not packaged, changes under rag_retrieval/train/ reach you only when you pull the repository, and there is no changelog described anywhere in the README covering what changed between versions. The inference package is the stable surface: it packages rag_retrieval.infer.reranker_models and depends on pydantic, tqdm, torch and transformers, a short list that keeps breakage risk low.
The licence is MIT, declared in both the LICENSE file and the pyproject.toml license field. MIT is permissive, so redistribution and commercial use are broadly allowed, but the repository also cites third-party methods and papers (MRL, Stella) and links to a paper on distilling SOTA embedding models. Whether the checkpoints you fine-tune carry their own licences is a separate question from this repository's licence, and the README does not answer it.
Editorial conclusion
Adopt RAG-Retrieval if you already know which base checkpoint you want to fine-tune and you need one repository that covers embedding, ColBERT and reranker training with DeepSpeed or FSDP. Do not adopt it if you want a managed retrieval service, a hosted index, or a one-command pipeline from raw documents to a running RAG application; the project stops at the model layer. Before committing, verify that the subdirectory README under rag_retrieval/train/embedding matches your GPU count and that your CUDA build is compatible with the torch version you install manually, since the installation notes warn that the automatically resolved torch may not match your local CUDA.
Frequently asked questions
What is RAG-Retrieval used for?
It provides end-to-end code for training, inference and distillation of RAG retrieval models, covering embedding models, ColBERT-style late-interaction models and reranker models. The inference side is a lightweight Python library that calls different reranker models through one interface.
How do I install RAG-Retrieval for reranker inference?
The README gives pip install rag-retrieval for prediction, and advises installing a torch build compatible with your local CUDA first. Training uses a separate path, a conda environment with python=3.8 followed by pip install -r requirements.txt from a checkout.
Does RAG-Retrieval support reranking?
Yes. The published library supports Cross Encoder Reranker and Decoder-Only LLM Reranker models, and handles long documents either by truncating to a maximum length or by splitting the document and taking the maximum score.
Can RAG-Retrieval distill a large model into a small one?
The README states that distillation of embedding and reranker models is supported, from a larger model down to a smaller one such as a 0.5B-parameter LLM or BERT-base. The examples directory includes distill_llm_to_bert_reranker and stella_embedding_distill as worked cases.
What licence does RAG-Retrieval use?
The repository is MIT licensed, declared in the LICENSE file and in the pyproject.toml license field. The licences of the individual base checkpoints you fine-tune are a separate matter that the README does not cover.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/novasearch-team-rag-retrieval)