IntelLabs/RAG-FiT: a four-stage pipeline for fine-tuning LLMs on RAG data
Framework for enhancing LLMs for RAG tasks using fine-tuning.
At a glance
- What is it?
- RAG-FiT (formerly RAG Foundry) builds RAG-augmented datasets, fine-tunes models on them with PEFT and TRL, then scores the output with RAG-specific metrics. It is a research harness, not a retrieval stack, and the Hydra configs are the whole interface.
- Who is it for?
- Adopt RAG-FiT if you need to fine-tune a model on retrieval-augmented examples and you already have a retrieval setup, since the framework consumes retrieval output rather than providing it. Skip it if you want a hosted pipeline or a retriever you do not have to build.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 114 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Who RAG-FiT is for, and the gap it fills
Most RAG work happens at inference time: retrieve passages, paste them into a prompt, generate. RAG-FiT targets the other half of the problem. The README describes a library that improves an LLM's ability to use external information by fine-tuning models on specially created RAG-augmented datasets. So the intended user is someone who has already decided that prompting alone is not enough and wants to train the model on examples that include retrieved context.
The audience is narrow on purpose. You need a retrieval source, a corpus, and enough compute to run parameter-efficient fine-tuning. The library does not supply a retriever, a vector store, or a serving layer. What it supplies is the middle: turning retrieval results into training records, running the training loop, generating predictions, and then scoring them. The repository layout reflects that split, with processing.py, training.py, inference.py and evaluation.py sitting at the top level as four separate entry points.
That separation is the main design decision. Each stage writes artifacts the next stage reads, so you can rerun evaluation without retraining, or regenerate a dataset without touching a checkpoint. The cost is that you own the glue between stages: paths, configs, and the discipline to keep the dataset schema stable.
The four modules and how data moves between them
The processing module does the heaviest lifting. According to the README, it handles dataset loading, column normalization, data aggregation for fewshot creation, information retrieval through external tools and frameworks, API integration, and template-based prompt creation. The output is persisted in what the README calls a consistent, model-independent, input-output format, along with other fields and metadata. That format is the contract between stages, and it is also where retrieval results, reasoning traces, citations and attributions live so that metrics can reach them later.
Training uses PEFT for efficient training and TRL, for example supervised fine-tuning, and the README states that training is done on the completions. Models can be pushed to the Hugging Face Hub. Inference generates predictions over the augmented datasets using either trained or untrained LLMs, which makes it possible to compare a fine-tuned checkpoint against its base model on the same inputs.
Evaluation is where the framework is most opinionated. It runs a list of metrics that you provide, and custom metrics can be implemented. The documented set includes EM, F1, ROUGE, BERTScore, Deepeval, RAGAS, the Hugging Face evaluate library, and classification. Metrics are classified as local (run per example) or global (run over the whole dataset, with recall as the example). The README notes that metrics can use any feature in the dataset, not just input and output text. That is the point of persisting retrieval results alongside the prompt: you can score whether the retrieved passages were actually used.
Installing RAG-FiT and running a first processing job
The README gives one installation path: clone the repository and install it in editable mode. Python 3.10 or newer is required according to pyproject.toml.
pip install -e .Two optional dependency groups exist. The haystack extra pulls haystack-ai and qdrant-haystack, and the deepeval extra pulls deepeval, which is the package behind the Deepeval metric family.
pip install -e .[haystack]
pip install -e .[deepeval]Every module is invoked as a script with Hydra options. The README shows the paper configs as the worked example, using the ASQA dataset. The -cp flag sets the config directory and -cn selects the config name.
python processing -cp configs/paper -cn processing-asqa-retrievalAny individual key can be overridden on the command line. The README demonstrates changing the output path and enabling the cache in the same call.
python processing -cp configs/paper -cn processing-asqa-retrieval \
output_path=/store/data/here \
cache=trueAfter this runs, you should have a dataset file at the path you gave in output_path, containing the RAG interactions in the input-output format described in docs/processing.md. The README points to the PubmedQA tutorial in docs/pubmed.md for a shorter end-to-end walkthrough, and to the configs/paper folder for the full ASQA reproduction.
Hydra configuration is the interface, and that is a real constraint
RAG-FiT does not expose a Python API in the README. The documented way to use it is configuration-as-code: Hydra instantiates Python classes based on the _target_ keyword in a config, and the configs folder holds defaults for each module. This buys hierarchical configs, CLI overrides, and the ability to run multiple jobs remotely through integrations with SLURM and Ray.
It also means the learning curve is Hydra's, not RAG-FiT's. If you have not worked with _target_ instantiation before, the configs are the documentation, and the prose docs in docs/processing.md, docs/training.md, docs/inference.md and docs/evaluation.md are the map. There is no builder function you can call from a notebook to get a dataset object back. For interactive experimentation this is friction.
The trade-off is deliberate and defensible for the stated use case. Reproducing a paper means running the same pipeline many times with small variations, and a config file per variation is easier to diff and archive than a notebook. If your work is exploratory rather than comparative, the config layer will feel like overhead.
Dependency pins and what they imply about maintenance
pyproject.toml pins aggressively. datasets is fixed at 2.16.1, transformers at 4.50.0 or newer, peft at 0.11.1, trl at 0.8.6, openai at 1.23.3, and bitsandbytes at exactly 0.42.0. torch is specified as 2.8.0 or newer. The exact pins on peft, trl and bitsandbytes are the ones most likely to fight with an existing environment, because those three move quickly and are frequently constrained by CUDA builds.
The last push to the default branch was on 2026-06-08, which is recent enough that the project is not abandoned, but the most recent release listed is v1.5.0 from 2024-11-12, and pyproject.toml still declares version 1.2.0. That mismatch between the released tag and the version string in the packaging file is worth knowing before you file a bug about version reporting.
Licensing is Apache-2.0, applied to the code in the repository. The README carries a disclaimer that this is not an official Intel product, which matters if you were planning to treat the project as vendor-supported. Apache-2.0 permits commercial use and modification, but the usual obligations around notices and attribution apply, and the third-party models and datasets you point the pipeline at carry their own terms. That part is not legal advice; check the licences of whatever you retrieve and train on.
Where RAG-FiT is the wrong tool, and what to use instead
RAG-FiT will not retrieve anything for you by default. The README describes information retrieval using external tools and frameworks, and the haystack extra adds haystack-ai with qdrant-haystack, so retrieval is something you wire in through config rather than something the library ships. If your problem is that you have documents and no search over them, this project is the wrong starting point. You would be better served by a retrieval framework such as Haystack on its own, or LlamaIndex, and then coming back to RAG-FiT once you have passages and a training set worth building.
A second mismatch is scale of ambition. If a prompt change gets you the accuracy you need, fine-tuning is a large amount of work for a marginal gain, and RAG-FiT makes that work reproducible rather than cheap. The framework also assumes you can run training, which means GPU memory and a checkpoint to manage.
The closest alternative in spirit is the evaluation and training tooling around the Hugging Face ecosystem itself: trl for supervised fine-tuning, peft for adapters, evaluate for metrics. RAG-FiT composes those same libraries, so the difference is not the underlying machinery. It is the dataset layer in between: RAG-FiT persists retrieval interactions in a schema that later stages can score against, including retrieved passages, reasoning, citations and attributions. If you build the loop yourself with trl and evaluate, you own that schema and its consistency. If you use RAG-FiT, you inherit one that the paper configs already exercise.
What to check before you commit
Start with the PubmedQA tutorial in docs/pubmed.md rather than the paper configs. It is the smallest complete path through all four modules, and it will surface dependency problems before you have invested in a corpus.
Then verify three things. First, that your retrieval output can be expressed in the input-output format the processing module persists, since every later stage depends on that schema. Second, that the pinned peft, trl and bitsandbytes versions resolve in your environment alongside torch 2.8.0 or newer. Third, that the metrics you care about are in the documented set or can be implemented as a custom metric, because the evaluation module takes a list you supply rather than a fixed suite.
If those three hold, the framework does something specific and useful: it makes the RAG fine-tuning loop reproducible from config files, and it keeps retrieval artifacts around long enough to measure whether the model used them.
Editorial conclusion
Adopt RAG-FiT if you need to fine-tune a model on retrieval-augmented examples and you already have a retrieval setup, since the framework consumes retrieval output rather than providing it. Skip it if you want a hosted pipeline or a retriever you do not have to build. Before committing, check that the pinned dependency set in pyproject.toml installs against your CUDA and PyTorch versions, and read docs/processing.md to confirm your corpus format fits the input-output schema it persists.
Frequently asked questions
What does RAG mean in the context of RAG-FiT?
RAG stands for retrieval augmented generation, and RAG-FiT treats it as a set of interactions to be captured and trained on rather than only a prompting pattern. The README describes the library as improving an LLM's ability to use external information by fine-tuning on specially created RAG-augmented datasets.
How do I install RAG-FiT?
Clone the repository and install it in editable mode with pip install -e . The README also lists optional extras, pip install -e .[haystack] and pip install -e .[deepeval], for the Haystack retrieval integration and the Deepeval metrics respectively.
Does RAG-FiT provide its own retriever?
No. The processing module performs information retrieval using external tools and frameworks, and the optional haystack extra adds haystack-ai with qdrant-haystack. You supply the retrieval setup and RAG-FiT persists the results into the training dataset.
What metrics can RAG-FiT evaluate with?
The README lists EM, F1, ROUGE, BERTScore, Deepeval, RAGAS, the Hugging Face evaluate library, and classification, and notes that custom metrics can be implemented. Metrics are either local, run per example, or global, run over the whole dataset such as recall.
Is RAG-FiT an official Intel product?
The README carries an explicit disclaimer that this is not an official Intel product, even though the repository lives under the IntelLabs organization. The code is licensed under Apache-2.0.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/intellabs-rag-fit)