Model or dataset
microsoft/rag-experiment-accelerator avatar
microsoft/rag-experiment-accelerator

RAG Experiment Accelerator: A Config-Driven Pipeline for Testing Azure AI Search and RAG Setups

The RAG Experiment Accelerator is a versatile tool designed to expedite and facilitate the process of conducting experiments and evaluations using Azure Cognitive Search and RAG pattern.

312 stars111 forksPythonNOASSERTION

At a glance

What is it?
microsoft/rag-experiment-accelerator is a Python tool that runs systematic experiments comparing search hyperparameters, document loaders, chunking strategies, and query types against Azure AI Search and Azure OpenAI. It generates reports and visualizations from the results, making it a structured alternative to ad-hoc RAG tuning.
Who is it for?
RAG Experiment Accelerator is appropriate for teams building RAG pipelines on Azure who need a structured way to compare retrieval configurations rather than tuning by intuition. It requires Azure AI Search at Basic tier or higher (for semantic search), Azure OpenAI or access to the OpenAI API, and optionally Azure Machine Learning for MLflow tracking.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 103 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 22, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What the Accelerator Is For and Who Uses It

RAG systems fail in non-obvious ways. The choice of chunk size, chunk overlap, embedding model, search type (pure vector, hybrid, semantic), re-ranking strategy, and query decomposition all affect retrieval quality. Testing each combination by hand is slow and produces results that are hard to compare across runs. There is also no standard for what counts as a fair comparison: different search types use different indexes, different embeddings produce different latency profiles, and different document loaders extract different text from the same file.

RAG Experiment Accelerator automates that comparison. It reads a configuration file that specifies which hyperparameters to vary, creates separate Azure AI Search indexes for each configuration, generates a question-and-answer set from the indexed documents, runs queries against each index, evaluates the results against ground-truth answers, and produces reports. Researchers, data scientists, and developers building Azure-hosted RAG systems are the described audience. The tool is config-driven: a single JSON file controls the experiment, and changing that file is how a new configuration variant is added.

The Four-Step Experiment Pipeline

The pipeline is organized into four numbered Python scripts, and a Makefile provides short commands for each:

bash
make index
make qnagen
make query
make eval

These correspond to `01_index.py` (load documents, chunk, embed, upload to Azure AI Search), `02_qa_generation.py` (generate question-answer pairs), `03_querying.py` (query each index with the LLM), and `04_evaluation.py` (compute metrics against ground truth). Each step accepts `--config_path` pointing to a JSON configuration file, with `config.sample.json` provided as a starting template. The data directory defaults to `./data` and can be overridden with `--data_dir`.

Running `make all` executes all four steps in sequence. Running `make query_eval` skips indexing and QA generation and runs only the query and evaluation steps, which is useful when re-running experiments without rebuilding the search index. The environment is loaded from a .env file, which the Makefile sources automatically before running any target. The template for that file is `.env.template` in the repository root.

Configuration-Driven Hyperparameter Search

The README describes the tool as config-driven. The config file controls the search index parameters (chunking strategy, embedding model, document loader), the search types to test, the query set, and the evaluation metrics. The tool creates multiple search indexes based on the hyperparameter combinations available in the configuration, runs all configured search types against each index, and collects results for comparison. A JSON schema for the config file (config.schema.json) is included in the repository, which means editors that support JSON Schema can validate the configuration before running it.

Supported search types include pure text, pure vector, cross-vector, multi-vector, hybrid, and more. The sub-querying feature evaluates each user query at runtime: if the query is complex enough, it breaks it into smaller sub-queries, retrieves context for each, and combines the results before generation. Re-ranking uses an LLM to re-score the results from Azure AI Search before passing them to the generation step, which adds latency but can improve relevance for complex queries. The query generation step (02_qa_generation.py) creates diverse and customizable question sets automatically from the indexed document chunks, so a ground-truth dataset does not need to be created by hand before the experiment can run.

Document Loaders and the Custom Document Intelligence Loader

The tool supports multiple document loaders. The standard path uses LangChain loaders. The Azure Document Intelligence path calls the Azure Document Intelligence API to extract structured content from PDFs and Office documents.

For the `prebuilt-layout` API model, the tool uses a custom loader rather than LangChain's built-in implementation. The README describes several behaviors of this custom loader: it formats tables with column headers into key-value pairs to improve readability for the LLM, excludes page numbers and footers, removes recurring patterns using regex, and chunks recursively by paragraph and line to avoid splitting table rows mid-row. When the `prebuilt-layout` model fails, the loader falls back to its simpler variant. All other API models use LangChain's Document Intelligence implementation.

Evaluation Metrics and Report Generation

The evaluation step computes metrics across two categories. End-to-end metrics compare the generated answers against ground-truth answers using distance-based measures, cosine similarity, and semantic similarity. Component metrics assess retrieval and generation performance using LLMs as judges, including context recall, answer relevance, and MAP@k for retrieval quality.

The tool integrates with Azure Machine Learning and MLflow for experiment tracking. Report generation is automated, producing visualizations that compare configurations across all evaluated metrics. The integration with MLflow means experiment runs are logged and comparable across sessions, which is important when running many configuration variants over time.

Limitations: Azure-Only and Dataset Size Constraints

The tool is built specifically for Azure AI Search. The requirements.txt includes azure-search-documents, azure-ai-ml, azure-ai-textanalytics, and azure.ai.documentintelligence as hard dependencies. There is no abstraction layer that would let the experiment pipeline target a different vector database or search service. Organizations that store documents in a non-Azure vector store, or that want to compare Azure AI Search against Elasticsearch or Pinecone, will find this tool does not support that comparison.

For large datasets, the sampling feature (added March 2024) clusters content and samples a specified percentage from each cluster for a representative subset. The README notes that results from the sample should be within approximately 10% of results on the full dataset, and recommends running the full dataset once a promising configuration is identified. Without sampling, full experiments on large document collections can take a significant amount of time and incur Azure API costs proportional to the number of configurations tested and the number of documents processed.

The multi-lingual support deserves a note. The tool supports language analyzers for individual languages and specialized language-agnostic analyzers for user-defined patterns on search indexes. This is a feature of Azure AI Search's index configuration that the tool can exercise through its config file, not a translation layer. Engineers working with non-English document collections should evaluate whether the language analyzer selection in the config is appropriate for their target language.

The tool requires Python 3.11 or higher (from setup.py). The current version in setup.py is 0.9. Setup is available through a dev container (which handles all dependencies automatically in a Docker environment) or a local install. The requirements.txt is comprehensive: it includes langchain, openai, mlflow, sentence-transformers, and dozens of supporting packages.

Maintenance and Getting Started

The last push to the repository was on 2026-06-19. The repository is not archived. The license field is NOASSERTION, though a LICENSE file is present; the repository is from Microsoft so the actual license should be verified in that file before redistribution.

The dev container path requires Docker Desktop and VS Code with the Remote-Containers extension. For Windows users, WSL 2 with Ubuntu is required for the dev container path. The local install path uses a .env file from the .env.template and runs setup.py. Azure AI Search at Basic tier or higher is required for semantic search features, and Azure OpenAI or the OpenAI API must be configured before the querying step will function.

Editorial conclusion

RAG Experiment Accelerator is appropriate for teams building RAG pipelines on Azure who need a structured way to compare retrieval configurations rather than tuning by intuition. It requires Azure AI Search at Basic tier or higher (for semantic search), Azure OpenAI or access to the OpenAI API, and optionally Azure Machine Learning for MLflow tracking. Teams not on Azure, or those who need to compare non-Azure vector databases, will find it a poor fit since the tool is tightly integrated with Azure AI Search. Before running full experiments, use the sampling feature with a small percentage of the dataset to validate that the configuration is correct and results are meaningful.

Frequently asked questions

What does RAG Experiment Accelerator primarily aim to solve?

It solves the problem of comparing RAG retrieval configurations systematically. Instead of manually testing chunking strategies, search types, and embedding models one at a time, it creates multiple Azure AI Search indexes from a single config file, generates query-answer pairs, runs queries, and computes evaluation metrics across all configurations in one pipeline.

How to do RAG evaluation with RAG Experiment Accelerator?

Run `make eval` (or `python3 04_evaluation.py`) after the query step completes. The evaluation step compares generated answers against ground-truth answers using distance-based, cosine, and semantic similarity metrics, and uses LLMs as judges for component metrics including context recall and answer relevance.

How to develop a RAG system using this tool?

Start with config.sample.json as a template, point it at your document data, and run the four steps in order: `make index`, `make qnagen`, `make query`, `make eval`. The tool creates and tests multiple search indexes based on the hyperparameter combinations in the config file and produces comparative reports through MLflow.

Official sources

  1. Issues
  2. microsoft/rag-experiment-accelerator on GitHub
  3. Project website
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/microsoft-rag-experiment-accelerator.svg)](https://hysenlabs.com/projects/microsoft-rag-experiment-accelerator)