# RagBuilder: Bayesian hyperparameter search over a RAG pipeline, now unmaintained

> A KruxAI toolkit that tunes chunking, retrieval and generation settings against an evaluation dataset, shipped under Apache-2.0 with a README that declares it is no longer actively maintained.

**KruxAI/ragbuilder** — A toolkit to create optimal Production-readyRetrieval Augmented Generation(RAG) setup for your data

- Repository: https://github.com/KruxAI/ragbuilder
- Website: https://ragbuilder.pages.dev
- Stars: 1,539 · Forks: 127
- Language: Python
- License: Apache-2.0
- Published: 2026-10-08 · Updated: 2026-10-08 · Language: en
- Canonical page: https://hysenlabs.com/projects/kruxai-ragbuilder

## A README that starts with its own obituary

The first line of the README is a maintenance status notice: RagBuilder is no longer actively maintained, this repository is kept for reference, and ongoing updates and support are not planned. Very few projects put that above their own install instructions, and the placement is the point. Anyone who lands on the page sees it before the logo, the badges, or the feature list.

The repository itself contradicts the notice in one measurable way. The last push was on 2026-10-02, which is recent, and the tree carries a `requirements.lock`, a hardened `Dockerfile` that runs as a non-root user with uid 10001, a `SECURITY.md`, and a compose file that requires explicit secrets. None of those artifacts are what a project that has stopped caring looks like. The release history tells a different story: v0.1.4 on 2024-12-31, then 0.0.22 in October 2024, then 0.0.21 in October 2024. Two releases on the same month, and nothing since the end of 2024.

So the accurate reading is that commits continue while releases stopped, and that the code in the repository has drifted well past the last published version. Both statements are true, and the second one is the one that should shape your expectations.

## The five minute path from a URL to a tuned pipeline

The quick start is short enough to quote in full, which is unusual for a project with this much configuration surface. The recommended install uses uv:

```bash
# Create a new venv
uv venv ragbuilder

# Activate the new venv
source ragbuilder/bin/activate

# Install
uv pip install ragbuilder
```

```python
from ragbuilder import RAGBuilder

# Initialize and optimize with defaults
builder = RAGBuilder.from_source_with_defaults(input_source='https://lilianweng.github.io/posts/2023-06-23-agent/')
results = builder.optimize()

# Run a query through the complete pipeline
response = results.invoke("What is HNSW?")

# View optimization summary
print(results.summary())
```

The shape of the API is the interesting part. `from_source_with_defaults` takes any input source the document loaders understand, `optimize()` runs the search, and the result object is itself a runnable pipeline with `invoke`. Passing a URL rather than a file in the flagship example is a deliberate choice: it lets someone see the whole loop run without assembling a corpus first.

You can set the models used throughout the search, which matters because the optimizer needs a generator to evaluate candidates:

```python
from langchain_openai import AzureChatOpenAI, AzureOpenAIEmbeddings

# Initialize with custom defaults
builder = RAGBuilder.from_source_with_defaults(
    input_source='data.pdf',
    default_llm=AzureChatOpenAI(model="gpt-4o", temperature=0.0),
    default_embeddings=AzureOpenAIEmbeddings(model="text-embedding-3-large"),
    n_trials=20  # Set number of optimization trials
)
```

Note that `n_trials` is the cost knob. The feature list describes Bayesian optimization, and the dependency list confirms the implementation with scikit-optimize and optuna, so a larger trial count means a longer and more expensive search rather than a smarter one.

## Optimizing one stage at a time instead of all at once

The v0.1.4 release added an SDK for module-wise optimization, which is the most useful thing in the project's history and the reason to read its documentation even now. Instead of calling `optimize()` and waiting for a full sweep, you can target a stage.

The advanced configuration example in the README shows the shape. Ingestion is configured as a `DataIngestOptionsConfig` naming the input source, the loaders, the chunking strategies, a chunk size range, and the embedding models:

```python
data_ingest_config = DataIngestOptionsConfig(
    input_source="data.pdf",
    document_loaders=[
        {"type": "pymupdf"},
        {"type": "unstructured"}
    ],
    chunking_strategies=[{
        "type": "RecursiveCharacterTextSplitter",
        "chunker_kwargs": {"separators": ["\n\n", "\n", " ", ""]}
    }],
    chunk_size={"min": 500, "max": 2000, "stepsize": 500},
    embedding_models=[{
        "type": "openai",
        "model_kwargs": {"model": "text-embedding-3-large"}
    }]
)
```

That `chunk_size` range is the clearest statement of intent in the whole repository. RagBuilder does not ask you to pick a chunk size, it asks you to bound the search from 500 to 2000 in steps of 500 and let the evaluation dataset decide. Retrieval works the same way, with a list of retrievers carrying their own weights, a reranker, and a `top_k` search space:

```python
retrieval_config = RetrievalOptionsConfig(
    retrievers=[
        {
            "type": "vector_similarity",
            "retriever_k": [20],
            "weight": 0.5
        },
        {
            "type": "bm25",
            "retriever_k": [20],
            "weight": 0.5
        }
    ],
    rerankers=[{
        "type": "BAAI/bge-reranker-base"
    }],
    top_k=[3, 5]
)
```

Generation is configured as candidate LLMs plus an optimization block and an evaluation config, with `evaluation_config={"type": "ragas"}` selecting the metric. Having several models in one search space is a slightly unusual design, since it optimizes for the pipeline rather than fixing a model, but it does let the search discover that a cheaper model scores as well.

The feature list also names pre-defined RAG templates as a headline capability, citing Graph retriever and Contextual chunker as examples that demonstrated strong performance. Those two names do not appear in the component reference, which lists the retrievers as `vector_similarity`, `vector_mmr`, `bm25`, `multi_query`, `parent_doc_full` and `parent_doc_large`, and the chunkers as `RecursiveCharacterTextSplitter`, `CharacterTextSplitter`, `MarkdownHeaderTextSplitter`, `HTMLHeaderTextSplitter`, `SemanticChunker` and `TokenTextSplitter`. The template names are marketing labels whose relationship to the API names is not documented, so plan to look them up in the docs rather than assume a mapping.

## A documented example that cannot run

The advanced configuration example ends with two lines that do not agree with each other:

```python
results = builder.optimization_results
response = adv_results.invoke("What is HNSW?")
```

`results` is assigned and then never used. `adv_results` is called but never assigned anywhere in the example, and it is not a name introduced by any of the configuration classes above it. Copy this block and it raises a `NameError` on the last line.

The earlier quick start uses the correct form, calling `invoke` on the object returned by `optimize()`, which makes it clear that `adv_results` is a typo or a leftover from an earlier draft rather than a real API name. It is a small thing, but it sits in the block a reader is most likely to paste, at the very end, after the long configuration section. A copy-and-paste failure at the last line of the hardest example is exactly the kind of detail that costs an afternoon.

This is the clearest single signal about documentation quality in the project. The component reference is thorough and the configuration classes are well named, but the worked example that ties them together has an error nobody caught, which is consistent with a project that has had no maintainer reviewing issues for the better part of a year.

## Dependency drift between the last release and the default branch

The last release was v0.1.4 on 2024-12-31. The packaging metadata on the default branch now requires `langchain>=1.4.3,<2`, `langchain-core>=1.6.6,<2`, `langchain-huggingface>=1.2.2,<2`, `langchain-openai>=1.6.7,<2`, `fastapi>=0.142.2,<1`, `ragas>=0.4.3,<0.5`, `chromadb>=1.5.9` and `langchain-chroma>=1.1.0`. Those floors are far ahead of anything resolvable at the time of the v0.1.4 tag.

That is not automatically a problem, since a maintainer bumping floors on an unmaintained branch is a normal side effect of dependency bot runs. It does change what installing from the default branch means: you are not getting v0.1.4 plus fixes, you are getting a source tree whose dependency contract has been rewritten by automation against a moving LangChain release train, with no release notes explaining what code changed alongside those bumps.

One line deserves particular attention. `langchain-community==0.4.1` is pinned to an exact version while every other LangChain package in the same list floats within a major version. An exact pin among floating neighbours is a deliberate constraint, and combined with the `>=1.6.6` floor on langchain-core it describes a specific compatibility window that nobody is going to widen for you now. The README offers no guidance on this, so if the resolver fights you, that line is the first thing to inspect.

The repository does carry a `requirements.lock`, which is the most reliable starting point for a reproducible environment, and the Dockerfile installs from exactly that file:

```
RUN pip install --no-cache-dir -r requirements.lock .
```

Version numbers in the Docker image are a separate curiosity. `pyproject.toml` declares `dynamic = ["version"]`, and the Dockerfile passes `SETUPTOOLS_SCM_PRETEND_VERSION=0.0.0` at build time, so an image built from source reports version 0.0.0 regardless of the git state it was built from.

## What running it as a service actually requires

The deployment path is not mentioned in the README at all, which makes `docker-compose.yml` the better documentation. It brings up two services. The first is Neo4j, built from a local `./neo4j` directory, bound to loopback on 7474 and 7687, with the APOC export and import file extensions explicitly disabled and the auth password required from the environment:

```yaml
NEO4J_AUTH: "neo4j/${NEO4J_PASSWORD:?Set a strong NEO4J_PASSWORD}"
NEO4J_apoc_export_file_enabled: "false"
NEO4J_apoc_import_file_enabled: "false"
```

The second service is the API itself, bound to `127.0.0.1:55003:8005`, with a token requirement that names its own length:

```yaml
RAGBUILDER_API_TOKEN: "${RAGBUILDER_API_TOKEN:?Set a random token of at least 32 characters}"
```

The compose syntax uses shell parameter expansion with the `:?` operator, so the stack refuses to start rather than booting with an empty password or token. That is careful work, and it is the kind of detail that distinguishes a project built for evaluation by a security reviewer from one that is not. Ports are loopback only, volumes persist `./data` and `./output`, and the container runs as uid 10001.

The knowledge graph half of the stack is optional at the package level, exposed as the `graph` extra pulling in `neo4j>=5.23.0` and `langchain-community[neo4j]`, with separate extras for vector stores (Elasticsearch, FAISS, Pinecone, Milvus, Qdrant, Weaviate) and document processors (pymupdf, python-docx, pikepdf, pandoc, pypdf, markdown, beautifulsoup4, unstructured). The base install already includes chromadb through langchain-chroma, which is the default store.

If you want to run this, read the compose file rather than the README, and expect to generate your own secrets before anything starts.

## Conclusion

RagBuilder is worth studying because it treats a RAG pipeline as a set of tunable knobs rather than a fixed recipe, and because the module-wise optimization API added in v0.1.4 lets you optimize one stage at a time instead of running the whole search end to end. Whether you should install it is a different question. The README opens by saying the project is no longer actively maintained and is kept for reference, and the dependency picture supports taking that seriously: the last release was v0.1.4 in December 2024, while the packaging metadata on the default branch now requires langchain 1.4.3, fastapi 0.142.2 and ragas 0.4.3, versions that could not have been resolved at release time. Start with the pinned `requirements.lock` if you need a reproducible environment, and read the config reference on docs.ragbuilder.io before assuming any component name matches.

## FAQ

### What does RagBuilder actually do?

It runs Bayesian hyperparameter optimization over a retrieval augmented generation pipeline. You give it a data source and, optionally, an evaluation dataset, and it searches chunking strategy, chunk size, retriever choice, reranker, top_k and even the generator model, evaluating candidate configurations against the dataset with ragas. The result is both an optimized configuration and a runnable pipeline object.

### Is RagBuilder still maintained?

The README states that RagBuilder is no longer actively maintained and that ongoing updates and support are not planned. The last release was v0.1.4 on 2024-12-31, though the repository has been pushed as recently as 2026-10-02. Plan around the frozen release and treat the default branch as unreviewed source.

### How do I optimize just one part of the pipeline, such as chunking?

Release v0.1.4 added module-wise optimization. Build a `DataIngestOptionsConfig` describing loaders, chunking strategies, a chunk size range and embedding models, pass it to `RAGBuilder`, then call `builder.optimize_data_ingest()`. There are matching `optimize_retrieval()` and `optimize_generation()` methods for the other two stages.

### Do I need a Neo4j instance to run RagBuilder?

Not for the base package. Neo4j is an optional extra under the name `graph`, pulling in the neo4j driver. It becomes necessary if you want to use the knowledge graph templates, and the supplied `docker-compose.yml` starts it alongside the API with both ports bound to loopback and the password required from the environment.

### What evaluation framework does RagBuilder score pipelines with?

Ragas, selected in the generation configuration through `evaluation_config={"type": "ragas"}`, and declared in the dependencies as `ragas>=0.4.3,<0.5`. You can also supply your own test dataset instead of letting the tool generate a synthetic one, which is the setting to reach for when your own queries are the ones you care about.

## Sources

- [KruxAI/ragbuilder on GitHub](https://github.com/KruxAI/ragbuilder)
- [License: Apache-2.0](https://github.com/KruxAI/ragbuilder/blob/main/LICENSE)
- [Project website](https://ragbuilder.pages.dev)
- [README](https://github.com/KruxAI/ragbuilder/blob/main/README.md)
- [Releases](https://github.com/KruxAI/ragbuilder/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/kruxai-ragbuilder
