# Ragbits: a modular Python stack for RAG, agents and prompt testing

> Ragbits is a set of Python packages from deepsense-ai that cover prompts, LLM access, vector stores, document ingestion, guardrails and a chat API. It is worth a look if you want those pieces as separate installable modules rather than one framework.

**deepsense-ai/ragbits** — Building blocks for rapid development of GenAI applications 

- Repository: https://github.com/deepsense-ai/ragbits
- Website: https://ragbits.deepsense.ai
- Stars: 1,668 · Forks: 145
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/deepsense-ai-ragbits

## What Ragbits solves for Python teams building GenAI features

Most GenAI prototypes start as a script that calls a model and prints a string. That works until you need the same prompt in a CLI, a test harness and a chat service, or until you want to swap the model provider without rewriting call sites. Ragbits targets that second stage. The README describes it as building blocks for rapid development of GenAI applications, and the repository layout backs that up: packages/ holds eight separate packages, each published on its own, plus a starter bundle named ragbits that pulls them together.

The intended audience is a Python developer who is comfortable with type hints and async code. The quickstart uses asyncio.run, Pydantic models and generic Prompt subclasses, so the library assumes that style rather than hiding it. If your team writes synchronous scripts with dictionary payloads, the examples will look foreign. The project is also not a hosted product: the homepage points to documentation, and the code lives in the repository, so you run it yourself.

One design decision stands out. Ragbits is modular by default. The pyproject.toml in the repository root lists the workspace members, and the README tells you that you can instead install individual components. That matters when a service only needs prompt handling and vector search, and you would rather not pull in Ray, Unstructured and a chat server as transitive dependencies.

## How the packages fit together: core, document search, agents, chat

The stack is layered. ragbits-core holds prompts, LLM clients and vector stores. ragbits-document-search builds ingestion and retrieval on top of it. ragbits-agents adds agent abstractions, ragbits-chat provides the full-stack conversational layer, ragbits-evaluate covers evaluation, ragbits-guardrails covers response safety, and ragbits-cli exposes the ragbits shell command.

The prompt mechanism is the clearest example of the data flow. You declare an input model with Pydantic, subclass Prompt with that input type and an output type, and put template text in system_prompt and user_prompt class attributes. The templates use Jinja-style placeholders such as {{ question }}. You then construct the prompt with an input instance and hand it to an LLM object's generate method. Because the output type is a generic parameter, the README describes this as type-safe LLM calls using Python generics to enforce strict type safety in model interactions. In practice that means the response type is part of the prompt class rather than something you cast afterwards.

Model access goes through LiteLLM, which the README says covers 100+ LLMs, with a separate path for local models. Embeddings follow the same pattern through LiteLLMEmbedder, and vector stores are pluggable: Qdrant and PgVector are named in the README, and the workspace dependency list also includes chroma, weaviate, fastembed and others as extras. Ingestion is the heaviest part. The README lists over 20 formats and two parser backends, Docling and Unstructured, with a custom parser option. For large corpora it points to Ray-based parallel processing. That is a real architectural commitment: distributed ingestion means a Ray cluster or at least the Ray runtime, which is not something you add casually to a small service.

## Installing Ragbits and running a first prompt

The README gives a single install command for the stable release, which installs the starter bundle containing all the main packages. Nightly builds come from the main branch with a pre flag and follow a version format like X.Y.Z.devYYYYMMDDHHMM; the README warns they may be less stable than official releases.

```bash
pip install ragbits
```

If you only want part of the stack, the README says you can install individual components instead of the bundle, and the repository lists the package directories under packages/.

The first real use is a prompt with a typed input. This example is adapted from the README quickstart, which defines a Pydantic input model, a Prompt subclass with a system and user template, and a LiteLLM instance.

```python
import asyncio
from pydantic import BaseModel
from ragbits.core.llms import LiteLLM
from ragbits.core.prompt import Prompt

class QuestionAnswerPromptInput(BaseModel):
    question: str

class QuestionAnswerPrompt(Prompt[QuestionAnswerPromptInput, str]):
    system_prompt = """
    You are a question answering agent. Answer the question to the best of your ability.
    """
    user_prompt = """
    Question: {{ question }}
    """
```

Running the prompt means constructing it with an input instance and awaiting generate on the LLM. The README's full example uses model_name="gpt-4.1-nano" and prints the response, so with a reachable provider and credentials in the environment you should see a plain string answer in the terminal.

The README also mentions a CLI path for testing prompts from your terminal, under the ragbits-vector-store command group documented in the CLI reference. It does not spell out the exact subcommand for prompt testing in the text we have, so check the CLI documentation before scripting against it.

## Where Ragbits gets in the way

The modular packaging is a trade-off, not a free win. The starter bundle's dependency list in the repository root is long: it includes extras for chroma, fastembed, local models, otel, logfire, qdrant, pgvector, weaviate, azure, gcs, hf, s3 and google_drive, plus unstructured and ray for document search, relari for evaluation, openai for guardrails, and a2a, cli and mcp for agents. Installing ragbits therefore pulls a wide surface. The README's answer is to install individual components, which means you now own a dependency graph and must keep those packages in step yourself.

Version drift is the second problem. Releases arrive frequently, with v1.6.0, v1.6.1 and v1.6.2 all published in March 2026, and the last push to main was on 2026-05-18. A fast cadence on a young API means upgrade work. The README does not document a rollback procedure, and the repository's release_checklist.md is a maintainer document rather than a user upgrade guide. If you need long-term API stability guarantees in writing, this is not the project for you.

There is also a scope limit worth naming. Ragbits gives you the pieces to build retrieval and agents, but it is not a hosted retrieval service and it does not ship a managed vector database. You supply Qdrant, PgVector or another backend, you supply the model credentials, and you operate the chat API yourself. Teams that want an end-to-end managed product should look elsewhere.

## Ragbits compared with LangChain and LlamaIndex

The obvious comparison is with LangChain and LlamaIndex, which also occupy the Python GenAI toolkit space. The difference in approach is packaging and typing. LangChain and LlamaIndex grew as broad frameworks with large integration surfaces and, in LangChain's case, a separate expression language for composing chains. Ragbits splits its functionality into eight separately installable packages and leans on Python's own type system instead: a Prompt subclass declares its input and output types as generic parameters, and Pydantic validates the input model.

That choice has consequences. You get less built-in orchestration and fewer prebuilt chain abstractions, so composition is ordinary Python rather than a framework DSL. In exchange, the surface you learn is smaller and the types are visible in your editor. The document search side is closer to the two alternatives in scope: Ragbits names Docling and Unstructured as parser backends, both of which are also common choices in other stacks, and it adds Ray-based distributed ingestion, which the README documents as a separate how-to page.

For agents, Ragbits points at the A2A protocol for interoperability and MCP for live tool access. That is a bet on open protocols rather than a proprietary agent runtime, and it means your agent code depends on those protocol implementations being available and current.

## Licence, maintenance and upgrade cost

Ragbits is MIT licensed, and the repository carries a LICENSE file plus check_licenses.sh and two whitelist files, .license-whitelist.txt and .libraries-whitelist.txt. Those scripts suggest the maintainers check dependency licences as part of their process. MIT is permissive, so the usual obligation is preserving the copyright notice and licence text; the repository does not offer legal advice and neither does this article. If your organisation has a dependency review step, the whitelist files are a useful signal about what the project considers acceptable, but they are the project's own policy, not a statement about your obligations.

On maintenance: the repository is not archived, and the most recent push to main was on 2026-05-18. The latest release listed is v1.6.2 from 2026-03-31. Those dates tell you the project was receiving commits in the months before this review, but they do not tell you the size of the maintainer team or how quickly issues are answered. Nothing in the README documents a support window or a deprecation policy.

Upgrade cost is the practical question. With three releases in March 2026 alone, pinning versions in your own lockfile is the sensible default. The repository uses uv with a uv.lock, and the workspace pyproject.toml pins several dev tools with compatible-release or range specifiers, which is a pattern you can copy for your own project. The README does not describe a migration guide between minor versions, so budget time to read release notes when you move.

## Conclusion

Ragbits fits teams already writing Python who want prompt definitions, LLM access and vector stores as installable packages instead of one opinionated framework. It is a poor fit if you need a managed service, or if you cannot accept that the README does not document rollback or upgrade paths between releases. Before adopting it, install the starter bundle, run the prompt example against a model you can reach, and confirm which optional extras your deployment actually needs.

## FAQ

### What is Ragbits used for?

It is a set of Python building blocks for GenAI applications, covering prompts, LLM access through LiteLLM, vector stores, document ingestion and search, agents, guardrails, evaluation and a chat API. The README describes it as building blocks for rapid development of GenAI applications.

### How do I install Ragbits?

The README gives pip install ragbits for the latest stable release, which installs the starter bundle of packages. Nightly builds from the main branch are available with pip install ragbits --pre and follow a version format like X.Y.Z.devYYYYMMDDHHMM.

### Which vector stores does Ragbits support?

The README names Qdrant and PgVector with built-in support and says you can bring your own vector store. The workspace dependency list also includes extras for chroma, weaviate and fastembed.

### Does Ragbits include a chat interface?

Yes. The ragbits-chat package is described in the README as full-stack infrastructure for building conversational AI applications, with an API, persistence and user feedback, and the repository includes a chat example directory.

## Sources

- [deepsense-ai/ragbits on GitHub](https://github.com/deepsense-ai/ragbits)
- [License: MIT](https://github.com/deepsense-ai/ragbits/blob/main/LICENSE)
- [Project website](https://ragbits.deepsense.ai)
- [README](https://github.com/deepsense-ai/ragbits/blob/main/README.md)
- [Releases](https://github.com/deepsense-ai/ragbits/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/deepsense-ai-ragbits
