Self-hosted service
SensAI-PT/RAGMeUp avatar
SensAI-PT/RAGMeUp

RAGMeUp: A Docker-First RAG Framework with BM25, Vector Search, and a React UI

Generic rag framework to apply the power of LLMs on any given dataset

679 stars99 forksPythonApache-2.0

At a glance

What is it?
RAGMeUp is an Apache-2.0 RAG framework from SensAI.PT that ships a complete four-service Docker stack: ParadeDB for hybrid BM25 and vector search, a Python Flask RAG backend, a Node.js auth and chat API, and a React client. It supports a hybrid mode that keeps the Python server on the host for GPU access.
Who is it for?
RAGMeUp is a good fit for teams that need to get a RAG application to production quickly with a full-stack Docker setup and do not want to wire together the database, auth layer, and frontend themselves. The use of ParadeDB for both BM25 and vector search in a single database is a concrete advantage over setups that require a separate keyword search service alongside a vector store.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 18 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 22, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What RAGMeUp Provides as a Complete Stack

Most RAG frameworks are libraries, not applications. They give you the building blocks but leave the deployment, authentication, document management, and UI to you. RAGMeUp ships all four as a working Docker Compose application that starts with three commands and exposes a React UI on port 80.

The architecture is four services communicating over an internal Docker network. ParadeDB handles the database layer with both pgvector for embedding search and pg_search for BM25 keyword search. The Python Flask server handles RAG logic. The Node.js Express server handles authentication, chat history, document uploads, and feedback. The React SPA serves the user interface through Nginx and also reverse-proxies `/api` requests.

The three persistent volumes, `paradedb_data`, `python_data`, and `uploads_data`, survive container restarts. The React client port is the only port exposed to the host in the default mode; all internal services communicate over the private Docker network.

Two Deployment Modes: Docker-Only and Hybrid

The full Docker mode runs everything inside containers, including the Python RAG server. The limitation documented in the README is that the Python server has no GPU or CUDA access inside Docker, so embeddings and inference run on CPU only.

The hybrid mode solves this by keeping Postgres, the Node.js API, and the React client in Docker while running the Python server directly on the host machine where it has full GPU access:

bash
cd server
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
python server.py

The Python server starts on port 5000. The Dockerized Node server connects to it via `PYTHON_SERVER_URL`, which defaults to `http://host.docker.internal:5000`. On Linux, `host.docker.internal` requires `--add-host=host.docker.internal:host-gateway` in the compose file, or setting `PYTHON_SERVER_URL=http://172.17.0.1:5000` in the env file.

The `server/.env` file must have `postgres_uri` pointing at the host-exposed Postgres port (default 6024) and `embedding_cpu=False` to enable GPU embeddings.

Installing and Starting RAGMeUp

The quickstart is three commands:

bash
git clone https://github.com/SensAI-PT/RAGMeUp.git
cd RAGMeUp
cp docker-compose.env.example docker-compose.env

After copying the example env file, two values must be set before the stack will start: `POSTGRES_PASSWORD` and `JWT_SECRET`. Then:

bash
docker compose --env-file docker-compose.env up --build -d

The React UI is accessible at `http://localhost` or at the port set in `HOST_PORT`. The Postgres container uses a health check that waits for the database to accept connections before dependent services start. Full documentation including architecture diagrams, API references, and setup guides is at ragmeup.sensai.pt. The repository also includes `ragmeup.drawio` and `ragmeup.drawio.svg` for the architecture diagram.

BM25 and Vector Search in ParadeDB

The database choice is notable. ParadeDB is a Postgres extension that adds both pgvector (dense vector similarity search) and pg_search (BM25 full-text search). Using a single database for both retrieval modes simplifies the architecture: there is no separate Elasticsearch or OpenSearch cluster to run alongside the vector database.

BM25 is important for retrieval accuracy in cases where the query uses exact vocabulary from the documents, such as proper nouns, product names, or technical terms. Dense vector search handles semantic similarity where the words differ but the meaning matches. Hybrid retrieval combines both; having both in the same database makes fusing the results simpler than querying two separate systems and re-ranking across their scores.

The RELATED SEARCHES for RAGMeUp include "Ragmeup bm25", which indicates users specifically notice this feature. The docker-compose.yml comment identifies the Postgres service as providing BM25 via `pg_search`.

Limitations Worth Knowing Before Adopting

The full Docker mode has a documented limitation: the Python server has no GPU or CUDA access. For embedding models and LLM inference, this means CPU-only execution. For large documents or high query volumes, CPU-only embedding will be significantly slower than GPU-accelerated alternatives. The hybrid mode restores GPU access but adds setup complexity, especially on Linux where `host.docker.internal` requires extra configuration.

RAGMeUp has no GitHub releases. Updates are tracked by commits on the main branch only. Teams that need a pinned, versioned release for reproducible production deployments will need to tag a commit themselves.

The README states the framework is modular: custom chunkers, vectorstores, and retrievers are supported. However, the README does not document the extension API or show examples of implementing a custom component. Teams that need to replace ParadeDB with a different vector database would need to consult the full documentation at ragmeup.sensai.pt or read the Python server source in `server/`.

Comparison with Verba by Weaviate

Verba is an open-source RAG application from Weaviate that also ships with a UI and wraps a complete retrieval pipeline. The key architectural difference is the underlying vector store: Verba is built on Weaviate, a dedicated vector database, while RAGMeUp uses ParadeDB, a Postgres extension. ParadeDB combines the operational simplicity of a familiar relational database with native BM25 and vector search, which means one less service to deploy and operate. Weaviate has a larger feature surface for purely vector-focused workloads. Teams already running Postgres in production may find the ParadeDB path easier to integrate than adopting a new database technology.

License and Maintenance Status

RAGMeUp is Apache-2.0 licensed, which permits use, modification, and distribution. The last push was on 2026-09-12, sixteen days before this review, indicating active maintenance. The project is used in production at SensAI.PT, which the README describes as an AI personal trainer application. That production use is cited as evidence of the framework's stability in real-world settings. The `docker-compose.env` file at the repository root is the required configuration entry point; a `docker-compose.env.example` file shows all available variables.

Editorial conclusion

RAGMeUp is a good fit for teams that need to get a RAG application to production quickly with a full-stack Docker setup and do not want to wire together the database, auth layer, and frontend themselves. The use of ParadeDB for both BM25 and vector search in a single database is a concrete advantage over setups that require a separate keyword search service alongside a vector store. Teams that need CPU-only Docker inference should verify that performance is acceptable before committing, since the full Docker mode disables GPU access for the Python server. The hybrid mode adds setup complexity but restores GPU access.

Frequently asked questions

Does RAGMeUp support GPU inference?

In the default Docker mode, the Python server runs CPU-only because Docker does not expose the host GPU to the container. The hybrid mode runs the Python server directly on the host, giving it full GPU and CUDA access while the other services remain in Docker.

What database does RAGMeUp use for retrieval?

RAGMeUp uses ParadeDB, a Postgres extension that adds both pgvector for dense vector search and pg_search for BM25 keyword search. This means both retrieval modes run inside a single Postgres database.

What services does the RAGMeUp Docker Compose stack include?

The stack includes ParadeDB (Postgres with pgvector and BM25), a Python Flask RAG server, a Node.js Express API for auth, chat, documents, and feedback, and a React SPA served by Nginx. Only the Nginx port is exposed to the host.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. README
  4. SensAI-PT/RAGMeUp on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/sensai-pt-ragmeup.svg)](https://hysenlabs.com/projects/sensai-pt-ragmeup)