# RAG Web UI: a self-hosted knowledge base Q&A stack in TypeScript and Python

> RAG Web UI pairs a Next.js frontend with a Python backend to ingest PDFs, DOCX, Markdown and text into a vector store, then answer questions with citations. It suits teams that want their own retrieval pipeline rather than a general chat interface.

**rag-web-ui/rag-web-ui** — RAG Web UI is an intelligent dialogue system based on RAG (Retrieval-Augmented Generation) technology.

- Repository: https://github.com/rag-web-ui/rag-web-ui
- Stars: 3,295 · Forks: 372
- Language: TypeScript
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/rag-web-ui-rag-web-ui

## What RAG Web UI actually replaces

Most teams that want retrieval-augmented generation end up writing the same three pieces: a document ingestion job, a vector store client, and a chat endpoint that stuffs retrieved chunks into a prompt. RAG Web UI ships those pieces as a product. The README describes it as an "intelligent dialogue system based on RAG (Retrieval-Augmented Generation) technology" that builds Q&A over your own knowledge base.

The intended user is a small team or a single engineer who wants a working knowledge base UI without building one. The repository separates frontend and backend, so the chat interface, document dashboard and API key management screens are already there. It also exposes OpenAPI interfaces, which matters if the knowledge base has to be called from something other than the bundled UI.

It is not a general chat client. The unit of work is a knowledge base, and the flow assumes you upload documents before you ask questions. If your use case is conversation with a model about no particular corpus, this project adds infrastructure you will not use.

## The ingestion and query path, as the flowchart shows

The README includes a mermaid flowchart that documents both halves of the system. Ingestion starts when a caller uploads a PDF, Markdown, TXT or DOCX file. The file goes to NFS storage, and the API immediately returns a job ID rather than blocking. A separate asynchronous process extracts and cleans text, splits it with segmentation and overlap, sends the chunks to an embedding service, and writes the vectors to the vector database. The client polls a job status endpoint that reports processing, completed or failed.

Query is a different path. User history feeds into the query, the query is embedded, vectors are retrieved, results pass through a cross-encoder re-ranking step, context is assembled, and the LLM generates the final response. The flowchart also shows the original query feeding into the re-ranking stage, which is the usual pattern for cross-encoder scoring.

The OpenAPI retrieval flow is separate again: an external caller can request context from the retrieval service, which reads the vector store directly and returns context without invoking the LLM. That is the piece worth noting. It means RAG Web UI can act as a retrieval backend for another application, not only as a chat product.

Two design choices stand out. Document processing is asynchronous with polling rather than a webhook, which is simpler but means the client owns retry logic. And NFS appears in the diagram as the file storage layer, while docker-compose.yml also defines a MinIO service, so the storage story in the diagram and the storage story in the deployment files are not described in the same terms.

## Installing RAG Web UI with Docker Compose

The repository ships a docker-compose.yml and a .env.example. The compose file defines backend, frontend, db (mysql:8.0), chromadb and a commented-out Qdrant service. Copy the example environment file first, because the backend reads it through env_file.

```bash
cp .env.example .env
```

Then edit .env. The two keys that decide everything are CHAT_PROVIDER and EMBEDDINGS_PROVIDER. The example defaults both to openai, which requires OPENAI_API_KEY, OPENAI_API_BASE, OPENAI_MODEL and OPENAI_EMBEDDINGS_MODEL.

```bash
CHAT_PROVIDER=openai
EMBEDDINGS_PROVIDER=openai
OPENAI_API_KEY=your-openai-api-key-here
OPENAI_API_BASE=https://api.openai.com/v1
OPENAI_MODEL=gpt-4
OPENAI_EMBEDDINGS_MODEL=text-embedding-ada-002
```

If you would rather keep inference local, the same file documents Ollama settings. The comments are explicit about the address problem: inside docker-compose on macOS use host.docker.internal, on a Linux server use the host machine's IP, and for a compiled installation use localhost.

```bash
CHAT_PROVIDER=ollama
OLLAMA_API_BASE=http://host.docker.internal:11434
OLLAMA_MODEL=deepseek-r1:7b
EMBEDDINGS_PROVIDER=ollama
OLLAMA_EMBEDDINGS_MODEL=nomic-embed-text
```

With .env populated, bring the stack up.

```bash
docker compose up -d
```

MySQL publishes on port 3306, ChromaDB publishes on 8001 and maps to 8000 inside the container. The db service has a healthcheck using mysqladmin ping with a 10 second start period, and the backend declares depends_on with condition: service_healthy for the database, so the backend waits for MySQL before starting. The backend also sets extra_hosts for host.docker.internal:host-gateway, which is what makes the Ollama address work on Linux. The first real use is to open the frontend, create a knowledge base, upload a document, and watch the job status move from processing to completed before asking a question against it.

## Where the deployment gets awkward

The compose file publishes MySQL on 3306 and ChromaDB on 8001 to the host by default. On a shared machine that is a collision waiting to happen, and it also exposes the database beyond the compose network. The backend and db are on an app_network, so the ports section is not strictly required for internal communication.

The restart configuration is inconsistent. The backend has restart: on-failure plus a deploy block with restart_policy, delay 5s and max_attempts 3. The frontend, db and chromadb have no restart policy at all. A host reboot leaves the stack partially up.

The database credentials are hardcoded in the compose file: MYSQL_ROOT_PASSWORD=root, MYSQL_DATABASE=ragwebui, MYSQL_USER=ragwebui, MYSQL_PASSWORD=ragwebui, with TZ fixed to Asia/Shanghai. These are development defaults, and nothing in the compose file suggests they are meant to be overridden through environment substitution. Anyone deploying this beyond a laptop should change them in the file itself.

The vector store switch is documented but manual. The Qdrant service in docker-compose.yml is commented out, with an instruction to remove the comment and start it. Nothing in the README explains how the Factory pattern picks between ChromaDB and Qdrant at runtime, so the switching mechanism has to be read from the backend source. That is a gap for anyone who wants Qdrant from day one.

## What RAG Web UI is not

The clearest alternative is Open WebUI, which appears in the related search terms people use around this project. The difference is architectural. Open WebUI is a chat interface first, and its RAG capability is configured against a document collection inside that interface. RAG Web UI is a retrieval pipeline first: ingestion is an asynchronous job with a status endpoint, retrieval is exposed over OpenAPI for external callers, and the chat UI is one consumer of that pipeline.

That distinction decides the choice. If you want a chat front end for a model and occasionally want it to read a document, Open WebUI is the smaller commitment. If you want a service that other applications call to retrieve context, and you want the ingestion lifecycle visible as jobs, RAG Web UI is built around that. The cost is the dependency set: MySQL, ChromaDB or Qdrant, MinIO, plus the backend and frontend containers.

Within its own category, the honest limitation is that the README does not document rollback, re-indexing after an embedding model change, or what happens to existing vectors when EMBEDDINGS_PROVIDER changes. Since the embedding model is set per provider in .env, switching from text-embedding-ada-002 to a HuggingFace model changes the vector space. The README is silent on whether the system detects that and re-embeds, or whether you must delete and re-upload. Treat that as an open question to test on a throwaway knowledge base before migrating real documents.

## Licence, maintenance and upgrade cost

The project is Apache-2.0, which permits commercial use and modification, and includes a patent grant. The LICENSE file is at the repository root. That is a permissive licence, but it governs the RAG Web UI code only. The compose file pulls mysql:8.0, chromadb/chroma:latest and, if enabled, qdrant/qdrant:latest. Those images carry their own licences, and the ChromaDB image is pinned to latest, which means an upgrade can change the vector store underneath you without a version change in this repository. Pin it to a digest or a specific tag if reproducibility matters.

Maintenance activity is visible in the release history. v0.8.0 added HuggingFace embeddings support and was pushed on 2026-04-06, the same day as the last push to main. Before that, 0.7.5 replaced passlib with a direct bcrypt implementation on 2025-11-15, and v0.7.3 updated Docker configurations for host.docker.internal on 2025-08-30. The gaps between releases are measured in months, and the most recent release is roughly five months before today. The repository is not archived, but the cadence suggests a project that moves in bursts rather than continuously.

Upgrade cost concentrates in two places. The .env keys are the contract between your deployment and the code, so a release that adds a provider or renames a key will surface as a startup failure rather than a silent one. And because the backend mounts ./backend into the container and runs from that path, the compose deployment is closer to a development setup than an immutable image. A production deployment would want to build the image without the bind mount.

## Conclusion

Adopt RAG Web UI if you need a self-hosted knowledge base with an OpenAPI surface and you are willing to run MySQL, ChromaDB and MinIO alongside it. Skip it if a single-binary chat UI with a built-in document store is enough, or if you cannot operate four containers. Before committing, verify that your chosen EMBEDDINGS_PROVIDER actually answers at the URL in .env, and confirm the docker-compose.yml Qdrant block matches the vector store you intend to run.

## FAQ

### What is RAG Web UI and who is it for?

It is an intelligent dialogue system based on RAG technology that builds Q&A over your own knowledge base, with a TypeScript frontend and a Python backend. It targets teams that want a self-hosted retrieval pipeline with an OpenAPI surface rather than a general chat client.

### Which LLM and embedding providers does RAG Web UI support?

The .env.example documents OpenAI, DeepSeek, MiniMax, Ollama and HuggingFace embeddings, selected through CHAT_PROVIDER and EMBEDDINGS_PROVIDER. Each provider requires its own API key, base URL and model keys in .env.

### What vector databases can RAG Web UI use?

The README lists ChromaDB and Qdrant, with switching handled through a Factory pattern. docker-compose.yml starts ChromaDB by default and includes a commented-out Qdrant service you uncomment to run instead.

### What document formats can be uploaded to RAG Web UI?

The README lists PDF, DOCX, Markdown and Text, with automatic chunking and vectorization and support for asynchronous processing and incremental updates.

## Sources

- [Issues](https://github.com/rag-web-ui/rag-web-ui/issues)
- [License: Apache-2.0](https://github.com/rag-web-ui/rag-web-ui/blob/main/LICENSE)
- [rag-web-ui/rag-web-ui on GitHub](https://github.com/rag-web-ui/rag-web-ui)
- [README](https://github.com/rag-web-ui/rag-web-ui/blob/main/README.md)
- [Releases](https://github.com/rag-web-ui/rag-web-ui/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/rag-web-ui-rag-web-ui
