Self-hosted service
Cinnamon/kotaemon avatar
Cinnamon/kotaemon

Kotaemon: A Configurable RAG UI for Document Question Answering

An open-source RAG-based tool for chatting with your documents.

25,769 stars2,156 forksPythonApache-2.0

At a glance

What is it?
Kotaemon is an Apache-licensed Python application that wraps a hybrid retrieval-augmented generation pipeline in a Gradio web interface. It targets two audiences: end users who need a working document QA interface and developers who want to build or extend RAG pipelines using the kotaemon library.
Who is it for?
Kotaemon suits teams that want a working RAG interface without assembling retrieval, re-ranking, and a PDF citation viewer from scratch, and who use either cloud LLM APIs or a local Ollama setup. It is less suited to deployments requiring a fully custom front end that goes beyond Gradio's component model.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 77 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 17, 2026, and from our analysis. They are not legal advice.

DEEP OPEN-SOURCE ANALYSIS

A Document QA Tool Built for Two Audiences

Kotaemon occupies a specific position in the RAG tool space: it is simultaneously a finished web application and a Python library that developers can import into their own projects. End users interact with a Gradio-based interface to upload documents, ask questions, and read answers with citations. Developers pull in the kotaemon library to build custom retrieval pipelines and surface them through the same interface without writing a front end from scratch.

The README describes this layering as three concentric audiences: end users of deployed apps, developers who build with the library, and contributors to the codebase. That design means the same repository serves a self-hosted document QA deployment and a foundation for custom pipeline work. End users access kotaemon as a zip archive from the releases page and follow the user guide. Developers work with the kotaemon and ktem packages defined in the uv workspace configuration, importing kotaemon components directly into their own Python projects.

The multi-user feature supports separate logins, private document collections, and shared chats. Teams can organize files into private or public collections and share specific conversations with other users on the same instance.

How the Hybrid Retrieval Pipeline Processes Documents

Kotaemon applies a hybrid retrieval strategy by default. It combines full-text search with vector retrieval and then applies a re-ranking pass over the combined candidate set. The README calls this combination the sane default for balancing recall and relevance. Full-text search finds documents containing exact query terms. Vector retrieval finds semantically related content that may not share vocabulary with the query. Re-ranking scores the merged set so the most relevant passages appear first.

Citations are tied directly to the in-browser PDF viewer. Each answer includes references with relevance scores, and the viewer highlights the exact passages the pipeline retrieved. When the pipeline returns results below a relevance threshold, the interface issues a warning. That signal lets users judge whether an answer is grounded in the document content or is the result of low-quality retrieval.

For questions that require reasoning across multiple document sections, the system supports question decomposition and agent-based reasoning. The README lists ReAct and ReWOO as available agent types, both selectable from the settings UI without changing code. The configurable settings panel also lets users adjust prompts, retrieval parameters, and the generation process directly from the interface.

Installing and Running Kotaemon via Docker

The recommended installation path is Docker. Kotaemon publishes three image variants through the GitHub Container Registry. The main-lite image covers PDF, HTML, MHTML, and XLSX files and suits the majority of use cases. The main-full image adds Tesseract OCR, LibreOffice, and Unstructured library support for Office formats such as .doc and .docx, at the cost of a larger image. The main-ollama image bundles Ollama for teams that want fully local inference with no data leaving their network.

The command for the full image is:

shell
docker run \
-e GRADIO_SERVER_NAME=0.0.0.0 \
-e GRADIO_SERVER_PORT=7860 \
-v ./ktem_app_data:/app/ktem_app_data \
-p 7860:7860 -it --rm \
ghcr.io/cinnamon/kotaemon:main-full

This starts the Gradio interface at http://localhost:7860/. The -v flag mounts a local directory so uploaded documents and settings persist across container restarts. To run on Apple Silicon, add --platform linux/arm64 before the image name.

For a non-Docker setup, start by cloning the repository:

shell
git clone https://github.com/Cinnamon/kotaemon
cd kotaemon

The README directs developers to uv as the preferred environment manager after cloning. API credentials go in a .env file copied from .env.example. The example file covers OpenAI (OPENAI_API_KEY, OPENAI_CHAT_MODEL), Azure OpenAI, Cohere, Mistral, VoyageAI, and local model variables. The LOCAL_MODEL and LOCAL_MODEL_EMBEDDINGS variables configure Ollama-based inference when running without a cloud API.

GraphRAG, Multi-Modal Parsing, and SSO

The repository includes a GraphRAG indexing pipeline as an example extension. GraphRAG builds a knowledge graph during indexing rather than a flat vector index, which the README notes can improve performance on multi-hop questions that require connecting information across separate sections. Enabling it requires setting GRAPHRAG_API_KEY, GRAPHRAG_LLM_MODEL, and GRAPHRAG_EMBEDDING_MODEL in the .env file. The Dockerfile installs the graphrag package only when TARGETARCH is amd64, so arm64 deployments do not include it.

Multi-modal document processing handles figures and tables inside documents. The README lists Azure Document Intelligence and PaddleOCR as configurable parsing options, both available from the settings UI. Adobe PDF Services is a third option for PDF extraction, requiring client credentials from Adobe's developer portal.

Authentication supports both Google OAuth and Keycloak for teams that need SSO. The AUTHENTICATION_METHOD variable in .env selects between the two. Setting it to KEYCLOAK requires KEYCLOAK_SERVER_URL, KEYCLOAK_CLIENT_ID, KEYCLOAK_REALM, and KEYCLOAK_CLIENT_SECRET.

Where Kotaemon Falls Short

Kotaemon is built on Gradio, which constrains how far UI customization can extend. The README acknowledges that developers can add or modify UI elements within Gradio's component model, but a product requiring a fully custom interface that shares nothing with Gradio's layout must replace the front end entirely. In that case the kotaemon library is still useful, but the bundled application is not.

The lite Docker image does not support .doc or .docx files. Adding Office document support requires the full image with Tesseract OCR and LibreOffice, which is substantially larger. File formats outside the core set require the Unstructured library, whose installation steps differ by operating system and are maintained separately from the kotaemon repository.

GraphRAG support is absent on arm64. Teams running kotaemon on Apple Silicon without x86 emulation cannot use GraphRAG indexing. The README does not document a migration path for the document index when moving between minor versions, nor does it describe rollback procedures for the vector store.

Kotaemon Versus a Hand-Assembled RAG Stack

Assembling a RAG pipeline from individual components means selecting a document parser, an embedding model, a vector database, a retrieval layer, a re-ranking step, and a front end independently. Each component is configurable to a degree that kotaemon's settings UI does not expose. Teams with a non-standard chunking strategy, a specific vector database requirement, or a tight integration need with an existing product often find that a hand-assembled stack is the better fit, even at higher initial setup cost.

Kotaemon trades fine-grained control for a working default. The hybrid pipeline, in-browser citation viewer, and multi-user login are built in and require no custom integration work. The kotaemon library also exposes its pipeline components for reuse outside the Gradio interface. A team that wants the underlying retrieval logic without the bundled UI can use it as a Python library dependency and build a different front end around it.

For local inference without a cloud API key, the main-ollama Docker image is the most direct option in this space: it bundles Ollama alongside the full kotaemon stack in a single container.

Editorial conclusion

Kotaemon suits teams that want a working RAG interface without assembling retrieval, re-ranking, and a PDF citation viewer from scratch, and who use either cloud LLM APIs or a local Ollama setup. It is less suited to deployments requiring a fully custom front end that goes beyond Gradio's component model. Before adopting it, verify that the lite Docker image covers your document formats, because Office documents require the larger full image, and GraphRAG indexing is only available on amd64. The last push was on 2026-07-14, and the project is distributed under Apache 2.0.

Frequently asked questions

How does kotaemon compare to Open WebUI?

Kotaemon is designed specifically for document question answering, with a hybrid retrieval pipeline, relevance-scored citations, and an in-browser PDF viewer. It indexes uploaded documents and retrieves specific passages rather than forwarding raw chat messages to an LLM.

Can kotaemon run without a cloud LLM API key?

Yes. The main-ollama Docker image bundles Ollama so kotaemon can run inference locally without sending data to an external API. The .env file uses LOCAL_MODEL and LOCAL_MODEL_EMBEDDINGS to configure the local model and embedding targets.

What document formats does the kotaemon Docker image support?

The lite image supports PDF, HTML, MHTML, and XLSX files. The full image adds .doc, .docx, and other formats supported by the Unstructured library, which includes Tesseract OCR and LibreOffice dependencies.

Official sources

  1. Cinnamon/kotaemon on GitHub
  2. License: Apache-2.0
  3. Project website
  4. README
  5. Releases
For maintainers

Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/cinnamon-kotaemon.svg)](https://hysenlabs.com/projects/cinnamon-kotaemon)
Community notes

Community notes