# Local RAG: Offline Retrieval Augmented Generation with Ollama and LlamaIndex

> Local RAG is an open-source Streamlit application for running retrieval augmented generation entirely on your own machine or network. It ingests local files, GitHub repositories, and websites, routes queries through local Ollama models, and keeps all data off third-party servers.

**jonfairbanks/local-rag** — Ingest files for retrieval augmented generation (RAG) with open-source Large Language Models (LLMs), all without 3rd parties or sensitive data leaving your network.

- Repository: https://github.com/jonfairbanks/local-rag
- Stars: 763 · Forks: 96
- Language: Python
- License: GPL-3.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/jonfairbanks-local-rag

## What Local RAG Solves and Who It Is For

Local RAG addresses a specific tension in retrieval augmented generation: most RAG implementations send document content to a cloud API for embedding and querying, which creates privacy and compliance concerns for sensitive documents. Local RAG keeps the entire pipeline on your machine or local network. Chat models run through Ollama, embeddings run through Ollama or a local Hugging Face embedding model, and indexed source content lives in a local volume.

The target users are developers and researchers who work with documents that cannot leave a controlled environment: internal company documents, legally sensitive materials, personal notes, or any corpus where sending content to a third-party API is unacceptable. The v2.0.0 release, dated 2026-05-18 on GitHub, is the current major version. The last push was on 2026-09-22.

The GPL-3.0 license means you can use and modify Local RAG freely, but any distributed derivative work must also be licensed under GPL-3.0.

## Ingestion Sources: Files, GitHub Repositories, and Websites

Local RAG supports three ingestion sources. Local files are ingested directly from your filesystem. GitHub repositories are ingested by URL, letting you index a codebase or documentation repository without manually downloading it first. Websites are ingested by URL, which means you can index a public page or documentation site for querying.

The README documents upload, URL, repository, and ingestion guardrails, which implies there are checks on what can be submitted for ingestion. This likely includes size limits and URL validation to prevent the ingestion process from being used against arbitrary internal network addresses.

The Streamlit interface provides browser-local settings persistence, meaning your configuration choices (model selection, endpoint URLs) persist in the browser without being sent to a server. Chat history export is also documented as a feature.

LlamaIndex handles the streaming RAG pipeline: document chunking, embedding, vector storage, retrieval, and generating streaming responses that appear incrementally in the interface.

## Running Local RAG with Docker

The repository includes three docker-compose variants for different hardware configurations. The default docker-compose.yml assumes an NVIDIA GPU:

```bash
docker compose up
```

The compose file configures the container with specific resource limits: mem_limit of 8g, 4 CPUs, and a pids_limit of 512. The container runs as an unprivileged user (appuser), drops all Linux capabilities, mounts the filesystem read-only, and uses tmpfs for temporary directories. The Streamlit interface listens on port 8501, bound to localhost only (127.0.0.1:8501:8501/tcp).

For CPU-only environments, the README documents a separate compose file:

```bash
docker compose -f docker-compose.yml-cpu up
```

For AMD GPU environments using ROCm:

```bash
docker compose -f docker-compose.yml-rocm up
```

The Ollama endpoint is configured via the LOCAL_RAG_OLLAMA_ENDPOINTS environment variable, which defaults to http://localhost:11434,http://127.0.0.1:11434. You can configure multiple endpoints as a comma-separated list. The container has an extra_hosts entry that maps localhost to the host gateway, which allows it to reach an Ollama process running on the host machine rather than inside Docker.

## Embedding Model Options: Ollama and Hugging Face

Local RAG supports two embedding backends. The first is Ollama embedding models, which means the same Ollama instance you use for chat can also serve embeddings without any additional setup. The second is local Hugging Face embedding models, which requires the embeddings optional dependency package.

The choice between the two has practical implications. Ollama embedding models run through the Ollama API and share the same endpoint configuration as chat models. Local Hugging Face models are loaded directly by the application process, which adds memory overhead but avoids any Ollama dependency for the embedding step.

The Dockerfile builds from python:3.14-slim and uses Pipenv for dependency management. The build process includes a fix_cusparselt_metadata.py script that patches a known metadata issue with the cuSPARSELt package. This suggests the application has been calibrated to work with specific GPU library versions and may require attention when upstream package versions change.

## Security Configuration in the Docker Setup

The docker-compose.yml applies several explicit security settings worth noting. security_opt: no-new-privileges:true prevents the container from acquiring new Linux privileges at runtime. cap_drop: ALL removes all Linux capabilities from the container, including network-binding and file system capabilities that are usually present by default. read_only: true makes the container filesystem read-only except for the explicitly defined tmpfs mounts.

The tmpfs mounts provide write space in /tmp (limited to 512MB, no executable, no setuid, no device) and /home/appuser/.cache (2GB for model cache, same restrictions). These settings mean the application cannot write to the container filesystem outside those two paths.

The health check in the Dockerfile pings the Streamlit health endpoint:

```bash
python -c "import urllib.request; urllib.request.urlopen('http://localhost:8501/_stcore/health', timeout=5)"
```

This gives Docker visibility into whether the Streamlit process is healthy and can respond to requests.

## Limitations and How Local RAG Differs from Cloud RAG Services

Local RAG's primary constraint is compute: the application requires up to 8GB of memory and 4 CPUs as configured in the compose file, and meaningful RAG response quality depends on the capability of the local Ollama model in use. A small local model will produce weaker responses than a large cloud model, which is the fundamental tradeoff of offline operation.

The application does not include its own vector database; LlamaIndex manages the in-process vector store. Large document collections may require significant memory or longer ingestion times, and the README does not document specific limits.

Cloud RAG services such as those built on OpenAI's API or Anthropic's API offload model execution to managed infrastructure, which provides access to larger models without local hardware requirements. The difference is data control: cloud services send document content to external servers, while Local RAG processes everything locally. For documents where data residency is not a concern and model quality matters most, a cloud-based pipeline may produce better results. For privacy-sensitive or air-gapped environments, Local RAG's local execution is the defining capability.

## Conclusion

Local RAG is the right choice for developers and researchers who need RAG over private or sensitive documents and cannot send that content to cloud APIs. The Docker deployment is the recommended path for most users and is straightforward, but the container requires up to 8GB of memory and 4 CPUs. The application depends on a running Ollama instance configured at LOCAL_RAG_OLLAMA_ENDPOINTS; if Ollama is not reachable, the application will not function. Before deploying, verify the correct compose variant for your hardware: the default compose file assumes an NVIDIA GPU, while docker-compose.yml-cpu and docker-compose.yml-rocm cover CPU-only and AMD GPU environments respectively.

## FAQ

### How do you set up a local RAG system with this project?

The setup documentation at docs/setup.md covers the full process. The quickest path is Docker: run docker compose up with the Ollama endpoint configured in LOCAL_RAG_OLLAMA_ENDPOINTS, open port 8501 in your browser, and use the interface to ingest files, GitHub repositories, or websites before asking questions.

### What is Local RAG and what does it do?

Local RAG is an open-source Streamlit application that runs retrieval augmented generation using local Ollama models and keeps all document content on your own machine or network. It supports ingesting local files, GitHub repositories, and websites, then answers questions using retrieved context from those sources.

### How do you make a local RAG pipeline from scratch?

Local RAG bundles the full pipeline: ingest sources via the Streamlit interface, which indexes them with LlamaIndex and stores embeddings locally. The pipeline uses Ollama or local Hugging Face models for embedding and an Ollama chat model for generation. The README documentation at docs/pipeline.md describes the pipeline architecture in detail.

### Is RAG still useful given advances in large context windows?

Large context windows reduce one common use of RAG (fitting more documents into a single prompt) but do not replace the retrieval step for very large document collections that exceed any context limit. Local RAG specifically targets the case where documents also cannot be sent to a cloud API, which is a separate constraint from context length.

## Sources

- [Issues](https://github.com/jonfairbanks/local-rag/issues)
- [jonfairbanks/local-rag on GitHub](https://github.com/jonfairbanks/local-rag)
- [License: GPL-3.0](https://github.com/jonfairbanks/local-rag/blob/develop/LICENSE)
- [README](https://github.com/jonfairbanks/local-rag/blob/develop/README.md)
- [Releases](https://github.com/jonfairbanks/local-rag/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/jonfairbanks-local-rag
