# localGPT: A Private On-Premise Document Intelligence Platform Built on Ollama

> localGPT is an MIT-licensed Python project that lets you query your own documents locally using Ollama-hosted LLMs, with a hybrid vector and full-text search engine, a cross-encoder reranker, and a smart router that chooses between RAG and direct LLM answering per query.

**PromtEngineer/localGPT** — Chat with your documents on your local device using GPT models. No data leaves your device and 100% private. 

- Repository: https://github.com/PromtEngineer/localGPT
- Stars: 22,199 · Forks: 2,454
- Language: Python
- License: MIT
- Published: 2026-09-21 · Updated: 2026-09-21 · Language: en
- Canonical page: https://hysenlabs.com/projects/promtengineer-localgpt

## A Document Intelligence Platform That Keeps Data on Your Machine

localGPT addresses a specific concern: organizations and individuals who want to ask questions of their own documents using LLM-based retrieval but cannot send those documents to an external API. Legal documents, medical records, internal strategy files, and source code with proprietary logic fall into this category. The README names the core promise directly: no data leaves your machine.

The system accepts PDF, DOCX, HTML, Markdown, and TXT files, processed by Docling. Documents are indexed into LanceDB, a local vector store. Queries are answered by Ollama-hosted language models running on the same machine. There is no cloud dependency in the serving path.

The prerequisites are specific. The README specifies Python 3.10 or later (with 3.11 recommended, matching the Docker images), Node.js 20 or later, and Ollama installed and running locally. A minimum of 8GB of RAM is listed, with 16GB or more recommended. Docker is listed as optional for containerized deployment.

The target user is a developer or technical team, not a casual consumer. The setup involves installing Ollama, pulling specific model weights, and starting multiple services.

## How the Hybrid Search, Reranker, and Smart Router Work Together

localGPT goes beyond a straightforward vector-similarity search. The README describes a pipeline with multiple distinct stages, each with a specific job.

The retrieval step uses hybrid search: dense vector embeddings stored in LanceDB are combined with LanceDB's native full-text search, with the results merged using Reciprocal Rank Fusion (RRF). The README describes this as a no-weights-to-tune approach, since RRF combines ranked lists without requiring a manually calibrated mixing coefficient.

After retrieval, a cross-encoder reranker performs a second pass over the combined candidate set. The reranker scores each candidate against the query directly, rather than against a query embedding. This is enabled by default. The reranker model is Qwen/Qwen3-Reranker-4B, as shown in the docker-compose.yml default configuration.

The smart router decides, per query, whether to use the full RAG pipeline or to answer directly from the LLM. Document overviews are written to index_store/overviews/<id>.jsonl and used by the router to inform this decision. Complex factual questions about the documents go to RAG; questions the LLM can answer without document context go directly to the model.

The default generation model is qwen3.5:9b and the contextual enrichment model is qwen3.5:4b. The embedding model is microsoft/harrier-oss-v1-0.6b. All four are configurable through environment variables.

## Setting Up localGPT with Docker

The README provides a Docker-based Quick Start as the primary setup path. First clone the repository, install Ollama on the host machine, and pull the required models:

```bash
git clone https://github.com/PromtEngineer/localGPT.git
cd localGPT
curl -fsSL https://ollama.ai/install.sh | sh
ollama pull qwen3.5:9b
ollama pull qwen3.5:4b
```

Then start Ollama in the background and launch the Docker stack:

```bash
ollama serve
./start-docker.sh
```

The application is then accessible at:

```bash
open http://localhost:3000
```

For Docker-in-container Ollama (when you do not want Ollama on the host), the README provides an alternative:

```bash
./start-docker.sh container
docker compose --profile with-ollama exec ollama ollama pull qwen3.5:9b
docker compose --profile with-ollama exec ollama ollama pull qwen3.5:4b
```

The docker-compose.yml defines four key environment variables for the RAG API service, with defaults shown:

```bash
GENERATION_MODEL=${GENERATION_MODEL:-qwen3.5:9b}
ENRICHMENT_MODEL=${ENRICHMENT_MODEL:-qwen3.5:4b}
EMBEDDING_MODEL=${EMBEDDING_MODEL:-microsoft/harrier-oss-v1-0.6b}
RERANKER_MODEL=${RERANKER_MODEL:-Qwen/Qwen3-Reranker-4B}
```

The backend API serves on port 8000, the RAG API on port 8001, and the Next.js frontend on port 3000.

## Document Processing: Formats, OCR, and the Docling Pipeline

localGPT uses Docling for document ingestion. The supported formats are PDF, DOCX, HTML, HTM, Markdown, and TXT. PDFs without a text layer (scanned documents) trigger an OCR fallback within Docling. The OCR engine selection is automatic: OcrMac on macOS, then EasyOCR, RapidOCR, tesserocr, or the tesseract CLI, chosen from whatever is installed.

Chunks are enriched at index time by a small LLM that generates context for each chunk. This contextual enrichment is described as inspired by the Contextual Retrieval approach and uses the enrichment model (qwen3.5:4b by default). A short per-document summary is also written to index_store/overviews/ and used by the smart router.

Late Chunking is available but ships disabled by default. The README describes a 2026-08-18 ablation that measured its impact at the noise floor for single-turn conversations, while it doubles the number of vectors written per index. A single flag, `retrieval.latechunk.enabled`, re-enables it for cases where multi-turn conversations with drifting phrasing might benefit.

The configuration architecture is notable: a single file, rag_system/main.py, holds every default and can be overridden by environment variables. The .env.example file in the repository documents the configurable variables with comments explaining each one.

## What localGPT Does Not Cover

localGPT is a single-user local tool. The README does not document multi-user authentication, role-based access controls, or document-level permissions. Teams with multiple analysts who need isolated document collections and access restrictions need to add those controls separately or evaluate a different platform.

The system has no versioned release track. The README shows no tagged releases, and the last push was on 2026-08-26. The requirements.txt pins transformers at 4.51.0 and torch at 2.4.1. An older torch pin on a newer CUDA driver setup can cause compatibility issues that require manual intervention.

The Answer Verification feature ships disabled. The README documents a specific reason: a 2026 ablation measured zero verdict flips from disabling it, meaning it appended confidence annotations without changing any answers. It can be re-enabled with `verification.enabled`, but it does not improve accuracy based on the documented evaluation.

Late Chunking and answer verification both exist as toggleable options, which means the default configuration reflects deliberate choices from measured evaluations rather than arbitrary defaults. The README points to the eval/decisions/ directory and eval/DECISIONS.md for the reasoning behind each enabled or disabled component.

## localGPT versus a Cloud-Based Document Q&A Service

The obvious alternative to localGPT is a cloud-hosted document intelligence service that processes your files on external servers. The trade-off is straightforward: cloud services require no local GPU or memory headroom, handle model updates automatically, and often provide polished multi-user interfaces. localGPT requires you to provision hardware that can run qwen3.5:9b comfortably and to manage the Ollama model stack yourself.

The README frames this trade-off as the core product decision: 100% private, with no data leaving the machine. For users under legal or contractual obligations that restrict data processing to on-premise infrastructure, cloud alternatives are not viable regardless of their convenience. localGPT's value is specifically for that constraint.

PrivateGPT is a frequently compared alternative, appearing in the related searches for this project. localGPT's README does not document a comparison with PrivateGPT, but the architectural difference from a purely cloud-based service is the central design choice in both projects.

## MIT License and Maintenance Notes

localGPT is released under the MIT license, which imposes no distribution restrictions and requires only attribution. The project has no tagged releases; development is tracked through commits on the main branch. The last push was on 2026-08-26, which is 33 days before today.

The project uses a multi-container Docker architecture with three separate Dockerfiles: Dockerfile.backend, Dockerfile.frontend, and Dockerfile.rag-api. The frontend is a Next.js application. The backend and RAG API are Python services. The docker-compose.yml wires them together and defines the service health checks and restart policies.

The requirements.txt file pins transformers at 4.51.0 and torch at 2.4.1, with torchaudio unpinned. When running outside Docker on a system with a newer CUDA toolkit, the pinned torch version may not match, and installation may fail or produce version conflict warnings. The Docker images use python:3.11-slim as the base, which avoids most host-level compatibility issues.

## Conclusion

localGPT is a strong fit for individuals and teams who need to query sensitive documents without sending them to an external API, and who are prepared to run the hardware and maintain the Ollama model stack locally. It is less appropriate for teams that need multi-user authentication, enterprise access controls, or a system with a stable versioned release track. Before deploying, confirm that Ollama is running and the qwen3.5:9b and qwen3.5:4b models have been pulled, as the system does not fall back gracefully if Ollama is unreachable at startup.

## FAQ

### What is localGPT?

localGPT is a private, on-premise document intelligence platform that lets you ask questions about your own files using Ollama-hosted language models. It uses hybrid search combining LanceDB vector search and full-text search, a cross-encoder reranker, and a smart router that decides whether each query needs RAG or can be answered directly.

### How do I install localGPT?

Clone the repository, install Ollama locally, pull the qwen3.5:9b and qwen3.5:4b models with `ollama pull`, then run `./start-docker.sh` to start the full stack. The frontend is then available at http://localhost:3000. Python 3.10 or later and Node.js 20 or later are required for non-Docker setups.

### How do I use localGPT?

After starting the application, open the web UI at http://localhost:3000, create an index by uploading PDF, DOCX, HTML, Markdown, or TXT files, and then start a chat session to ask questions. Each answer includes the source chunks it was grounded in for traceability.

### How does localGPT compare to Ollama?

Ollama is the LLM inference layer that localGPT depends on for running language models locally. localGPT builds a full document intelligence system on top of Ollama, adding document ingestion via Docling, LanceDB vector storage, hybrid search, reranking, and a smart query router. Ollama alone handles model serving but provides no document indexing or retrieval pipeline.

### What are the alternatives to localGPT?

The README does not document alternatives. PrivateGPT is a frequently searched comparison, as both projects target private, local document Q&A. Cloud-hosted document intelligence services are the primary alternative class, with the key difference being that cloud services process your documents on external servers rather than on your own hardware.

## Sources

- [Issues](https://github.com/PromtEngineer/localGPT/issues)
- [License: MIT](https://github.com/PromtEngineer/localGPT/blob/main/LICENSE)
- [PromtEngineer/localGPT on GitHub](https://github.com/PromtEngineer/localGPT)
- [README](https://github.com/PromtEngineer/localGPT/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/promtengineer-localgpt
