Model or dataset
TilmanGriesel/chipper avatar
TilmanGriesel/chipper

Chipper: a self-hosted Ollama RAG interface built on Haystack and Elasticsearch

✨ AI interface for tinkerers (Ollama, Haystack RAG, Python)

485 stars48 forksPythonMIT

At a glance

What is it?
Chipper is a Dockerized Python service that adds retrieval, document chunking, web scraping and an Ollama API proxy to a local model stack. It is a personal project aimed at tinkerers, and the README says so.
Who is it for?
Adopt Chipper if you already run Ollama locally and want a self-hosted knowledge base behind a web UI, CLI or third-party Ollama client, and you accept the README's own warning that it is a personal project not designed for commercial or production use. Skip it if you need a supported platform with a release cadence you can plan around, or if you have no Ollama instance and no interest in running Elasticsearch alongside it.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 134 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Chipper adds to a plain Ollama install

Ollama gives you a model endpoint. It does not give you a place to put documents, split them, embed them, retrieve them and paste the results into a prompt. Chipper is the layer that does that, plus a web interface and a CLI on top. The README describes it as "a web interface, CLI, and a modular, hackable, and lightweight architecture for RAG pipelines, document splitting, web scraping, and query workflows".

The audience is narrow and stated plainly. The author explains the project began as a personal tool to help his girlfriend explore characters and ideas for a book using local RAG and LLMs, keeping the work off cloud services. That origin explains the feature set: audio transcription, web scraping, document chunking and a chat UI are the things you need when you are working through a manuscript, not the things an enterprise search team asks for. The README also carries a note that this is a personal project not designed for commercial or production use, and asks anyone deploying it in production to conduct their own due diligence. Take that at face value.

If your goal is a hosted assistant with a support contract, this is the wrong repository. If your goal is a local, inspectable pipeline you can modify, the scope is right.

How the retrieval pipeline is assembled

The stack is named in the README: Haystack for the pipeline, Ollama for inference, Hugging Face for remote models, Elasticsearch for vector storage, Docker for packaging, TailwindCSS for the interface. The data flow follows from those pieces. Documents are ingested and split into chunks, the chunks are embedded, and the vectors land in Elasticsearch. At query time the pipeline retrieves matching chunks and feeds them to the selected model along with your system prompt.

The parts you can change are model selection, query parameters and system prompts, which the README lists under customizable RAG pipelines. Chunking is a named feature rather than a fixed internal step, so splitting behaviour is something you configure rather than something hidden.

Two features are worth separating from the RAG story. The Ollama API proxy sits between an Ollama client and an Ollama instance, so a client such as Enchanted or Open WebUI can talk to Chipper instead of directly to Ollama and get retrieval-augmented answers without any client-side change. And distributed processing lets you chain multiple Chipper instances together for workload distribution. The README lists that capability but does not walk through a topology, so treat it as something to read the project site for before you design around it.

Installing Chipper and running a first query

The README points installation at the Quickstart page on chipper.tilmangriesel.com rather than reproducing steps inline, so the authoritative sequence lives there. What the repository does show is the shape of the deployment: a docker directory, an examples/docker_compose directory, and a setup.py that generates configuration and probes for a reachable Ollama instance.

The setup script defines two Ollama URLs. One is the internal address used when Ollama runs as a container alongside Chipper, and one is the external address used when Ollama runs on your host.

python
DEFAULT_INTERNAL_OLLAMA_URL = "http://ollama:11434"
DEFAULT_EXTERNAL_OLLAMA_URL = "http://host.docker.internal:11434"
DEFAULT_EXTERNAL_LOCAL_OLLAMA_URL = "http://localhost:11434"

Those constants tell you the expected port and hostnames. If you run Chipper in Docker and Ollama on the host, the container reaches it through host.docker.internal, which the setup script itself flags as not Linux compatible.

The script also carries a placeholder API key constant, which is the value the generated configuration uses until you replace it.

python
EXAMPLE_API_KEY = "EXAMPLE_API_KEY"

Once the service is up, the README says to run the /help command to learn how to switch models, update the embeddings index and more. That is the first thing to type, because it enumerates the available commands from the running instance rather than from a page that may lag the code. If you would rather see the interface before deploying anything, the README links a live demo at demo.chipper.tilmangriesel.com.

The GPU profile check that decides where Ollama runs

setup.py does more than write a config file. It inspects the platform and a GPU profile argument, then decides whether Ollama can run inside the same compose stack or has to be reached externally. The logic is explicit: Darwin, a cpu profile and amd-linux all set a flag requiring an external Ollama.

python
is_darwin = system in ["Darwin"]
is_cpu_profile = gpu_profile == "cpu"
is_amd_linux = gpu_profile == "amd-linux"
requires_external = is_darwin or is_cpu_profile or is_amd_linux

This is a real constraint, not a formality. On macOS there is no containerized GPU passthrough for Ollama, so the container cannot host the model. On a CPU-only profile the same reasoning applies for practical reasons. On AMD Linux the project does not attempt to bundle the GPU runtime. Before the script gets that far it calls a socket check against the Ollama URL with a five second timeout, so a misconfigured host fails early with a connection error rather than at first query.

The upshot is that "fully containerized" in the README describes Chipper, not necessarily your model server. Plan for two moving parts on most laptops.

Where Chipper stops being the right tool

The README's own production warning is the first limitation, and it is not boilerplate. A personal project with a single maintainer carries different expectations around upgrade paths, migration and incident response than a platform you buy. If your use case involves other people's data under a compliance regime, the due diligence the README asks for is not a formality either.

Elasticsearch is the second constraint. It is a JVM service with its own memory footprint and its own operational surface. Adding it to a laptop setup to answer questions about a few hundred pages of text is a heavy trade. Projects that keep embeddings in a local file or an embedded index avoid that weight at the cost of scale, and for a single-user knowledge base that trade is often the better one. Chipper's README frames Elasticsearch as the choice for "scalable indexing", which is a fair description of what you get and also of what you pay for.

The third gap is documentation depth in the repository itself. Installation lives on the project site, not in the README. Distributed processing is listed as a feature with no configuration example in the README. If you need to understand a mechanism before you trust it, expect to read services/ and tools/ rather than a specification. That is consistent with a project that calls itself hackable, but it is a cost you should price in.

Chipper versus wiring Haystack yourself

The obvious alternative is not another product. It is the underlying library. Haystack is a Python framework for building retrieval pipelines, and Chipper is one opinionated assembly of it with a UI, a CLI, a Docker layout and an Elasticsearch backend already chosen. Building directly on Haystack means you pick your own document store, your own chunking strategy and your own interface, and you own the deployment. The difference is not capability, it is who makes the decisions and who maintains the result.

A second comparison is closer to home: running Ollama with a client such as Open WebUI and no retrieval layer at all. That setup is smaller and has fewer failure points, and it is the correct answer if your questions are about general knowledge rather than about your own documents. Chipper's proxy mode is designed for exactly this situation, letting an existing Ollama client gain server-side knowledge base embeddings without being replaced. So the honest framing is that Chipper competes with your own glue code, not with your chat client.

A third option worth naming is the Hugging Face API path that Chipper itself supports. The README lists local models through Ollama and remote models through the Hugging Face API as two supported modes. Choosing the remote mode removes the local GPU question entirely, at the cost of sending your queries off your machine, which is the opposite of the privacy motivation the project started from.

Licence, upgrades and what maintenance looks like

Chipper is MIT licensed. In practice that means you can fork it, modify it and ship it inside your own product, provided you keep the copyright notice and permission notice with the distribution. It does not give you warranty, and the README independently disclaims production readiness. Nothing here is legal advice; read LICENSE in the repository root if the terms matter to your organisation.

The upgrade picture is uneven. Releases in the repository are v.2.3.0 in February 2025, v2.4.0 in March 2025 and v2.5.0 in June 2025. The last push to the default branch was on 2026-05-19, roughly four months before today's date, so the repository is not archived but the recent release cadence is slower than the spring 2025 stretch. Treat the release notes and the project site as the source of truth for what changed between versions rather than assuming a smooth migration.

Upgrade cost concentrates in three places: the Elasticsearch index format, the generated configuration that setup.py writes, and the API key handling. If a release changes the index schema, reindexing is your problem, and the README does not document a migration path. Pin your Chipper image tag, keep your documents in a form you can re-ingest, and read the release notes for the version you are moving to before you pull it.

Editorial conclusion

Adopt Chipper if you already run Ollama locally and want a self-hosted knowledge base behind a web UI, CLI or third-party Ollama client, and you accept the README's own warning that it is a personal project not designed for commercial or production use. Skip it if you need a supported platform with a release cadence you can plan around, or if you have no Ollama instance and no interest in running Elasticsearch alongside it. Before committing, read the Quickstart on chipper.tilmangriesel.com, check whether your GPU profile is one of the external-Ollama cases that setup.py detects, and confirm the MIT licence terms and the API key handling fit your setup.

Frequently asked questions

What is TilmanGriesel/chipper?

It is a Dockerized Python service that provides a web interface, a CLI and a modular architecture for RAG pipelines, document splitting, web scraping and query workflows, built on Haystack, Ollama, Hugging Face and Elasticsearch. It can also act as a proxy between an Ollama client and an Ollama instance.

Does Chipper run without an internet connection?

The README lists an offline web UI as a feature, describing it as working without an internet connection using vanilla JavaScript and TailwindCSS, and Edge TTS is described as a WebAssembly-based client-side generator. Local models through Ollama are one of the two supported model modes, the other being remote models through the Hugging Face API.

Can Chipper add retrieval to an existing Ollama client?

Yes. The README describes an Ollama API proxy that extends Ollama with retrieval capabilities and lists Enchanted and Open WebUI as example clients, allowing server-side model selection, query parameters and system prompt overrides. The proxy is also described as supporting API key-based and bearer token service authentication.

Is Chipper suitable for production use?

The README states that it is a personal project and not designed for commercial or production use, and asks anyone deploying it in a production environment to conduct their own due diligence. The MIT licence supplies no warranty.

Official sources

  1. License: MIT
  2. Project website
  3. README
  4. Releases
  5. TilmanGriesel/chipper on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/tilmangriesel-chipper.svg)](https://hysenlabs.com/projects/tilmangriesel-chipper)