PrivateGPT: the open-source API layer between your app and a local model
Complete API layer for private AI applications on local models: RAG, skills, tools, MCP, text-to-sql, and more. Works with any OpenAI-compatible inference server.
At a glance
- What is it?
- PrivateGPT is not a model runner. It is a Claude-shaped API layer that sits in front of any OpenAI-compatible inference server and supplies retrieval, tools, MCP and database access. Here is what it does, how to install it, and where it stops.
- Who is it for?
- Adopt PrivateGPT if you are building an application and already run, or plan to run, an OpenAI-compatible inference server such as Ollama, llama.cpp or vLLM, and you want retrieval, tools, MCP and database access behind one API instead of assembling them yourself. Skip it if you want a finished chat product for non-technical users, or if you need prompt caching or OAuth and organisations, which the README's compatibility table marks as not supported.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 7 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.
DEEP OPEN-SOURCE ANALYSIS
The gap PrivateGPT fills between a local model and a usable application
Running a model locally gets you a completion endpoint. It does not get you file ingestion, retrieval with citations, tool calling, MCP connectors or a way to query a database. Those are the parts every private AI application rebuilds, usually badly, and they are what this project packages.
The README states the positioning plainly: PrivateGPT is the API layer that turns local models into production AI applications, and it does not run models itself. It connects to any server implementing /v1/chat/completions and /v1/models via OPENAI_API_BASE. That constraint is also the design: the inference layer is someone else's problem, and PrivateGPT owns everything above it.
The audience is developers building a product, not end users looking for a chat window. The README says the API is the actual product and that the bundled workbench UI at /ui is a demonstrator, polished enough for demos, videos and internal pilots, but not the core offering. If you were hoping to hand this to a colleague as a private ChatGPT replacement, the project is telling you that is not the shape it was built in.
Architecture: a Claude-shaped API in front of an OpenAI-shaped server
The data flow shown in the README is three layers: your app, agent, workflow or UI on top; the PrivateGPT API in the middle; an OpenAI-compatible inference server underneath. Requests enter as Anthropic-style API calls and are translated down to whatever server you pointed OPENAI_API_BASE at.
That translation layer is the interesting decision. The API reference follows the Anthropic API spec, not the OpenAI one, so the surface you code against is Claude's messages API with streaming, async processing and token counting. Underneath, the project depends on llama-index-core for the retrieval and orchestration work, FastAPI for the web layer, and injector for dependency wiring, according to pyproject.toml.
The compatibility table in the README is the honest part of the documentation. Models, messages, streaming, token counting, files, PDF ingestion, retrieval with citations, embeddings, tool use, streaming tools, built-in web search, web fetch, custom tools, MCP in the API and remote MCP servers are all marked supported. Database querying and CSV analysis are built in rather than routed through tools, which is a genuine difference from the Claude API. Structured outputs are marked supported but inference-dependent, and vision is supported but model-dependent. Skills are marked basic, prompt caching is marked not supported, and OAuth and organisations are marked not supported. Read that table before you design around a capability.
Installing PrivateGPT on macOS, Linux and Windows
The prerequisite is not PrivateGPT itself but a running OpenAI-compatible LLM server. The README names Ollama as the easiest starting point and gives a concrete model pair: qwen3.5:35b for the LLM at roughly 24 GB and mxbai-embed-large for embeddings at roughly 670 MB.
On macOS, installation goes through a Homebrew tap. This is the shortest path and installs the private-gpt command directly.
brew tap zylon-ai/tap
brew install private-gptOn Linux, the README installs uv first and then PrivateGPT as a uv tool, pinned to Python 3.11 and pulling wheels from the project's own wheel index. Note the [core] extra, which selects the feature set baked into the install.
curl -LsSf https://astral.sh/uv/install.sh | sh
uv tool install --python 3.11 \
--find-links https://wheels.privategpt.dev/packages/ \
"private-gpt[core]"Windows uses the same uv tool install through a PowerShell bootstrap. The README shows the backtick line continuation instead of the shell backslash.
powershell -ExecutionPolicy ByPass -c "irm https://astral.sh/uv/install.ps1 | iex"
uv tool install --python 3.11 `
--find-links https://wheels.privategpt.dev/packages/ `
"private-gpt[core]"Before starting the server, pull the models and start Ollama. The README gives these three commands as the example setup.
ollama pull qwen3.5:35b # LLM (~24 GB)
ollama pull mxbai-embed-large # Embeddings (~670 MB)
ollama serveThen launch PrivateGPT with two environment variables pointing at the LLM endpoint and the embedding endpoint. The README writes the ports as placeholders, so substitute the port your server actually listens on.
OPENAI_API_BASE=http://localhost:<llm-port>/v1 \
OPENAI_EMBEDDING_API_BASE=http://localhost:<embedding-port>/v1 \
private-gpt serveThe API then answers on port 8080 and the workbench UI is at http://localhost:8080/ui. From the UI you can send messages, select models from /v1/models, upload documents, test retrieval with citations, enable tools per chat, configure databases, MCP connectors, skills and custom tools, and inspect requests and responses in the API Debugger. For a first real use, upload a document and ask a question about it: the retrieval-with-citations path is the one worth validating before you build anything on top.
Where PrivateGPT is the wrong tool
The clearest limitation is stated by the project itself: PrivateGPT does not run models. If you have no inference server, you have nothing. The README points at Ollama, llama.cpp and vLLM as examples, and the requirement is specific: the server must implement /v1/chat/completions and /v1/models. A server that only exposes a native completion route will not work.
The second limitation is the feature table. Prompt caching is marked not supported, and OAuth and organisations are marked not supported. For a single-team deployment behind a token, that is fine. For a multi-tenant product where each customer needs separate identity and access, you are building that layer yourself. Skills are marked basic rather than supported, so do not plan a skills-heavy design on the assumption that the implementation is complete.
The third is hardware. The README's own example model is roughly 24 GB, which is a real machine, not a laptop with integrated graphics. Nothing in the project changes that; it is a consequence of the model you choose, and the API layer is indifferent to it.
Finally, the UI. It is described as a demonstrator. If your requirement is a chat product for people who will not read an API reference, the workbench is not that product, and the README says so.
PrivateGPT compared with Open WebUI and Ollama
The related searches pair PrivateGPT against Open WebUI, ChatGPT, GPT4All and Ollama, and the differences are not subtle. Ollama is the inference server in this stack, not a competitor: the README's own quickstart pulls models with ollama pull and starts ollama serve, then points PrivateGPT at it. If you only need to talk to a local model, Ollama alone is enough and PrivateGPT adds a layer you will not use.
Open WebUI is a chat interface. PrivateGPT ships a UI too, but the README is explicit that the API is the product and the UI is a demonstrator. The comparison that matters is therefore not interface against interface but application layer against interface: Open WebUI gives you a place to chat, PrivateGPT gives you an API with file ingestion, retrieval with citations, tool use, MCP connectors and built-in database and CSV access for your own code to call.
Against ChatGPT the split is architectural rather than feature-by-feature. ChatGPT is a hosted service; PrivateGPT is an Apache-2.0 layer you run next to your own inference server, with the compatibility table above telling you which Claude API capabilities are present, partial or absent. If you need the absent ones, the local route costs you that functionality today.
Maintenance, releases and the Apache-2.0 licence
The repository is not archived. The most recent push recorded is 2026-06-18, which is the same date as the v1.0.1 release, following v1.0.0 and v1.0.0-rc7 on 2026-06-03. The 1.0 line is recent, and the CHANGELOG.md at the repository root is where release-by-release changes are recorded; the README does not document a rollback or downgrade procedure, so pinning a version before upgrading is your own decision rather than a documented one.
Upgrade cost is dominated by dependency surface, not by PrivateGPT's own code. pyproject.toml pins requires-python to >=3.11,<3.12, which means a Python 3.12 environment is out of range, and it constrains llama-index-core to >=0.14.21,<0.15.0 alongside fastapi, injector, pyyaml, arq and others. Optional extras are split by capability: sdk-openai, llm-openai, llm-openai-compatible, llm-mistral, llm-anthropic and tokenizer-local among the ones visible. If you build from the Dockerfile, the EXTRAS build argument decides which system packages are installed, including libreoffice, pandoc and ghostscript for documents, ffmpeg for media, and a set of Chromium libraries for web scraping. A heavier extra means a heavier image and a longer rebuild.
The licence is Apache-2.0, which permits commercial use and modification and includes an explicit patent grant. That is a permissive licence, not a copyleft one, so distributing a modified PrivateGPT does not force you to publish your changes. This is a description of the licence text, not legal advice; if you are embedding it in a product, have your own counsel read the terms.
Editorial conclusion
Adopt PrivateGPT if you are building an application and already run, or plan to run, an OpenAI-compatible inference server such as Ollama, llama.cpp or vLLM, and you want retrieval, tools, MCP and database access behind one API instead of assembling them yourself. Skip it if you want a finished chat product for non-technical users, or if you need prompt caching or OAuth and organisations, which the README's compatibility table marks as not supported. Before committing, verify three things against your own setup: that your inference server implements both /v1/chat/completions and /v1/models, that your embedding endpoint is reachable through OPENAI_EMBEDDING_API_BASE, and that the extras you need are present in your build, since the Dockerfile decides which system packages get installed from the EXTRAS argument.
Frequently asked questions
What is PrivateGPT?
It is an open-source API layer that turns local models into AI applications, providing retrieval, tools, MCP connectors and database access. The README states that PrivateGPT does not run models itself and connects to any OpenAI-compatible inference server via OPENAI_API_BASE.
How do I install PrivateGPT?
On macOS the README uses a Homebrew tap and brew install private-gpt. On Linux and Windows it installs uv first, then runs uv tool install --python 3.11 with the project's wheel index and the private-gpt[core] extra.
How do I install PrivateGPT on Windows?
The README bootstraps uv through PowerShell with an irm command, then installs the package with uv tool install --python 3.11, pointing --find-links at https://wheels.privategpt.dev/packages/ and requesting private-gpt[core]. Environment variables are set with $env: syntax before running private-gpt serve.
Is PrivateGPT free?
The project is released under Apache-2.0, which permits commercial use and modification and includes a patent grant. The README does not describe a paid tier for the software itself; the cost you carry is the hardware and inference server you run it against.
How does PrivateGPT compare with Ollama?
They occupy different layers. Ollama is the inference server in the README's own quickstart, where models are pulled with ollama pull and served with ollama serve; PrivateGPT sits above it and adds ingestion, retrieval with citations, tools and MCP connectors. If you only need to talk to a local model, Ollama alone is sufficient.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/zylon-ai-private-gpt)
Community notes