AI LocalBase: a self-hosted RAG stack built around Qdrant and Ollama
一个本地优先的AI知识库系统(RAG),用于把本地文档接入辅导搜索与大模型对话流程。目前支持md、txt、pdf(文本)、xlsx、cvs类型。支持mcp服务
At a glance
- What is it?
- AI LocalBase is a Go and React knowledge base that indexes local documents into Qdrant and answers questions through Ollama or an OpenAI-compatible API. It is a single-instance system with an embedded MCP server, and the README is explicit about where that design stops working.
- Who is it for?
- Adopt AI LocalBase if you want a self-hosted RAG prototype with a web UI, Qdrant-backed retrieval and an MCP endpoint for external agents, and you are willing to run one backend replica. Do not adopt it if you need horizontal scaling, since the README states the application layer is designed for a single instance and SQLite, the state file and the in-memory MCP jobs do not support shared writes across replicas.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 13 days ago.
- What is it written in?
- Mainly Go, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem AI LocalBase solves, and for whom
A folder of PDFs, Markdown notes and spreadsheets is easy to accumulate and hard to query. AI LocalBase takes that folder, splits and embeds the text, stores the vectors in Qdrant, and puts a React web interface on top so you can ask questions and get answers grounded in the retrieved chunks. The README names the intended audience directly: individuals or small teams who want a working knowledge base question-answering system in a local or self-hosted environment. It is not aimed at organisations that need a managed multi-tenant service. The supported input formats are TXT, Markdown, PDF (text), xlsx and csv, so scanned PDFs without a text layer are outside what the parser handles. Chat history is persisted in a local SQLite database, and model plus knowledge base configuration is written to a local JSON file, which tells you the deployment target is one machine, not a cluster.
How retrieval works: Qdrant collections, reranking and MMR
The mechanism is a fairly standard RAG pipeline with several optional stages layered on top. Documents are uploaded through the UI, split into chunks, embedded in batches, and written to a Qdrant collection. At query time the backend performs vector search, then applies dynamic candidate recall and a keyword coverage rerank. MMR (maximal marginal relevance) is used to drop redundant chunks from the context window, and a low-confidence path triggers a second, wider recall. Embedding results are cached, and the README lists an optional semantic cache as well. Optional stages include Hybrid Search, a semantic reranker, query rewrite and context compression. Two configuration details matter here. The README states that QDRANT_VECTOR_SIZE must match the embedding model's dimension, and that changing to a model with a different dimension requires either a new QDRANT_COLLECTION_PREFIX or rebuilding the old collection. It also recommends switching to a new prefix and rebuilding the index before enabling ENABLE_HYBRID_SEARCH, because the Qdrant collection then needs named dense and sparse vectors. Those are not settings you can flip on an existing index without a rebuild.
Installing AI LocalBase with Docker Compose and answering your first question
The README's shortest path is three steps. Copy the environment template, bring up the stack, and open the front end on port 4173.
cp .env.example .env
docker compose up --buildAfter the build finishes, the front end is at http://localhost:4173, the backend at http://localhost:8080, and the Qdrant HTTP API at http://localhost:6333. The first thing you do in the browser is open Settings and configure a Chat model and an Embedding model; without both, retrieval and answering have nothing to call. Ollama is the simplest local option. The README gives these values for the chat side: provider ollama, base URL http://localhost:11434, model qwen2.5:7b or llama3.2, API key left empty. For embeddings: provider ollama, base URL http://localhost:11434, model bge-m3 or nomic-embed-text, API key left empty. If you prefer a hosted or gateway model, set the provider to openai and point the base URL at your compatible endpoint, for example https://your-api.example.com/v1, with the matching model name and API key. Once both are saved, create a knowledge base, upload a document, and ask a question; the answer should cite retrieved content. For a prebuilt image instead of a local build, the README gives:
AI_LOCALBASE_IMAGE_TAG=v1.4.6 docker compose -f docker-compose.prod.yml up -dOne deployment note worth reading before you start: if ENABLE_AUTH=true and AUTH_PASSWORD is unset, the first visit to the web page opens an initialisation wizard, so on a server you should set AUTH_PASSWORD or AUTH_SETUP_TOKEN in advance rather than leaving that window open.
The embedded MCP server and its JSON-only transport
AI LocalBase ships an MCP server inside the backend rather than as a separate process. It is disabled by default: ENABLE_MCP defaults to false. When enabled, it exposes GET /mcp, GET /mcp/tools and POST /mcp, and supports initialize, ping, tools/list, tools/call and initialisation notifications. Tools are graded read-only, write, or dangerous, and the dangerous ones require a one-time confirmation. Authentication uses API keys with mcp:* scopes; the older MCP token is deprecated and only accepted when ENABLE_MCP_LEGACY_TOKEN=true, which the README frames as a temporary migration switch and notes is equivalent to full MCP permissions. Rate limiting, request timeouts and audit logging are configurable through MCP_REQUESTS_PER_MINUTE (default 120) and MCP_REQUEST_TIMEOUT_SECONDS (default 15). The transport is the part most likely to trip up a client. The README states that the implementation returns JSON, does not issue an MCP session ID, and does not offer SSE long connections, so clients must accept application/json and should not declare only text/event-stream. Cherry Studio is documented as an example, with the URL http://127.0.0.1:8080/mcp, an Authorization: Bearer header carrying a scoped API key, and an Accept header listing both application/json and text/event-stream. The README also says to enable ENABLE_AUTH before turning on MCP on a server.
Single-instance by design: the limitation that decides deployments
The README is unusually blunt about this. The application layer is designed around a single instance: chat records live in local SQLite, application state lives in a file, and MCP jobs are held in memory, none of which support shared writes from multiple backend replicas. The README explicitly says not to run docker compose scale backend=2, and warns against relying on the latest tag, pointing instead at AI_LOCALBASE_IMAGE_TAG for upgrades and rollbacks. That rules out the usual answer to a slow or overloaded deployment. You cannot add a second backend behind a load balancer without first replacing the persistence layer, and the project does not offer that replacement. There is a second boundary in the Qdrant exposure defaults. Qdrant binds to 127.0.0.1 by default; if you set QDRANT_BIND_ADDRESS to a non-loopback address without QDRANT_API_KEY, the compose entrypoint exits with an error rather than starting an open vector store. That is a sensible guard, but it means any remote access to Qdrant requires you to manage an API key and firewall rules yourself. Finally, the default upload ceiling is MAX_UPLOAD_BYTES=26214400 (25 MiB) per file, and the front-end proxy limit NGINX_CLIENT_MAX_BODY_SIZE must stay above it to absorb multipart overhead. Large documents need both values raised.
AI LocalBase compared with a plain Qdrant and script setup
The obvious alternative is assembling the same pieces yourself: Qdrant for vectors, a chunking and embedding script, and a thin chat client against Ollama. That approach gives you full control over chunking strategy, metadata filters and deployment topology, and it scales the way you write it to scale. What it does not give you is the parts AI LocalBase already assembled: a document upload UI with asynchronous indexing jobs that can be queried and cancelled, chat history persisted to SQLite, an evaluation dataset generator that builds RAG test samples from existing knowledge base documents, and a scoped MCP endpoint that external agents can call. The trade-off is the inverse of the DIY route. With AI LocalBase you inherit its single-instance persistence model and its fixed pipeline stages, and you configure retrieval through environment variables rather than code. If your retrieval needs are unusual, or you need multiple backend replicas, the script approach is the better fit. If you want a working UI and an MCP surface this week, AI LocalBase saves the assembly work at the cost of the scaling ceiling.
Maintenance, upgrades and what the MIT licence leaves to you
The repository is not archived and the last push was on 2026-09-02, with v1.4.10 released the same day. That is recent enough that the project is being changed, but the release cadence in the listed history is dense: v1.4.7 on 2026-08-23, v1.4.9 on 2026-09-01, v1.4.10 on 2026-09-02. Frequent patch releases mean you should not treat an upgrade as a no-op. The README's own guidance is to pin AI_LOCALBASE_IMAGE_TAG to a specific version, back up the data directories before upgrading or migrating, and use the development compose files for local source changes rather than depending on latest. Those data directories cover uploaded files, application state, the chat SQLite database and the Qdrant persistence volume, so a backup that omits the Qdrant volume loses the index. On licensing: the repository is MIT, which permits commercial use and modification. That is the extent of what can be said here; the licence text is in the LICENSE file and questions about your own obligations belong with your legal counsel, not with a README.
Editorial conclusion
Adopt AI LocalBase if you want a self-hosted RAG prototype with a web UI, Qdrant-backed retrieval and an MCP endpoint for external agents, and you are willing to run one backend replica. Do not adopt it if you need horizontal scaling, since the README states the application layer is designed for a single instance and SQLite, the state file and the in-memory MCP jobs do not support shared writes across replicas. Before committing, verify that your embedding model's dimension matches QDRANT_VECTOR_SIZE, decide whether you need ENABLE_HYBRID_SEARCH before the first index is built, and confirm your MCP client accepts application/json rather than only text/event-stream.
Frequently asked questions
What is AI LocalBase used for?
It is a local-first RAG knowledge base that indexes local documents into Qdrant and answers questions through Ollama or an OpenAI-compatible API, with a web UI for knowledge base management, document upload and chat history. The README also describes using it as an MCP backend that external agents can call.
Can I run AI LocalBase with my own local models?
Yes. The README documents Ollama as a provider for both chat and embeddings, with base URL http://localhost:11434 and example models qwen2.5:7b or llama3.2 for chat and bge-m3 or nomic-embed-text for embeddings. You configure both in the Settings page after the stack starts.
Which models does AI LocalBase support for chat and embeddings?
Two provider types are documented: ollama and an OpenAI-compatible endpoint. For Ollama the README lists qwen2.5:7b or llama3.2 for chat and bge-m3 or nomic-embed-text for embeddings; for the openai provider you supply your own base URL, model name and API key.
Does AI LocalBase store my documents in a database?
Documents are parsed, chunked and embedded into a Qdrant collection, and chat messages are saved to a local SQLite database. Application state and model configuration are kept in a local JSON file, and the README says to back up the upload, state, SQLite and Qdrant directories before upgrading.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/veyliss-ai-localbase)