Self-hosted service
2dogsandanerd/Knowledge-Base-Self-Hosting-Kit avatar
2dogsandanerd/Knowledge-Base-Self-Hosting-Kit

Knowledge Base Self-Hosting Kit: A Docker RAG Stack That Reads Code and Prose Differently

A Docker-powered RAG system that understands the difference between code and prose. Ingest your codebase and documentation, then query them with full privacy and zero configuration.

356 stars35 forksPythonNOASSERTION

At a glance

What is it?
A self-hosted retrieval layer built on ChromaDB, Docling and a FastAPI backend, with an MCP server for agents. It installs with docker compose, but the repository has not been pushed since 2026-03-14 and the README opens with an authorship dispute.
Who is it for?
Adopt it if you want a local ChromaDB plus Docling ingestion stack that you can read end to end, and if you are comfortable pinning the image tags yourself. Do not adopt it if you need a maintained dependency with releases, or if you are building a product on top of the MCP connector, since the README states the author has stopped publishing open source work.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Activity is slowing. The repository last received commits 6 months ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem Knowledge Base Self-Hosting Kit targets

Most retrieval-augmented generation demos treat every file as prose. You chunk a Python module at 512 characters with 128 overlap and the retriever happily returns half a function body with the docstring cut off. The README for this project claims the opposite approach: it describes a system that "understands the difference between code and prose", and its ingestion pipeline is built on ChromaDB plus Docling, with hybrid chunking that fuses vector search and BM25. That hybrid step matters more than the marketing line. Keyword search catches an exact symbol or config key that an embedding model will smooth away; vector search catches the paraphrased question. Running both and merging scores is the standard fix, and the query endpoint is documented as returning citation data alongside the answer.

The intended user is an engineer or small team that wants a private memory layer for local models and agents. The default configuration points at Ollama with llama3:latest and nomic-embed-text, and the compose file mounts the host's ~/.ollama directory into the container so models are shared rather than re-downloaded. If you already run Ollama on the host, that mount is the single most useful line in the stack. If you do not, you pay the download once inside the container instead.

It is worth being precise about what the code-versus-prose claim does and does not cover. The README's feature list puts "hybrid chunking (vector + BM25)" under ingestion modes, but the mechanism it names is a retrieval mechanism: two scoring paths merged at query time. Whether the parser applies different splitting rules to a .py file than to a Markdown page is not stated anywhere in the README. The safest reading is that Docling normalises everything into the same pipeline and the code-awareness lives in how results are ranked and cited. If you adopt this expecting syntax-aware chunk boundaries for source files, check that assumption against the actual ingestion code before you load a large repository.

How the stack is wired

There is no application server in the compose file beyond the gateway. ChromaDB runs as chromadb/chroma:0.5.23 on an internal network, exposed only on port 8000 inside the Docker network, with IS_PERSISTENT=TRUE and ANONYMIZED_TELEMETRY=FALSE. Ollama sits beside it on port 11434, also internal. An nginx:alpine container is the only thing with a published port, mapped from ${PORT:-8080} to port 80, and it serves the static frontend from ./frontend while proxying to the API. That means the browser, the OpenAPI docs at /docs and the health endpoint at /health all arrive through one origin.

The backend is FastAPI, and the API surface is explicit: collections, document upload, folder ingestion, per-collection stats, query, and a keyword-only search endpoint. The split between /api/v1/rag/query and /api/v1/rag/search is the interesting design decision. Query runs the full retrieval plus generation path and returns an answer with sources. Search skips the LLM entirely and returns matches immediately. When you are debugging why a document is not being found, hitting search first tells you whether the problem is retrieval or generation, which is a distinction many RAG stacks make you guess at.

A second FastAPI surface, /api/v1/config, reads and writes the .env values through a ConfigService, and /api/v1/config/test validates downstream connections before you commit a change. That is what backs the Agent Configuration tab. Writing configuration from the UI into a .env file is convenient and also the part I would watch: the README does not describe what happens to container environment variables that were already read at startup, so a provider change made in the browser may not reach a running process until the containers are recreated.

Installing Knowledge Base Self-Hosting Kit and running a first folder ingest

The README gives a five step quick start. Clone the repository, copy the example environment file, bring the stack up, then check the health endpoint. The .env.example ships with sensible defaults, so the only value most people will want to change before the first run is DOCS_DIR, the host folder that gets mounted for folder ingestion.

bash
git clone https://github.com/2dogsandanerd/Knowledge-Base-Self-Hosting-Kit.git
cd Knowledge-Base-Self-Hosting-Kit
cp .env.example .env
docker compose up -d
curl http://localhost:8080/health

The first four commands clone, enter the directory, create your local .env and start the three containers detached. The curl at the end is the check the README specifies; a healthy backend answers on port 8080. If that call fails, the gateway container is the one to inspect, because it is the only service with a published port and the only route to the API.

Once the health check passes, open http://localhost:8080 in a browser. The README points to four entry points behind that single origin: the web UI with the Agent Configuration tab at the root, the OpenAPI docs at /docs, the health endpoint, and the API root at /api/v1/rag. From the UI you can create a collection, upload PDF, MD or TXT files, or point the ingest-folder endpoint at the mounted DOCS_DIR. The README notes that the host path is mounted at /host_root inside the container, so a folder scan targets that path, not the host path you typed into .env.

For the connector, the README gives one command for OpenClaw and a three step build for the server itself.

bash
openclaw mcp add --transport stdio knowledge-kit npx -y @knowledge-kit/mcp-server
cd mcp-server
npm install
npm run build
npm start

The first line registers the published connector with an MCP-aware agent over stdio. The remaining three build and start the connector from the checked-out source in mcp-server/, which is the route to take if you want to read what the connector sends before an agent does.

What the folder ingestion path actually does

Folder ingestion is the feature that separates this from a chat-with-a-PDF toy. The .env.example sets DOCS_DIR to ./data/docs and describes it as the local folder to mount into the container for ingestion. The README states that this host path is mounted at /host_root, and the ingest-folder endpoint kicks off a scan of it. That path translation is the thing to get right: if you point the API at the host path instead of /host_root, the container will not see your files, and the failure will look like an empty collection rather than an error.

The ingestion pipeline itself is where the code and prose claim has to hold up. Chunking is configured by CHUNK_SIZE and CHUNK_OVERLAP, defaulting to 512 and 128, with INGEST_BATCH_SIZE at 10. Docling handles the parsing, which is what lets PDFs and Markdown and plain text enter the same pipeline. The README lists PDF, MD and TXT as the upload formats for the document endpoint.

One constraint worth flagging comes from the environment comments rather than the feature list. The context window is capped at 8192 tokens in code, explicitly to prevent out of memory failures, and the example file recommends llama3.2 (3b) or llama3 (8b) for 8GB of VRAM. That is a real ceiling. If your question requires synthesising five long documents, the generator will be working with a truncated context and the citations may not cover everything the answer draws on. The embedding side has its own cost: nomic-embed-text is the default, and the README documents that collections are created with embedding metadata, which is what makes a later model swap a re-ingestion rather than a config edit.

Connecting an agent through the MCP server

The companion connector lives in mcp-server/ and is published as @knowledge-kit/mcp-server. The README gives the OpenClaw invocation directly, and the Agent Configuration tab exists to surface the same values in copyable form so you do not have to edit .env by hand. The tab also lets you set KNOWLEDGE_BASE_API_URL, timeouts and log levels for the MCP server, and test connectivity before handing the configuration to an agent.

This is the part of the project with the least verifiable surface. The MCP server is a Node package built from a subdirectory, and the README does not state a version, a registry, or a Node version requirement. The repository has no releases, so there is no published changelog for the connector independent of the 1.2.0 badge in the README. If your agent workflow depends on this connector, the responsible move is to build it from the checked-out source with the documented npm install, npm run build, npm start sequence rather than trusting an npm resolution you have not inspected. The smithery/ directory and smithery.yaml suggest a second deployment route, but the README only says they exist for running the MCP server inside Smithery.

The design intent behind the config API is sound: change the connector's target URL and log level from a browser tab, test the connection, then hand the values to the agent. The gap is persistence. The README says configuration persists in .env while the UI edits it through /api/v1/config, and it does not say how a running container picks up the new value. Plan on a docker compose restart after any change you make through that tab, and treat the test endpoint as a check on the downstream service rather than proof that the running server is using your new settings.

Where this project is the wrong tool

The maintenance situation is the first honest limitation. The last push to the default branch was on 2026-03-14, which is more than six months before today. There are no retrieved releases. The README itself states that the author is "out of this game" and will not publish more open source, and the top of the file is an authorship dispute naming another GitHub account. Whatever you think of that dispute, it changes the calculus for adoption: you are not choosing a project with a roadmap, you are choosing a snapshot.

That has concrete consequences. The compose file pins chromadb/chroma:0.5.23 while pulling ollama/ollama:latest, so one dependency is frozen and the other moves under you. A Chroma upgrade that changes the persistence format will not be handled by this repository. The license situation is also unresolved: the README badge says MIT and links to the MIT text, but the repository's license is reported as NOASSERTION, which usually means GitHub could not match the LICENSE file to a known template. Read the LICENSE file yourself before you depend on it.

Beyond maintenance, the stack is wrong for you if you need multi-tenant isolation, audit trails, or access control. Nothing in the documented API surface suggests per-user permissions on collections. It is also wrong if you want a managed vector store with backups and replication; ChromaDB here is a single container with a named volume, and the README does not document a backup or restore procedure. And it is wrong if your corpus is large enough that a full re-ingest is expensive, because the embedding metadata recorded at collection creation is what forces that re-ingest whenever you change EMBEDDING_MODEL.

How it compares to wiring the pieces yourself

The obvious alternative is not another product but the assembly: run ChromaDB, run Ollama, and write the FastAPI layer yourself. The difference in approach is that this project has already made the decisions you would otherwise make one at a time. It chose Docling for parsing, chose hybrid vector plus BM25 retrieval with score fusion, chose a single nginx gateway so the UI and API share an origin, and chose to expose configuration over HTTP so the UI can edit it. Rebuilding that is a week of work, and the parts you would get wrong are the parts that are already written down here.

The trade-off runs the other way too. A hand-rolled stack lets you upgrade ChromaDB on your own schedule and swap the parser without touching a compose file you did not write. Here, changing the embedding model means changing EMBEDDING_MODEL and re-ingesting, because the collection metadata records embedding information at creation time. The README documents that collections are created with embedding metadata but does not describe a migration path for existing collections when that model changes. That is a real operational gap, and it is the kind of gap that only shows up after you have loaded a few thousand documents.

There is a third option worth naming: skip the vector store entirely and put your documentation into a wiki or a static site with full-text search. That answers "where is this documented" but not "answer this question with citations", which is the whole point of the query endpoint. If your team's actual problem is findability rather than synthesis, this stack is more machinery than the problem requires.

If your priority is a stack you can read in an afternoon and adapt, the pre-assembled version wins. If your priority is a dependency you can patch on your own cadence, the assembly wins, and this repository becomes a reference for the architecture rather than the thing you deploy.

Licence and upgrade cost

The README carries an MIT badge and links to opensource.org/licenses/MIT, while the repository license field resolves to NOASSERTION. Those two facts are in tension and the README does not resolve it. If you need MIT terms in writing, open the LICENSE file at the repository root and read it; do not rely on the badge. Nothing here is legal advice, and I have not examined the file's contents.

Upgrade cost is dominated by two pinned decisions. First, the ChromaDB image tag is fixed at 0.5.23, so moving forward means testing a new tag against your existing chroma_data volume, and the README does not document a migration procedure. Second, the embedding model is recorded in collection metadata, so a change to EMBEDDING_MODEL implies re-ingestion rather than a config flip. The CHUNK_SIZE and CHUNK_OVERLAP values have the same property: they affect how documents were split when they were ingested, so changing them does not retroactively re-chunk anything. Budget the re-ingest, not just the edit.

The cheap upgrades are the ones the UI already supports. LLM_PROVIDER and LLM_MODEL can move to an OpenAI-compatible endpoint by setting LLM_PROVIDER=openai_compatible and pointing OPENAI_BASE_URL at a local server, and the config test endpoint exists to validate that connection before you rely on it. That path also sidesteps the 8192 token context cap only in the sense that you choose the server; the cap is described as living in code, so the new backend inherits the same limit unless you change it.

What to check before you deploy it

Start with the two files that decide everything: docker-compose.yml and .env. Confirm the Chroma tag matches what your existing data was written with, and confirm DOCS_DIR points at a folder you are willing to have parsed and embedded. If that folder contains secrets or credentials, remember that everything in it becomes retrievable text in the vector store, and the documented API surface has no per-collection access control to keep it out of an agent's reach.

Then run the health check and the OpenAPI docs before you load anything. The docs at /docs enumerate the endpoints the running build actually exposes, which is a faster way to spot a gap between the README and the code than reading either one alone. After that, ingest a small sample and query it with both /api/v1/rag/search and /api/v1/rag/query. If search finds the passage and query does not use it, the problem is in generation or context assembly, not in retrieval, and you have narrowed the fault to one layer in a single step. The examples/ directory ships ingest_my_code.py and sample_knowledge.md, which is a reasonable first corpus for exactly that test.

Editorial conclusion

Adopt it if you want a local ChromaDB plus Docling ingestion stack that you can read end to end, and if you are comfortable pinning the image tags yourself. Do not adopt it if you need a maintained dependency with releases, or if you are building a product on top of the MCP connector, since the README states the author has stopped publishing open source work. Before you commit, verify that the pinned chromadb/chroma:0.5.23 image is the version your existing collections use, and confirm whether the npm package @knowledge-kit/mcp-server actually resolves from your registry.

Frequently asked questions

What is a knowledge base used for in Knowledge Base Self-Hosting Kit?

It stores ingested documents as embeddings plus keyword index entries so you can query them and get answers with citation data. The README describes it as a self-hosted RAG memory layer that pairs with local LLMs, chains and autonomous agents.

Is self-hosting Knowledge Base Self-Hosting Kit legal?

The README shows an MIT licence badge and links to the MIT text, so the code is presented as free to run yourself. The repository's licence field is reported as NOASSERTION, so read the LICENSE file at the repository root rather than relying on the badge.

Is Knowledge Base Self-Hosting Kit free?

The README presents it under MIT terms, which permits self-hosted use at no cost. There are no retrieved releases and no pricing information anywhere in the repository, so the licence file is the only statement of terms available.

How do I create my own knowledge base with Knowledge Base Self-Hosting Kit?

Clone the repository, copy .env.example to .env, set DOCS_DIR to the folder you want to ingest, then run docker compose up -d. Once the health check at http://localhost:8080/health responds, open the UI at http://localhost:8080 and use the ingestion tooling or the ingest-folder endpoint.

Official sources

  1. 2dogsandanerd/Knowledge-Base-Self-Hosting-Kit on GitHub
  2. Issues
  3. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/2dogsandanerd-knowledge-base-self-hosting-kit.svg)](https://hysenlabs.com/projects/2dogsandanerd-knowledge-base-self-hosting-kit)