OpenKB: an LLM knowledge base that compiles documents into a wiki instead of a vector index
OpenKB: Open LLM Knowledge Base
At a glance
- What is it?
- OpenKB is a Python CLI that turns PDFs, Office files, URLs and directories into an interlinked Markdown wiki using LiteLLM models and PageIndex tree indexing. It is a good fit if you want persistent, Obsidian-readable knowledge; it is the wrong tool if you need a hosted service or a stable 1.x API.
- Who is it for?
- Adopt OpenKB if you already keep documents on disk, are comfortable editing .openkb/config.yaml, and want a wiki you can open in Obsidian. Skip it if you need a hosted service with an SLA, or if your corpus is short notes where a plain-text grep already answers the question.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 70 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 26, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem OpenKB targets: retrieval that forgets everything between queries
OpenKB's README states the design premise directly: traditional RAG rediscovers knowledge from scratch on every query, and nothing accumulates. The project positions itself against that by compiling documents once into a persistent wiki, then keeping it current. Cross-references are computed at compile time rather than at query time, and the README claims contradictions are flagged and synthesis reflects everything consumed. The stated inspiration is a concept described by Andrej Karpathy, where an LLM generates summaries, concept pages and cross-references that are maintained automatically so knowledge compounds.
The audience is narrower than the tagline suggests. This is a CLI tool for developers and researchers who already have a document pile (papers, specs, internal docs) and want a browsable artifact from it. The wiki is plain .md files with cross-links, so it opens in Obsidian for graph view. If your goal is a chatbot over a handful of PDFs, the compile step is overhead. If your goal is a knowledge base you will read, edit and extend over months, the compile step is the product.
How OpenKB works: markitdown for short documents, PageIndex trees for long PDFs
The architecture has two layers according to the README: a wiki foundation that compiles and maintains knowledge, and generators (query, chat, Skill Factory) that turn it into output. Ingestion branches on document length. Short documents go through markitdown into Markdown, images are extracted inline with pymupdf, and the LLM reads the full text. Long documents, defined in the README's table as PDFs of 20 pages or more, go through PageIndex instead, which produces a tree index plus summaries, and the LLM reads document trees rather than raw text. Images in that path are extracted by PageIndex.
Both paths converge on the same output: a summary plus concept pages. Entity pages for people, organizations, places and products are auto-extracted and kept in sync, and wiki pages follow the Google Open Knowledge Format specification for knowledge sharing. Retrieval is described as vectorless and reasoning-based, which is the substantive difference from a conventional embedding pipeline: there is no vector database in the loop, so there is no embedding model to choose, no chunk-size parameter to tune, and no index to rebuild when the embedding model changes. The trade-off is that retrieval quality now depends on the tree structure PageIndex builds and on the reasoning model you point at it.
Installing OpenKB and running a first query
The README gives a single install command from PyPI. Python 3.10 or newer is required per pyproject.toml, and the package is published as openkb.
pip install openkbThe README also lists two alternatives: pip install git+https://github.com/VectifyAI/OpenKB.git for the latest from GitHub, and a source install with git clone followed by pip install -e . for development.
Before anything works you need a model and a key. OpenKB routes through LiteLLM, so providers are addressed in provider/model form, for example anthropic/claude-sonnet-4-6; OpenAI models can omit the prefix, for example gpt-5.4. You set the model during openkb init or later in .openkb/config.yaml, and the key goes in a .env file:
LLM_API_KEY=your_llm_api_keySubscription providers that authenticate through an OAuth device flow, such as chatgpt/* or github_copilot/*, do not need an API key, and OpenKB skips the missing-key warning for them.
The README's quick start then creates a knowledge base, adds documents from a file, a directory or a URL, and asks a question:
mkdir my-kb && cd my-kb
openkb init
openkb add paper.pdf
openkb add ~/papers/
openkb add https://arxiv.org/pdf/2509.11420
openkb query "What are the main findings?"Expect the add step to be the slow one, since it is where compilation happens. The README also shows openkb chat for a multi-turn session with persisted sessions you can resume, and optional generators: openkb skill new for a portable agent skill, openkb visualize for an interactive knowledge graph, and openkb deck new for a single-file HTML slide deck.
There is also a bundled web interface. Install the extra and start the server:
pip install "openkb[web]"
openkb-webThe README says this serves the API and the Knowledge Workbench at http://127.0.0.1:7566/. Auth is off by default because the design is local-first; set OPENKB_API_TOKEN to require a bearer token before exposing the server. The REST API variables in .env.example are OPENKB_API_TOKEN and OPENKB_KB_ROOT, the latter controlling where REST /init creates knowledge bases, defaulting to ~/.config/openkb/kbs.
Where OpenKB breaks down or is the wrong choice
The most concrete limitation is stated in pyproject.toml rather than the README: every dependency is pinned exactly, and the file explains why, citing supply-chain caution and the litellm package-poisoning incident, with the instruction to bump deliberately after vetting each release. That is a defensible stance, but it means you inherit the maintainers' pinning decisions. The same file documents two live compatibility constraints: litellm 1.87.2 is the version that fixes the chatgpt/* provider returning empty Responses output and auto-injects GitHub Copilot IDE-auth headers, and openai is held below 2.45.0 because 2.45.0 added a required cache_write_tokens field that openai-agents 0.17.3 does not set, which crashes usage parsing on every response. If your environment forces a newer openai, you have a conflict to resolve yourself.
The project also classifies itself as Development Status :: 3 - Alpha in pyproject.toml. Treat the CLI surface and the config format as moving. The README does not document rollback for a bad compile, and it does not describe how to remove a document and its derived pages from the wiki once added, so plan on keeping the source documents and the wiki under version control.
The Web UI is explicitly local-first. Auth is off by default, and the README's guidance is to set OPENKB_API_TOKEN before exposing the server. If you need multi-user access control, per-user isolation or an audit trail, this is not that product. Finally, cost and latency scale with the corpus: compilation sends document text or document trees to your chosen provider, and long PDFs are read as trees by a reasoning model, so a large backfill is a metered operation rather than a one-time local index build.
OpenKB compared with a conventional vector-database RAG stack
The obvious alternative is the standard pipeline: chunk documents, embed them, store vectors in something like a vector database, and retrieve nearest neighbours at query time. The difference in approach is not cosmetic. In the vector stack, the artifact is an index that is rebuilt whenever the embedding model, chunker or metadata scheme changes, and the answer to a question is assembled fresh from retrieved chunks with no memory of previous questions. In OpenKB, the artifact is a set of Markdown files, and the expensive work happens at ingestion: summaries, concept pages, entity pages and cross-links are written once and then maintained. Retrieval is reasoning-based over PageIndex trees rather than similarity search over embeddings.
That shift buys persistence and readability at the cost of flexibility. You can open the wiki in Obsidian, diff it in git, or hand-edit a page, none of which applies to a vector index. What you give up is the mature tooling around embeddings: no off-the-shelf evaluation harnesses for retrieval quality, no ability to swap in a domain-tuned embedding model, and no straightforward way to serve low-latency lookups at high query volume, since a reasoning model sits in the retrieval path. If your workload is high-QPS semantic search over millions of short chunks, a vector database remains the better fit. If your workload is a slowly growing document collection you want to interrogate and read, the compile-once model is the stronger shape.
Licence, maintenance and the cost of upgrading
OpenKB is Apache-2.0, declared both in the repository LICENSE file and in pyproject.toml as license = {text = "Apache-2.0"}. That is a permissive licence, and it matters here for a specific reason: the Skill Factory distills redistributable agent skills from your wiki, so the licence on the tool is separate from whatever rights you hold over the documents you feed it and the skills you generate. Nothing in the repository changes the copyright status of your source material, and OpenKB does not claim to. This is not legal advice; if you plan to redistribute generated skills, check the terms of the underlying documents yourself.
On maintenance, the last push to the default branch was on 2026-07-22, and the most recent release listed is v0.4.5 from 2026-07-20, preceded by v0.4.4 and v0.4.4-rc1 on 2026-07-10. The repository is not archived. Releases are frequent enough that the version numbers move in the fourth position, which is consistent with the alpha classifier.
The upgrade cost is dominated by the exact pins. Because litellm, openai, openai-agents, pageindex, markitdown and the rest are pinned with ==, installing OpenKB into an environment that already has different versions of those packages will either fail to resolve or force a downgrade. The practical approach is a dedicated virtual environment per knowledge base, and reading the dependency block in pyproject.toml before any upgrade rather than after. The comments in that block are the closest thing to a changelog for the compatibility constraints.
Editorial conclusion
Adopt OpenKB if you already keep documents on disk, are comfortable editing .openkb/config.yaml, and want a wiki you can open in Obsidian. Skip it if you need a hosted service with an SLA, or if your corpus is short notes where a plain-text grep already answers the question. Before committing, verify that your chosen LiteLLM provider string works with your key, check the pinned dependency set in pyproject.toml against your own environment, and confirm what openkb init writes into .openkb/ in a throwaway directory.
Frequently asked questions
What is OpenKB?
OpenKB is an open-source CLI that compiles raw documents into a structured, interlinked wiki-style knowledge base using LLMs. It is powered by PageIndex's vectorless, reasoning-based retrieval for long documents, and the resulting wiki is plain Markdown with cross-links.
Does ChatGPT have a knowledge base?
The OpenKB README does not describe ChatGPT as having a knowledge base. It does note that OpenKB supports subscription-based providers that authenticate via OAuth device flow, such as chatgpt/* and github_copilot/*, which need no API key.
What is an example of a knowledge base built with OpenKB?
The README's quick start builds one from a directory: you run openkb init, add a PDF, a folder of papers, or a URL such as an arXiv PDF, and the LLM compiles summaries, concept pages and entity pages into the wiki. The result is a set of .md files with cross-links that opens in Obsidian.
Can you give me an example of knowledge-based AI?
OpenKB is one: the README describes it as compiling documents into a persistent wiki, then answering queries over that wiki with reasoning-based retrieval instead of rediscovering knowledge on every query. It also offers chat sessions and a Skill Factory that distills agent skills from the wiki.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/vectifyai-openkb)