Knowhere: parsing dirty documents into navigable memory for AI agents
Knowhere extracts, parses, and outputs structured chunks ready for AI Agents and RAG.
At a glance
- What is it?
- Ontos-AI's Knowhere is an Apache-2.0 document parsing and retrieval system that turns PDFs and slide decks into hierarchy-native chunks with citations. The interesting part is the dual-track design; the weak part is that the README never explains how to install the server yourself.
- Who is it for?
- Adopt Knowhere if your pipeline is stuck on messy PDFs and slide decks and you want one chunk schema shared by a vision path and a text path, with citations back to source pages. Do not adopt it if you need a documented self-host install path today, or if you want a retrieval engine that decides the search strategy for you: the project deliberately hands that decision to the agent.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem Knowhere targets: dirty documents as agent context
Most retrieval stacks assume the hard part is the vector index. Knowhere assumes the hard part is upstream of it. The README states the project's premise directly: traditional OCR and document intelligence pipelines try to extract every element before a model can understand the document, and on dirty PDFs and slide decks mistakes in reading order, layout, tables, or hidden text layers accumulate into unreliable model context.
That failure mode is familiar to anyone who has fed scanned filings or exported slide decks into a chunker. A table read column-wise, a two-column paper flattened into alternating lines, a text layer that duplicates the visible glyphs: each of these produces chunks that look fine in a count and are wrong in content. The downstream embedding step cannot repair them, because the damage happened before tokenization.
Knowhere's answer is to stop treating perfect element extraction as a precondition for retrieval. It is aimed at teams building agentic RAG over local or offline document collections, and at anyone who wants the output of parsing to carry its own provenance: every result, per the README, stays connected to its document, section, source pages, and related assets. The audience is narrower than "anyone doing RAG". If your corpus is clean HTML or well-formed Markdown, the machinery here is more than you need.
Two tracks, one schema: how the parsing pipeline is put together
Knowhere 2.0 runs two complementary parsing tracks. The Text Track preserves precise text and native structure where they are reliable. The Vision Track sends a page or slide to a frontier vision model and asks it to understand the page as a whole. The README is explicit that the tracks differ in how they understand the source, not in how agents consume the resulting memory.
The routing rule is stated in the README: PDF and .pptx uploads through the V2 Jobs API use the Vision Track, and other supported formats use the Text Track. That is a format-and-API-generation decision, not a per-document quality heuristic. It also means the Vision Track is reachable only through one API generation, which is a constraint worth noticing before you design around it.
Both paths then normalize into what the README calls a hierarchy-native chunk and metadata schema. The build step stores navigation trees, linked assets, citations, and cross-document relationships. The repository layout is consistent with that description: pyproject.toml declares a uv workspace with three members, apps/api, apps/worker, and packages/shared-python, so the request path and the heavy parsing path are separate processes sharing a common package. Type checking is configured across apps/api/app, apps/worker/app, and packages/shared-python/shared, which suggests the shared schema really is shared code rather than a convention.
The retrieval half is deliberately thin. The README describes Knowhere as providing the substrate: one corpus schema and tools for document outlines, structural filters, exact search, fuzzy recall, full reading, assets, and cross-document relationships. The agent decides how to search, traverse, read, and cite. In the September 8, 2026 release notes, the project describes this as MapNav evolving from a fixed navigation workflow into a corpus-native foundation, with the same foundation exposed to external agents through MCP. If you expected a retriever that picks the strategy for you, this is the opposite design.
Installing Knowhere and running the first parse
The README does not document a self-host installation. It points to a separate repository, Ontos-AI/knowhere-self-hosted, for self-deployment, and to knowhereto.ai for a managed API. What the main repository does document is the development toolchain. The project requires Python 3.11 or newer, and the Makefile drives uv.
To set up a working tree and run the checks the maintainers run, install uv, clone the repository, and use the Makefile targets. The lint target runs ruff over apps and packages, and check runs lint plus pyright:
make lint
make checkIf you want to run the document agent test suite specifically, the Makefile exposes a target that changes into the worker app and runs pytest against tests/document_agent:
make test-doc-agentThe pyproject.toml pins the interpreter floor and declares the workspace members, which is the part that matters if you intend to run the API and worker locally:
requires-python = ">=3.11"
[tool.uv.workspace]
members = [
"apps/api",
"apps/worker",
"packages/shared-python",
]For an actual first parse, the documented path is the V2 Jobs API, and the README states that PDF and .pptx uploads through that API use the Vision Track. The README does not give the endpoint, the request body, or the response shape, so you will have to read the API source or the docs site at docs.knowhereto.ai before writing a client. Container images are published to ghcr.io/ontos-ai/knowhere, which is the closest thing to a documented deployment artifact in the main README.
Where Knowhere is the wrong tool
The clearest limitation is documented rather than hidden: the README does not describe how to install and run the platform from this repository. Self-hosting is delegated to a second repository, and the managed API is the path the README promotes with a note about free credits. For an open source project, that is a real gap. You can read the parsing architecture, run the lint and type checks, and inspect the worker, but the README alone will not get a server answering requests.
The Vision Track has a cost profile the README does not quantify. It routes PDFs and pptx files through frontier vision models, page by page. The README mentions that the pipeline can process long-form PDFs with hundreds of pages, for example 300, 500, or more, and that technical atlases or drawing collections go through a dedicated layout-aware parser. Nothing in the README states latency, throughput, or token cost for that path, so plan for it as an unknown rather than assuming it behaves like a text extractor.
The agent-neutral retrieval design is also a limitation in disguise. Knowhere supplies tools and a schema; the agent decides how to explore. If your team wants a retrieval component with a fixed, tunable strategy that behaves the same way on every query, this is the wrong layer to adopt, because the variability moves into your agent's decision-making. And if your documents are already clean, structured text, the dual-track machinery buys you nothing that a straightforward chunker would not.
How this differs from a conventional document intelligence pipeline
The obvious comparison is a classic OCR plus layout-analysis pipeline, the kind that runs detection, recognition, and reading-order reconstruction before anything is indexed. The difference is where each system places its bet. A conventional pipeline bets that better extraction produces better context, so it invests in recovering every element correctly. Knowhere bets that extraction will sometimes fail and that retrieval should survive that failure, which is why the Vision Track indexes pages through summaries, entities, source text, and hierarchy even when OCR or layout extraction cannot reliably recover every component.
A second comparison is a plain vector RAG stack built on a text splitter. There, chunks are produced by character or token boundaries and carry whatever metadata the loader attaches. Knowhere instead emits hierarchy-native chunks that retain their document, section, source pages, and related assets, and it builds navigation trees and cross-document relationships at ingest time. The practical difference shows up when an agent needs to cite: with a plain splitter you reconstruct provenance after the fact, while here it is part of the stored record.
The third difference is the retrieval contract. LangChain-style retrievers, and Knowhere's topics list includes langchain, typically expose a retriever object that takes a query and returns documents. Knowhere exposes corpus tools and lets the agent choose among outlines, structural filters, exact search, fuzzy recall, and full reading. That is a different division of labor, and it explains why the project also lists MCP support: the same substrate is meant to be driven by an external agent rather than by an internal query planner.
Licence, release cadence and the cost of upgrading
Knowhere is licensed under Apache-2.0, with the LICENSE file at the repository root and a NOTICE file alongside it. Apache-2.0 permits commercial use and modification and includes an express patent grant. It also requires that you preserve copyright and licence notices and that you carry the NOTICE contents forward when you redistribute. If you embed Knowhere in a product, the practical obligation is attribution hygiene in your distribution, not a change to your own licensing. This is a description of the licence text, not legal advice; have counsel review your specific distribution model.
On cadence, the release history shows v1.2.2 on 2026-09-08, v1.2.3 on 2026-09-09, and v1.2.4 on 2026-09-10, with the last push to the default branch on 2026-09-10. Three patch releases inside three days suggests active work on the 1.2 line, and the September 8 news entry about Retrieval 2.0 and the September 2026 entry about Document Parsing 2.0 line up with that window. The repository is not archived.
That pace is the upgrade cost. Rapid patch releases on a parsing pipeline mean the chunk and metadata schema is the thing to watch, because the README's central promise is that both tracks converge on one contract. If that schema shifts between minor versions, your stored memory and your retrieval tools can drift apart. There is no migration guide or versioning policy for the schema in the README, so pin a version and read the diff between releases before moving.
Editorial conclusion
Adopt Knowhere if your pipeline is stuck on messy PDFs and slide decks and you want one chunk schema shared by a vision path and a text path, with citations back to source pages. Do not adopt it if you need a documented self-host install path today, or if you want a retrieval engine that decides the search strategy for you: the project deliberately hands that decision to the agent. Before committing, verify three things in the repository rather than the README: what apps/api actually exposes as endpoints, whether the V2 Jobs API is the only route that triggers the Vision Track, and what the self-hosted repository contains, since the main README links to it instead of documenting deployment inline.
Frequently asked questions
What does Knowhere mean?
The repository does not define the name. It is shared with a location in Marvel comics and films, which is what most searches for the term refer to, and that is a different subject from this project.
What is Knowhere?
Knowhere is a document parsing and retrieval system that ingests unstructured documents and produces persistent, navigable memory for AI agents. It runs parsing, hierarchy reconstruction, multi-modal structuring, and graph construction in a single pipeline, and the output stays connected to its document, section, source pages, and related assets.
Is Knowhere a Celestial?
That question is about the Marvel location, not this project. Knowhere the software is a Python document parsing and retrieval system licensed under Apache-2.0.
What is Knowhere made of?
The repository does not describe the fictional location. Knowhere the software is built as a uv workspace in Python with three members: apps/api, apps/worker, and packages/shared-python.
What is Knowhere in Marvel?
That question refers to the Marvel location, which is unrelated to this project. Knowhere here is an Apache-2.0 document parsing and retrieval system for AI agents, with a managed API at knowhereto.ai and a self-hosted repository at Ontos-AI/knowhere-self-hosted.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/ontos-ai-knowhere)