Model or dataset
VectifyAI/PageIndex avatar
VectifyAI/PageIndex

PageIndex: a vectorless, reasoning-based RAG engine that indexes documents as trees

PageIndex: Document Index for Vectorless, Reasoning-based RAG. Inspired by AlphaGo, we propose **PageIndex**, a **vectorless**, **reasoning-based RAG** system that builds a **hierarchical tree index** from long documents, and uses LLMs to **reason** *over that index* for **agentic, context-aware retrieval**.

35,850 stars3,158 forksPythonMIT

At a glance

What is it?
PageIndex replaces the vector index with a hierarchical tree and lets an LLM reason over it, which makes retrieval traceable but shifts the cost to the query model. Here is how the SDK works, where it breaks, and who should adopt it.
Who is it for?
Adopt PageIndex if your corpus is long professional documents where an answer must point back to a page, and you are willing to pay for a strong chat model on every query. Skip it if you need sub-second retrieval over millions of short, loosely related texts, or if you cannot send document content to a hosted model.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 5 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.

DEEP OPEN-SOURCE ANALYSIS

The problem PageIndex targets: similarity is not relevance

Vector RAG retrieves by embedding similarity. The README puts the objection plainly: "similarity ≠ relevance". On a 200-page regulatory filing, the passage that answers a question often shares little vocabulary with the question itself, while a semantically close chunk may be a boilerplate disclaimer. PageIndex is built for that gap.

The intended audience is narrow and specific. The README lists financial reports, legal documents, regulatory filings, technical manuals, medical literature, and academic textbooks. What those share is length, internal structure, and a reader who needs to know which page a claim came from. If your corpus is short support articles or product descriptions, the problem PageIndex solves is not your problem.

Tree index instead of chunks, LLM reasoning instead of nearest neighbours

PageIndex splits retrieval into two stages. Indexing produces a hierarchical tree per document, where the leaves are natural sections rather than fixed-size chunks. Retrieval is an agentic walk of that tree: the chat model inspects nodes, decides which branch to descend, and reads the section it lands on.

The repository layout reflects this. The package is pageindex/, there is a tests/ directory, a cookbook/, and run_pageindex.py at the top level. The index model and the chat model are separate parameters, and the README is explicit that they should be chosen differently: the tree structure comes from the document's own layout information, so the index model mainly summarizes and refines it, which a basic model handles. The chat model does the searching, so the README says to use the best model you can afford.

PageIndex Flash, announced in the updates, generates tree structure from PDFs in seconds by extracting structure heuristically from layout rather than building it with an LLM. That is the same design decision stated as a feature: structure extraction is not a reasoning task, so it does not need a reasoning model. The comparison table in the README frames the result as traceable to explicit references, against what it calls opaque "vibe retrieval".

Installing PageIndex and running a first query locally

The SDK installs from PyPI. The README gives `pip install -U pageindex`, and notes that this now ships a local mode where you index, retrieve, and chat entirely on your machine with your own LLM key. Python 3.10 or newer is required, according to pyproject.toml.

bash
pip install -U pageindex

The quickstart then sets an OpenAI key in the environment and constructs a client. Note that the quickstart uses the short spellings `index=` and `chat=`, while the usage guide uses `index_model=` and `chat_model=`; the README states that either spelling works.

python
import os
from pageindex import PageIndexClient

os.environ["OPENAI_API_KEY"] = "your-openai-key"

client = PageIndexClient(
    index_model="gpt-5.6-luna",
    chat_model="gpt-5.6-sol",
    storage_path=".pageindex",
)
doc_id = client.submit_document("report.pdf")["doc_id"]
answer = client.chat("What was the 2023 operating margin, and where is it stated?", doc_id=doc_id)
print(answer)

`submit_document` returns a dict containing `doc_id`, which you pass to `chat`. `storage_path` is where indexed documents are kept locally. The model names follow LiteLLM's naming convention, so the provider prefix pattern in the README is what you change when you point the client at a different provider.

To get page-level citations, the README passes a system message alongside the question instructing the model to cite only statements supported by tool outputs, using a `<cite doc="{docName}" page="{pageNumber}"/>` tag. The model fills in the name and page, producing output like `<cite doc="report.pdf" page="12"/>`.

Where the design costs you: query price, layout dependence, and the missing rollback story

The clearest limitation is stated by the project itself. The README says to use the best chat model you can afford, because the chat model searches the tree. In vector RAG, the expensive model is optional; here it is on the critical path of every question. A cheap embedding lookup becomes a multi-step LLM traversal. The README links a section titled "Query cost and accuracy" from the model recommendations, which tells you the project treats this as a known trade-off rather than a detail.

The second constraint follows from Flash. If tree structure is extracted heuristically from the document's own layout, then documents without reliable layout information are the weak case. Scanned pages, transcripts, and plain text without headings give the heuristic less to work with. The README does not document how the pipeline degrades on those inputs.

Third, the README does not document rollback or re-indexing behaviour when a document changes, and it does not describe what happens to an existing tree when the index model or the layout extractor is upgraded. The pyproject.toml classifier still reads "Development Status :: 3 - Alpha", which is worth weighing against the version number. The last push to the repository was on 2026-08-25, and v0.2.11 was released the same day.

Finally, local mode means your own LLM key and your document content going to whichever provider that key belongs to. The README does not describe a fully offline path with a self-hosted model.

PageIndex versus vector RAG: what actually differs in approach

The honest comparison is not "tree versus vectors" as data structures. It is where the intelligence sits. A vector pipeline spends its budget at write time: embed every chunk once, then answer queries with a cheap nearest-neighbour lookup. PageIndex spends its budget at read time: the tree is cheap to build, and each question pays for a reasoning walk.

That inverts the economics. Vector RAG amortizes cost across many queries and gets cheaper per query as volume grows. PageIndex gets more expensive per query as questions get harder, because harder questions mean more traversal. For a workload of a few hundred questions a day against a stable corpus, that is fine. For a workload of millions of short lookups, it is the wrong shape entirely.

The second difference is auditability. A vector result gives you a chunk and a similarity score, which is hard to explain to a reviewer. A tree walk gives you the path the model took and, with the citation system message, a page number. In regulated review work, that difference is often the whole reason to switch.

Licence and upgrade cost

PageIndex is MIT licensed, per both the LICENSE file and the pyproject.toml `license = "MIT"` field. MIT permits commercial use and modification with attribution and no warranty. That covers the SDK code. It does not cover the model you point it at, and it does not cover the hosted PageIndex Cloud service, which the README describes as an alternative endpoint for the same client using an API key. Those are separate terms you would need to read on their own.

Upgrade cost is dominated by the dependency list, not by the SDK surface. pyproject.toml pins minimum versions with comments explaining why: `openai-agents = ">=0.18.1"` because older releases crash on current openai before the request is sent, `urllib3 = ">=1.26"` for MCP retries, and `claude-agent-sdk = ">=0.1.53"` because older releases break string prompts with SDK MCP servers. The Anthropic and Claude extras are optional; the `openai` extra exists and is deliberately empty so that `pip install "pageindex[openai]"` stays valid. If you pin your own LLM SDK versions for other parts of your stack, expect to reconcile them here.

Who this is for

PageIndex is for teams whose retrieval failures are traceable to a reasoning gap, not a latency gap. Analyst tooling over filings, contract review, technical documentation search where the answer spans sections: these are the cases the README names, and they are the cases where a page-level citation is worth more than a millisecond.

It is not for high-volume, low-stakes lookup, and it is not for teams that cannot send document text to a hosted model. The SDK is Python only, so a non-Python service would be calling it through a sidecar or a subprocess. If you are already running a vector store and your retrieval quality is adequate, the migration buys you traceability and costs you query budget; that is a real trade, not an upgrade.

Editorial conclusion

Adopt PageIndex if your corpus is long professional documents where an answer must point back to a page, and you are willing to pay for a strong chat model on every query. Skip it if you need sub-second retrieval over millions of short, loosely related texts, or if you cannot send document content to a hosted model. Verify three things before committing: that your documents have usable layout structure, that your chat model follows the citation system message, and what your per-query token bill looks like at the volume you expect.

Frequently asked questions

What are the key differences between PageIndex and a vector database?

PageIndex uses a hierarchical tree index built from a document's natural sections, while a vector database stores fixed-size chunks as embeddings. Retrieval differs too: PageIndex has an LLM reason over the tree, whereas vector RAG does a semantic similarity search. The README frames the result as traceable to explicit references rather than opaque similarity matches.

How does PageIndex work in RAG?

It runs in two steps. First it generates a tree-structure index for each document, then it agentically searches that tree using LLM reasoning. The index model summarizes and refines the tree, and the chat model performs the search.

Where can I find my PageIndex API key?

The README does not document where to obtain a PageIndex Cloud API key. It shows that the same SDK client can be pointed at PageIndex Cloud with an API key, and links to pageindex.ai/developer for the API and MCP, but the key retrieval flow is not described in the repository files.

How do I use PageIndex locally?

Install with pip install -U pageindex, then construct a PageIndexClient with your own LLM key set in the environment and a storage_path for local storage. The README states that local mode lets you index, retrieve, and chat entirely on your machine. Document content still goes to whichever model provider your key belongs to.

Is PageIndex open source?

Yes. The repository is VectifyAI/PageIndex and it is MIT licensed, according to both the LICENSE file and the license field in pyproject.toml. The hosted PageIndex Cloud service is a separate offering with its own terms.

Is PageIndex better than vector RAG?

The README reports 98.7% accuracy on FinanceBench for financial document QA, which it says outperforms vector-based RAG, and argues that similarity is not relevance. That claim is scoped to professional documents needing contextual understanding. The README also says to use the best chat model you can afford, so the accuracy comes with a per-query cost that a vector index does not carry.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
  4. Release notes
For maintainers

Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/vectifyai-pageindex.svg)](https://hysenlabs.com/projects/vectifyai-pageindex)
Community notes

Community notes