Knowledge retrieval

PageIndex vs vector RAG: choosing retrieval for long PDFs

Compare PageIndex’s document-tree retrieval with a realistic vector RAG pipeline. Check benchmark scope, citations and total costs before choosing for long PDFs.

By Clairevue · · 4 min read

An open manual beside connected section cards and a magnifying glass over individual passages.
AI-generated illustration of document-tree and passage retrieval; not a PageIndex interface or benchmark result.

Suppose a supplier’s manual describes a general warranty in one section and an exception for missed maintenance in another. An AI assistant that finds only the warranty can give a plausible answer while missing the clause that changes it.

PageIndex offers an open-source approach that uses a document’s hierarchy to guide a model towards relevant pages. A vector-based system searches representations of passages by similarity. Both need to deliver the evidence that answers the question.

For a small team building a manuals or policy assistant, PageIndex deserves a trial when section structure helps locate that evidence. Its “vectorless” label, on its own, gives you no reason to replace a working search pipeline.

What the tree changes

Retrieval-augmented generation, or RAG, supplies a model with material to use when answering. The retrieval stage decides which material reaches it.

PageIndex organizes a document into a tree of sections. Its Flash indexer reads the layout of text-based PDFs and uses a model for summaries. The tree records section titles and page ranges; at question time, a model uses that structure to find and read relevant content.

In the fictional warranty example, the model could inspect the warranty section and follow a reference to maintenance requirements. Preserving those sections gives it a way to search beyond a passage whose wording resembles the question. But it can also choose an unhelpful section or stop before finding the exception.

PageIndex describes this as avoiding chunking. It still divides the document into sections with bounded page ranges. The model searches those sections through the hierarchy; it still works within a limited context window.

A vector pipeline commonly splits text into passages, converts them into numerical representations called embeddings, and retrieves nearby representations for a query. If a decisive exception never enters the candidate set, a later answer model has nothing to cite from it.

Azure AI Search’s hybrid documentation shows a pipeline combining keyword and vector search, with optional semantic reranking and metadata filters. Exact product codes can benefit from keyword matching; filters can restrict results to the applicable manual version. Compare PageIndex with the pipeline you would actually deploy, not an intentionally bare similarity search.

The benchmark needs its small print

PageIndex’s local open-source benchmark contains 62 lookup questions over 34 PDFs, with answers stated in running text. It excludes charts, tables and calculations, as well as documents the local indexer refuses to process.

The team reports 53 correct answers out of 62, or 85.5%, for the gpt-5.6-luna answering model with the reasoning-effort setting none, and 60 out of 62, or 96.8%, with high. Both use the same saved trees. The company uses a model judge to compare responses with reference answers; we haven't rerun the evaluation.

The result shows that the answering configuration changes outcomes even when the index stays fixed. It doesn’t establish how well the system handles a scanned contract, a table-heavy report or an unsupported file. Nor does that table contain a vector baseline that would establish a winner between architectures.

PageIndex’s headline FinanceBench claim comes from Mafin 2.5, a system built on PageIndex. Its result doesn’t automatically describe the local SDK configuration you install.

In the original FinanceBench study, the researchers’ 2023 experiments still produced wrong answers when models received the evidence pages. Finding the right page and interpreting it correctly are different tasks; an answer score combines them.

Run a comparison that can change your decision

Start with authorized copies of documents your staff really consult, including awkward files you expect the assistant to handle. Write the questions and identify their supporting passages before looking at either system’s answers. Include the warranty exception, if your documents contain one, and questions the files cannot answer.

For each question, record whether retrieval found all the necessary evidence. Then check whether the answer follows it. Keep a separate record of unsupported claims and appropriate refusals, so a fluent wrong answer doesn’t count as success.

Citations help reviewers inspect the result. PageIndex’s chat documentation supports requested page-level citations locally and block-level citations in cloud mode. A page reference still needs checking against the claim it accompanies. Vector pipelines can also retain source locations; citation validity belongs in both evaluations.

Where practical, use the same answering model and instructions, while documenting the retrieval budgets and any architecture-specific changes. Keep some questions aside until you’ve finished tuning. If an approach refuses a required document, record that failure instead of quietly removing the file.

Compare complete response times and total costs. Include document parsing and initial indexing, plus the cost of rebuilding when a manual changes. For repeated questions, count retrieval calls and answer generation together; the PageIndex benchmark’s per-question cost excludes its shared indexing work. Record caching conditions too.

Before using confidential documents, check the data flow. The PageIndex client reference separates local index storage from the model used for indexing and chat. Choosing local mode while configuring a remote model API doesn’t make inference local. Verify which document content reaches that provider and under what terms.

Ask both systems to point to the clause that changes the answer in the applicable document version. When that evidence is missing, expect the assistant to say it can’t answer from the files.