PageIndex SDK Goes Local PageIndex SDK now supports fully local reasoning-based RAG workflows, allowing users to build tree indexes, store documents, run retrieval, and chat with long documents on their own machines using their preferred LLM and API key. The update, announced by PageIndex, removes the vector database from the retrieval path entirely, replacing it with a hierarchical tree index and LLM-based tree search that produces page-level citations. Local mode is available via `pip install -U pageindex` and is designed for text-heavy PDFs and private development workflows. Today, we are announcing a major update to the PageIndex SDK : the complete reasoning-based RAG workflow can now run locally. PageIndex SDK users can now build tree indexes, store documents, run reasoning-based retrieval, and chat with long documents on their own machines. The local workflow uses your preferred LLM and API key, produces page-level citations, and works with the same agent frameworks and model APIs already supported by the SDK. This is not a separate local product or a new client. It is a major expansion of the existing PageIndex SDK: choose local or cloud indexing when you initialize PageIndexClient , while keeping the rest of your application workflow consistent. Install it with one command: pip install -U pageindex Local mode is designed for text-heavy PDFs and private development workflows. It uses your own model API key, stores indexes on your machine, and requires no vector database. Why PageIndex? Most RAG systems split a document into fixed-size chunks and retrieve them by vector similarity. This approach is useful, but similarity is not the same as relevance. In long professional documents, the passage that answers a question may use completely different language from the query. A semantically similar passage may also be nearby in meaning while being irrelevant to the actual task. Financial reports, regulations, technical manuals, and textbooks often require context, domain knowledge, and multi-step reasoning to identify the right evidence. PageIndex takes a different approach. It organizes each document as a hierarchical tree index , then lets an LLM reason through that tree in the same way a human reader uses a table of contents and section structure to find the right pages. Retrieval happens in two stages: Index: Build a tree that preserves the document's natural structure. Retrieve: Use LLM-based tree search to identify and read the relevant sections.