Introducing PageIndex Flash PageIndex has released PageIndex Flash, a fully open-sourced, local tree-indexing engine for text-based PDFs that builds hierarchical indexes from a PDF's own layout instead of using a vision model, and it is now the default indexer in the PageIndex SDK's local mode. The engine runs entirely on the user's machine, requires no vector database, and supports reasoning-based RAG workflows with page-level citations, making it suitable for private and regulated document workflows. PageIndex Flash is available now via `pip install -U pageindex`, and for scanned or image-heavy documents, PageIndex recommends its cloud service, PageIndex Cloud. Today we are introducing PageIndex Flash , a fast tree-indexing engine for long, text-based PDFs . PageIndex Flash runs entirely on your own machine — your documents never leave it — and it is available now in the PageIndex SDK. PageIndex Flash builds a hierarchical tree index by reading a PDF's own layout, rather than asking a vision model to infer the whole outline from scratch. That one change makes indexing fast, cheap, and predictable enough to run across every text-based document you hold. PageIndex Flash is fully open-sourced and is the default indexer in the SDK's local mode, which brings the complete reasoning-based RAG workflow to your machine. Build tree indexes, store documents, run reasoning-based retrieval, and chat with long documents using your preferred LLM and API key, with page-level citations and no vector database. Install it with one command: pip install -U pageindex PageIndex Flash runs locally by design. It uses your own model API key, keeps every index on your own disk, and requires no vector database, which makes it suitable for private and regulated document workflows. The local pipeline reads the PDF directly and runs no OCR, so it expects text-based PDFs . For scanned files and image-heavy documents, we recommend PageIndex Cloud https://developer.pageindex.ai/ , which runs OCR and image understanding before the tree is built. What is PageIndex? Most RAG systems split a document into fixed-size chunks and retrieve them by vector similarity. This approach is useful, but similarity is not the same as relevance. In long professional documents, the passage that answers a question may use completely different language from the query. A semantically similar passage may also be nearby in meaning while being irrelevant to the actual task. Financial reports, regulations, technical manuals, and textbooks often require context, domain knowledge, and multi-step reasoning to identify the right evidence. PageIndex takes a different approach. It organizes each document as a hierarchical tree index , then lets an LLM reason through that tree the way a human reader uses a table of contents and section structure to find the right pages. That splits retrieval into two stages — building the tree, then searching it: