Radar
← All projectsTuesday, September 29, 2026
PageIndex
Vectorless RAG that builds a tree over long PDFs and lets the model reason to the right section — no chunking, no vector DB.
Why it matters
PageIndex indexes a document as a hierarchical tree instead of embedding chunks into a vector store. At query time an LLM walks that tree the way a person flips through a report — section by section — so retrieval is driven by relevance and context, not cosine similarity. Local SDK mode runs on your machine with your own model keys; Cloud handles OCR-heavy and scanned docs. It is aimed at finance filings, legal packs, manuals, and other long professional PDFs where naive RAG usually fails.
If your product answers questions over long PDFs, vector RAG still misses the section that matters and returns lookalike paragraphs. PageIndex is a different retrieval job: structure first, then reason. That is the unlock founders need for reliable doc agents without babysitting chunk sizes.
How it works
Install with pip install pageindex, point PageIndexClient at index and chat models, submit a PDF, then chat against the doc id. Indexing is cheap and one-time; later questions reuse the tree. Same client can talk to PageIndex Cloud with an API key, and you can drop the tools into OpenAI or Claude agent SDKs.
Most RAG stacks still mean embeddings plus a vector database. PageIndex skips both and retrieves by LLM reasoning over a document tree, with citations you can actually audit.
Capabilities
- API / SDK surface
- MCP server / client support