VectifyAI/PageIndex
65.4
Adequate · 18 September 2026
25k
lines of production code
Python
primary language
1
measurement over time
What this system is
This system is a Python SDK for PageIndex that enables local and cloud-based document indexing, retrieval, and chat. It features a 'Flash' mode for fast, LLM-free PDF structure extraction and supports multimodal retrieval with block-level citations. The SDK integrates with major agent frameworks like OpenAI and Anthropic, allowing document Q&A within agent workflows via a unified client.
Features
Introduce PageIndex Flash for LLM-free PDF structure extraction
Adds the PageIndex Flash module, a new local indexing mode that builds a hierarchical tree structure from PDF layout statistics without requiring an LLM for the initial extraction. The \page\_index\_flash\ API accepts a PDF path or stream and returns a structured tree with optional LLM-generated node summaries and deterministic optimization (merge or full expand). It automatically detects table of contents sources (layout, embedded bookmarks, or hybrid), handles edge cases like scanned or unreadable PDFs, and includes a rejection policy to prevent indexing documents that lack a viable structure.
pageindex/flash · high confidence
New PageIndex documentation notebooks for citations, local indexing, and multimodal retrieval
The cookbook now includes three new Jupyter notebooks demonstrating key PageIndex capabilities. The 'Citations & Source Highlighting' notebook shows how to use block-level citations to trace answers back to specific regions in source documents. The 'Flash Demo' notebook illustrates building vectorless RAG using the local SDK and tree-indexing model. The 'Multimodal Retrieval & Understanding' notebook demonstrates retrieving and interpreting document images using a vision-capable chat model.
cookbook · high confidence
PageIndex SDK v0.2.x: Local mode, agent tools, and multi-provider chat
The PageIndex SDK now supports a local indexing and chat mode alongside the existing cloud surface, allowing users to index PDFs and chat against them using local models (configurable via \index\_model\ and \chat\_model\) without requiring a cloud API key. The SDK introduces a unified \PageIndexClient\ that routes to either cloud or local backends, and adds agent tool integrations for OpenAI Agents SDK, Anthropic SDK, and Claude Agent SDK, enabling document retrieval and Q&A within agent workflows. Chat functionality is expanded with streaming support (\ChatStream\), multiple protocols (\chat\_completions\, \messages\, \responses\), and LiteLLM integration for multi-provider LLM support, including model passthrough and backend connection overrides. The package also includes lazy loading for performance, improved error handling with \PageIndexAPIError\, and configuration via \config.yaml\ for indexing parameters.
pageindex · high confidence
Behavioural changes
PageIndex SDK introduces local mode with Flash indexing as the default
The PageIndex SDK now supports a local mode that allows users to index, retrieve, and chat entirely on their own machine using their own LLM keys, or connect to PageIndex Cloud. The primary indexing method has shifted to 'PageIndex Flash', a fast tree index generation process for text-based PDFs which is now the default in local mode. This change replaces the previous standard indexing approach and is accessible via the new \run\_pageindex.py\ CLI entry point, which defaults to Flash mode and includes options for optimization and summary generation.
(repo-wide) · high confidence
Test coverage
Added test fixtures and contract data for agent tools and flash layout; Initial test suite for PageIndex SDK and Flash indexing.
Dependencies
Update PageIndex SDK dependencies and add pyproject.toml
The PageIndex Python SDK (v0.2.10) now includes a pyproject.toml manifest and updates its core dependencies in requirements.txt. Key changes include adding openai-agents (\>=0.18.1) and mcp (\>=1.19.0,\<3) as base dependencies, bumping litellm to 1.97.0, and adding pyyaml, regex, sortedcontainers, and pypdfium2. Optional dependencies for claude-agent-sdk and anthropic are also defined for extended agent support.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Baseline
- First survey — no prior run to compare against. CAI 65.
Lenses
- Code Health 75
- Architecture 100
- Maturity 51
- Readiness 78
- Security 79
Changes since last survey
- 300 commits — 250 feature/other, 50 fixes
By area
- (root) — 136 commits
- (repo) — 38 commits
- pageindex/page_index.py — 25 commits
- .github/workflows — 13 commits
- pageindex/client.py — 12 commits
- pageindex/flash — 12 commits
- pageindex/agent_tools.py — 10 commits
- cookbook/agentic_retrieval.ipynb — 5 commits
- pageindex/utils.py — 5 commits
- cookbook/pageindex_RAG_simple.ipynb — 4 commits
- cookbook/vision_RAG_pageindex.ipynb — 4 commits
- examples/agentic_vectorless_rag_demo.py — 4 commits
- pageindex/integrations — 4 commits
- cookbook/pageindex-flash-demo.ipynb — 3 commits
- pageindex/init.py — 3 commits
- pageindex/local_chat.py — 3 commits
- pageindex/page_index_md.py — 3 commits
- pageindex/tree_optimize.py — 3 commits
- tests/test_agent_tools.py — 3 commits
- examples/tutorials — 2 commits
Notable commits
- fix: Fix backfill pagination: use raw count instead of filtered count
- fix: Fix backfill: replace gh issue list with gh api for pagination
- fix: Fix contact link and clean up README
- fix: Fix issues from Copilot review: 403 retry, comments pagination, backfill pagination
- fix: Fix list_index variable shadowing in fix_incorrect_toc
- fix: Merge pull request #132 from VectifyAI/fix/backfill-dedupe-pagination
- fix: Merge pull request #133 from VectifyAI/fix/allow-bot-trigger
- fix: Merge pull request #142 from VectifyAI/fix/allow-all-users-dedupe
- fix: Merge pull request #167 from VectifyAI/fix/list-index-shadowing
- fix: Merge pull request #370 from VectifyAI/fix-flash-readme
- fix: Merge pull request #371 from VectifyAI/fix-flash-readme
- fix: Merge pull request #374 from VectifyAI/fix/flash-requirements
- fix: Merge pull request #375 from VectifyAI/fix/default-merge
- fix: Merge pull request #505 from VectifyAI/readme-fix
- fix: Merge pull request #63 from luojiyin1987/fix/api-error-return
- fix: Merge pull request #65 from luojiyin1987/fix/extract-toc-infinite-loop
- fix: Review fixes: unify format validation, pin key set, self-check exception list
- fix: fix
- fix: fix agent integration
- fix: fix agent integration
- …and 280 more
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
VectifyAI/PageIndex was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 18 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 9a8dd6658278fec90347e8ac3388a205305667a3 — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-5d04157a340d.