Skip to content
CAI
Software that uses CAICheck a score

VectifyAI/PageIndex

65.4

Adequate · 18 September 2026

25k

lines of production code

Python

primary language

1

measurement over time

CAI band scale
CAI lens gauges

What this system is

This system is a Python SDK for PageIndex that enables local and cloud-based document indexing, retrieval, and chat. It features a 'Flash' mode for fast, LLM-free PDF structure extraction and supports multimodal retrieval with block-level citations. The SDK integrates with major agent frameworks like OpenAI and Anthropic, allowing document Q&A within agent workflows via a unified client.

Features

Introduce PageIndex Flash for LLM-free PDF structure extraction

Adds the PageIndex Flash module, a new local indexing mode that builds a hierarchical tree structure from PDF layout statistics without requiring an LLM for the initial extraction. The \page\_index\_flash\ API accepts a PDF path or stream and returns a structured tree with optional LLM-generated node summaries and deterministic optimization (merge or full expand). It automatically detects table of contents sources (layout, embedded bookmarks, or hybrid), handles edge cases like scanned or unreadable PDFs, and includes a rejection policy to prevent indexing documents that lack a viable structure.

pageindex/flash · high confidence

New PageIndex documentation notebooks for citations, local indexing, and multimodal retrieval

The cookbook now includes three new Jupyter notebooks demonstrating key PageIndex capabilities. The 'Citations & Source Highlighting' notebook shows how to use block-level citations to trace answers back to specific regions in source documents. The 'Flash Demo' notebook illustrates building vectorless RAG using the local SDK and tree-indexing model. The 'Multimodal Retrieval & Understanding' notebook demonstrates retrieving and interpreting document images using a vision-capable chat model.

cookbook · high confidence

PageIndex SDK v0.2.x: Local mode, agent tools, and multi-provider chat

The PageIndex SDK now supports a local indexing and chat mode alongside the existing cloud surface, allowing users to index PDFs and chat against them using local models (configurable via \index\_model\ and \chat\_model\) without requiring a cloud API key. The SDK introduces a unified \PageIndexClient\ that routes to either cloud or local backends, and adds agent tool integrations for OpenAI Agents SDK, Anthropic SDK, and Claude Agent SDK, enabling document retrieval and Q&A within agent workflows. Chat functionality is expanded with streaming support (\ChatStream\), multiple protocols (\chat\_completions\, \messages\, \responses\), and LiteLLM integration for multi-provider LLM support, including model passthrough and backend connection overrides. The package also includes lazy loading for performance, improved error handling with \PageIndexAPIError\, and configuration via \config.yaml\ for indexing parameters.

pageindex · high confidence

Behavioural changes

PageIndex SDK introduces local mode with Flash indexing as the default

The PageIndex SDK now supports a local mode that allows users to index, retrieve, and chat entirely on their own machine using their own LLM keys, or connect to PageIndex Cloud. The primary indexing method has shifted to 'PageIndex Flash', a fast tree index generation process for text-based PDFs which is now the default in local mode. This change replaces the previous standard indexing approach and is accessible via the new \run\_pageindex.py\ CLI entry point, which defaults to Flash mode and includes options for optimization and summary generation.

(repo-wide) · high confidence

Test coverage

Added test fixtures and contract data for agent tools and flash layout; Initial test suite for PageIndex SDK and Flash indexing.

Dependencies

Update PageIndex SDK dependencies and add pyproject.toml

The PageIndex Python SDK (v0.2.10) now includes a pyproject.toml manifest and updates its core dependencies in requirements.txt. Key changes include adding openai-agents (\>=0.18.1) and mcp (\>=1.19.0,\<3) as base dependencies, bumping litellm to 1.97.0, and adding pyyaml, regex, sortedcontainers, and pypdfium2. Optional dependencies for claude-agent-sdk and anthropic are also defined for extended agent support.

(dependencies) · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

Baseline

  • First survey — no prior run to compare against. CAI 65.

Lenses

  • Code Health 75
  • Architecture 100
  • Maturity 51
  • Readiness 78
  • Security 79

Changes since last survey

  • 300 commits — 250 feature/other, 50 fixes

By area

  • (root) — 136 commits
  • (repo) — 38 commits
  • pageindex/page_index.py — 25 commits
  • .github/workflows — 13 commits
  • pageindex/client.py — 12 commits
  • pageindex/flash — 12 commits
  • pageindex/agent_tools.py — 10 commits
  • cookbook/agentic_retrieval.ipynb — 5 commits
  • pageindex/utils.py — 5 commits
  • cookbook/pageindex_RAG_simple.ipynb — 4 commits
  • cookbook/vision_RAG_pageindex.ipynb — 4 commits
  • examples/agentic_vectorless_rag_demo.py — 4 commits
  • pageindex/integrations — 4 commits
  • cookbook/pageindex-flash-demo.ipynb — 3 commits
  • pageindex/init.py — 3 commits
  • pageindex/local_chat.py — 3 commits
  • pageindex/page_index_md.py — 3 commits
  • pageindex/tree_optimize.py — 3 commits
  • tests/test_agent_tools.py — 3 commits
  • examples/tutorials — 2 commits

Notable commits

  • fix: Fix backfill pagination: use raw count instead of filtered count
  • fix: Fix backfill: replace gh issue list with gh api for pagination
  • fix: Fix contact link and clean up README
  • fix: Fix issues from Copilot review: 403 retry, comments pagination, backfill pagination
  • fix: Fix list_index variable shadowing in fix_incorrect_toc
  • fix: Merge pull request #132 from VectifyAI/fix/backfill-dedupe-pagination
  • fix: Merge pull request #133 from VectifyAI/fix/allow-bot-trigger
  • fix: Merge pull request #142 from VectifyAI/fix/allow-all-users-dedupe
  • fix: Merge pull request #167 from VectifyAI/fix/list-index-shadowing
  • fix: Merge pull request #370 from VectifyAI/fix-flash-readme
  • fix: Merge pull request #371 from VectifyAI/fix-flash-readme
  • fix: Merge pull request #374 from VectifyAI/fix/flash-requirements
  • fix: Merge pull request #375 from VectifyAI/fix/default-merge
  • fix: Merge pull request #505 from VectifyAI/readme-fix
  • fix: Merge pull request #63 from luojiyin1987/fix/api-error-return
  • fix: Merge pull request #65 from luojiyin1987/fix/extract-toc-infinite-loop
  • fix: Review fixes: unify format validation, pin key set, self-check exception list
  • fix: fix
  • fix: fix agent integration
  • fix: fix agent integration
  • …and 280 more

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

VectifyAI/PageIndex was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 18 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit 9a8dd6658278fec90347e8ac3388a205305667a3 — the exact code this score is about.
  • Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-5d04157a340d.