Skip to content
CAI
Software that uses CAICheck a score

run-llama/liteparse

52.8

Adequate · 29 September 2026

58k

lines of production code

Rust

with Python

2

measurements over time

CAI band scale
CAI trend line
CAI lens gauges

What this system is

LiteParse is a high-performance document parsing library that extracts structured content from PDFs, DOCX, XLSX, and images. It provides a unified Rust core with bindings for Node.js, Python, and WebAssembly, enabling execution across servers, edge runtimes, and browsers. The system supports comprehensive extraction of text, layout, forms, and images, enhanced by configurable OCR engines and advanced markdown conversion.

Features

Add EasyOCR and PaddleOCR server implementations

New standalone OCR services are introduced for EasyOCR (port 8828) and PaddleOCR (port 8829), each exposing a FastAPI server that conforms to the LiteParse OCR API specification. These services accept image uploads via a POST /ocr endpoint, support language selection (with PaddleOCR normalizing aliases like 'zh' to 'ch'), and return structured results including text, bounding boxes, and confidence scores. The PaddleOCR service includes optimizations for large images via configurable detection limits and preserves polygon data for rotated text detection.

ocr/paddleocr · high confidence

Automated pdfium binary distribution for Python wheels

Added build scripts to streamline the inclusion of the pdfium shared library in the Python package. The new \download-pdfium.sh\ script fetches prebuilt binaries for various platforms (macOS, Linux, Windows) and stages them into the \liteparse\ directory for inclusion in the wheel, while \copy-pdfium.sh\ handles runtime deployment by copying the library to the package directory to ensure it is found via platform-specific paths like \@loader\_path\ or \$ORIGIN\.

packages/python/scripts · high confidence

Initial WebAssembly support for browser-based PDF parsing

The pdfium-sys crate now supports building for WebAssembly targets, enabling PDF parsing directly in the browser. The build script detects wasm32 targets and links the pdfium library statically along with necessary WASI runtime libraries (libc, libc++, etc.), while non-wasm targets continue to use dynamic loading. This change introduces the necessary bindings and build configuration to allow Rust-based PDF processing in web environments.

crates/pdfium-sys · high confidence

Introduce LiteParse Node.js bindings with bounded-memory parsing and worker pools

The \packages/node\ directory now provides the official Node.js/TypeScript bindings for LiteParse, exposing a \LiteParse\ class that supports JSON, text, and Markdown output formats. The API introduces opt-in extraction capabilities for structured data including annotations, AcroForm fields, tagged-PDF structure trees, vector graphics, and embedded images. To handle large documents efficiently, the bindings add bounded-memory batch parsing via an async iterator and a persistent worker-pool architecture with configurable timeouts, allowing high-throughput services to bypass the underlying PDFium process-global lock.

packages/node · high confidence

Introduce LiteParse Node.js library with worker pools and extensive extraction options

The Node.js package now exposes a new \LiteParse\ API that supports parsing via a process-isolated worker pool with configurable timeouts and automatic worker recycling. Users can enable a wide range of extraction features including embedded images, PDF annotations, AcroForm fields, tagged-structure trees, vector graphics, and page complexity signals. The library also supports OCR hedging for lower latency, screenshot generation, and detailed text metadata. A new CLI (\liteparse\) is included for command-line usage, supporting JSON, text, and markdown output formats.

packages/node/src · high confidence

Introduce LiteParse Python package with lazy loading and process-isolated worker pools

The \packages/python/liteparse\ directory now contains the complete Python package for LiteParse, providing a public API via \\_\init\\_.py\ that uses lazy imports to significantly reduce startup time. The package exposes document parsing capabilities through the \LiteParse\ class and \search\_items\ function, wrapping native Rust bindings. It includes a \WorkerPool\ implementation (\\_pool.py\) that manages process-isolated worker subprocesses for parallel parsing, ensuring stability and performance. The package also defines comprehensive type structures (\types.py\) for extracted content such as text items, layout blocks, annotations, form fields, and structure trees, along with a CLI entry point (\cli.py\) for command-line usage.

packages/python/liteparse · high confidence

Introduce LiteParse browser-based PDF parser demo site

Adds a new web-based demo application for LiteParse, allowing users to parse PDFs directly in the browser. The site features a drag-and-drop interface for uploading PDF files, status indicators for loading and processing states, and controls for configuring parsing options, providing a visual way to interact with the library's capabilities.

wasm-demo-site · high confidence

Introduce Python bindings for the LiteParse document parsing engine

This change adds the \liteparse-python\ crate, providing Python bindings (via PyO3) and a CLI tool (\lit\) for the core document parsing library. Users can now parse PDFs, DOCX, XLSX, and images from Python or the command line, accessing structured outputs including text, word-level bounding boxes, layout blocks, and optional features like screenshots, embedded image extraction, and AcroForm field parsing. The bindings also handle platform-specific linking for the underlying \libpdfium\ dependency.

crates/liteparse-python · high confidence

Introduce Surya OCR 2 service with GPU-accelerated single-container deployment

Adds a new OCR service in the \ocr/suryaocr\ directory that wraps the Surya 2 multilingual OCR model. The service exposes a FastAPI server (\server.py\) with \POST /ocr\ and \GET /health\ endpoints, conforming to the LiteParse OCR API specification. It supports GPU inference via a bundled llama.cpp backend (offloading all layers via \LLAMA\_CPP\_NGL=99\) or vLLM, and includes a Dockerfile that combines the API and model server in a single image for simplified deployment. Tests (\test\_server.py\) verify block-level result mapping, HTML stripping, and endpoint behavior.

ocr/suryaocr · high confidence

Introduce WebAssembly bindings for browser-based PDF parsing

Adds a new \liteparse-wasm\ crate that exposes a JavaScript-facing API mirroring the Node package, allowing users to parse PDFs directly in the browser. The \LiteParse\ class accepts a configuration object with options for OCR, output formats (JSON, text, markdown), and extraction of images, links, annotations, and form fields. To support the underlying PDFium library, the crate includes WASI stubs for libc functions and fallback implementations for setjmp/longjmp exception handling, enabling the native parsing engine to run within the WebAssembly environment.

crates/liteparse-wasm/src · high confidence

Introduce block-level markdown layout module with advanced table handling

The \liteparse\ crate now includes a new \markdown\_layout\ module that classifies page content into structured markdown blocks (headings, paragraphs, lists, code, tables, and figures). This module introduces robust table detection, including a cross-region re-merge pass to reconstruct tables split by layout cuts, support for merged cells (rowspan/colspan) via \SpanCell\, and fallback rendering for ambiguous tabular regions. It also features improved heading detection with TOC suppression, header/footer stripping, and proper handling of inline styles and links.

_crates/liteparse/src/markdown\layout · high confidence

Introduce new pdfium binding crate with typed accessors and extraction capabilities

A new \crates/pdfium\ library has been added, providing a safe, typed Rust wrapper around the PDFium FFI. This crate introduces a process-global serialization lock to ensure thread safety when interacting with the underlying PDFium library. It exposes core document and page models, including \Document\, \Page\, and \TextPage\, along with specialized accessors for extracting embedded images, vector paths, and AcroForm fields. The implementation also supports tagged PDF structure trees, allowing consumers to access semantic document structure, and provides bitmap rendering utilities with format conversion (BGRA to RGB/RGBA). This change establishes the foundational API for PDF parsing and content extraction within the product.

crates/pdfium · high confidence

Introduces N-API bindings for the LiteParse document parser

This change adds the \crates/liteparse-napi\ module, exposing the Rust \liteparse\ engine to JavaScript/Node.js via N-API. It provides a \LiteParse\ class with methods to parse documents from file paths or raw buffers, take page screenshots, and assess page complexity. The bindings also support bounded-memory batch parsing through \open\_batch\_session\, allowing large documents to be processed in configurable page chunks. Configuration is exposed via \JsLiteParseConfig\, enabling users to control OCR settings, output formats (JSON, text, markdown), and extraction options for images, form fields, and structure trees.

crates/liteparse-napi/src · high confidence

New HTTP, OAR, and Tesseract OCR engines with resilience and multi-language support

The OCR subsystem now supports three distinct backends: an HTTP-based engine that communicates with external OCR services (supporting both standard and production worker response formats), an optional native ONNX backend via the \oar-ocr\ crate (with auto-download support for PP-OCRv6 models), and a Tesseract-based engine. The HTTP engine includes configurable retry/backoff policies and request hedging to improve resilience against transient failures and reduce tail latency. The Tesseract engine now supports multi-language recognition (e.g., \eng+fra\) and automatically downloads missing language data on first use. All engines conform to the \OcrEngine\ trait and return structured results with bounding boxes, confidence scores, and optional polygons for rotated text detection.

crates/liteparse/src/ocr · high confidence

New JSON and text output formatters for parsed PDF data

The \liteparse\ crate now includes dedicated modules for formatting parsed PDF results into structured JSON and plain text. The new JSON output (\json.rs\) serializes page content, text items with optional metadata (such as font details, rotation, and colors), extracted images, annotations, form fields, structure trees, and vector graphics, while gating verbose fields like \content\_bounds\ and creator/producer metadata to API-only usage to maintain stable default CLI output. The text output (\text.rs\) provides a simple formatter that joins page text with page headers. These changes introduce new serialization logic and data structures for the parse result, enhancing the library's ability to expose detailed PDF extraction data in standard formats.

crates/liteparse/src/output · high confidence

New build, release, and testing scripts for LiteParse

Added a suite of build and release scripts to the \scripts/\ directory. \build-glibc-cli.sh\ and \build-glibc-node.sh\ compile the CLI and Node.js bindings inside an old-glibc Debian container to ensure compatibility with older hosts like AWS Lambda. \build-musl-node.sh\ and \build-musl-py.sh\ build the musl-based Node and Python artifacts, bundling required C++ runtime libraries. \bump-version.py\ automates version updates across all packages. Testing and release tooling includes \compare-dataset.sh\ and \compare-outputs.sh\ for regression testing, \create-dataset.sh\ for generating baseline data, \smoke-test-musl-node.sh\ for validating musl builds, and \upload-dataset.sh\ for publishing datasets to HuggingFace. Documentation tooling includes \generate-api-docs.sh\ and \strip-impls-from-api-docs.py\ for generating API references, and \sync-docs-to-developer-hub.sh\ for syncing docs.

scripts · high confidence

New core parsing library with comprehensive PDF extraction capabilities

The \crates/liteparse/src\ directory now contains the core implementation of the LiteParse library, replacing previous ad-hoc parsing logic. This new module introduces a configurable parsing pipeline (\LiteParseConfig\) that supports OCR (local and HTTP-based), embedded image extraction, AcroForm field extraction, and tagged PDF structure tree extraction. It adds robust handling for right-to-left text via a new \bidi.rs\ module, repairs orphaned form widgets in malformed PDFs (\acroform\_repair.rs\), and extracts document provenance metadata (\document\_metadata.rs\). The library also supports page orientation corrections, vector graphics extraction, and screenshot generation, providing a unified entry point for converting various document formats to PDF and extracting structured content.

crates/liteparse/src · high confidence

New dataset evaluation and benchmarking toolkit for PDF parsers

The \dataset\_eval\_utils\ package introduces a comprehensive toolkit for generating ground-truth datasets and evaluating PDF text extraction quality. It provides CLI tools (\lp-process\, \lp-evaluate\, \lp-benchmark\) to benchmark multiple parsers—including LiteParse, PyMuPDF, PyPDF, MarkItDown, and OpenDataLoader—using LLM-based QA evaluation and performance metrics. The toolkit supports generating structured QA pairs via Anthropic's Claude vision capabilities and produces detailed HTML reports with pass rates and latency breakdowns.

_dataset\_eval\utils · high confidence

Behavioural changes

Improved setjmp/longjmp handling in WASM builds

The WASM build now detects whether the linked pdfium library provides a real libsetjmp.a. If present, it uses the native longjmp implementation; otherwise, it falls back to custom stubs. This prevents FreeType's error recovery from aborting the module when the wrong longjmp definition is linked.

crates/liteparse-wasm · high confidence

Runtime loading of pdfium shared library

The pdfium-sys crate now loads the pdfium shared library at runtime using libloading instead of linking it at compile time. This change avoids rpath issues when liteparse is used as a library dependency in other Rust projects, allowing the library to locate and load the pdfium binary dynamically.

crates/pdfium-sys/src · high confidence

WASM module now runs in the browser via WASI stub injection

A new post-build script patches the generated WASM glue code to replace unresolvable WASI and 'env' imports with JavaScript stubs. This allows the WASM module to instantiate and run in browser environments by providing no-op or error-returning implementations for system calls (like file I/O) and logging stderr to the console, effectively bridging the gap between the native WASI expectations and the browser's sandboxed environment.

packages/wasm/scripts · high confidence

Fixes

Fix runtime library loading for bundled PDFium in N-API bindings

The N-API bindings for liteparse now correctly locate the bundled libpdfium at runtime by setting the rpath to the directory containing the .node binary (@loader\_path on macOS, $ORIGIN on Linux). Previously, the build script incorrectly included the build-time pdfium directory in the rpath, which could cause issues in release artifacts; this change ensures the native library is found relative to the installed module.

crates/liteparse-napi · high confidence

Test coverage

Added Python wrapper tests for parsing, batch processing, screenshots, and worker pools; Added browser-based WASM compatibility tests for LiteParse; Added edge runtime compatibility test for LiteParse WASM; Added integration test fixtures for PDF parsing edge cases; Added integration tests for LiteParse parsing and screenshot capabilities; Added tests for dual ESM/CJS exports and worker pool functionality.

Dependencies

LiteParse v2.14.7 release with updated dependencies and platform bindings

This update releases LiteParse v2.14.7 across all language bindings (Node.js, Python, WebAssembly) and the core Rust library. The release includes dependency bumps for the Python evaluation utilities (pypdf, pillow) and OCR services (python-multipart, pytest, pillow), as well as updates to the Rust core's build dependencies (ureq, tar, bindgen). The Node.js package now explicitly supports musl-based Linux environments alongside existing platforms, and the Python bindings bundle the native PDFium library directly into the wheel for easier installation.

(dependencies) · high confidence

Housekeeping

Initial repository scaffolding and documentation

The repository is initialized with core documentation (README, CHANGELOG, AGENTS.md, CONTRIBUTING.md, SECURITY.md, OCR\_API\_SPEC.md), build and CI infrastructure (Dockerfiles, Makefile, GitHub Actions workflows), and development tooling (Prettier, ESLint, .gitignore). This establishes the project structure and configuration for the LiteParse PDF parsing library.

(repo-wide) · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

This is the PUBLIC form of this artifact. Findings are listed in full, but the details of SECURITY findings — which rule fired, in which file, on which line, and how to fix it — are deliberately withheld, and any secret-scanner results are excluded entirely. Where detail is absent here it was REMOVED FOR PUBLICATION; it is not missing from the analysis. The complete artifact is available from the repository owner.

Score

  • CAI 55 → 53 (-2.7)
  • Rubric changed (rubric-2026.09.8 → rubric-2026.09.17) — scores are not directly comparable.

Lenses

  • Code Health 44 → 44 (-0.2)
  • Architecture 97 → 98 (+0.6)
  • Maturity 78 → 79 (+0.3)
  • Readiness 70 → 57 (-12.8)
  • Security 67 → 75 (+8.1)
  • Accessibility 56 → 56 (+0.0)
  • Performance 60 (new)

Resolved (18)

  • Change coupling: native.ts ↔ init.py (packages/node/src/native.ts)
  • Documentation: no architecture or design documentation (README.md)
  • Documentation: no project overview (README.md)
  • Edited copy of a member (19 corresponding lines) (crates/pdfium/src/bitmap.rs)
  • FunctionTooLong: liteparse::ocr_merge::ocr_and_merge_rendered (crates/liteparse/src/ocr_merge.rs)
  • Hotspot: crates/liteparse/src/conversion.rs (crates/liteparse/src/conversion.rs)
  • Hotspot: crates/liteparse/src/markdown_layout/headings.rs (crates/liteparse/src/markdown_layout/headings.rs)
  • Hotspot: crates/liteparse/src/markdown_layout/repetition.rs (crates/liteparse/src/markdown_layout/repetition.rs)
  • Hotspot: crates/liteparse/src/ocr/tesseract.rs (crates/liteparse/src/ocr/tesseract.rs)
  • Hotspot: crates/liteparse/src/render.rs (crates/liteparse/src/render.rs)
  • Hotspot: packages/node/src/lib.ts (packages/node/src/lib.ts)
  • Medium: security finding (details withheld)
  • Medium: security finding (details withheld)
  • Medium: security finding (details withheld)
  • Medium: security finding (details withheld)
  • Medium: security finding (details withheld)
  • liteparse::ocr_merge::ocr_and_merge_rendered (cognitive 67) (crates/liteparse/src/ocr_merge.rs)
  • liteparse::ocr_merge::ocr_and_merge_rendered (cyclomatic 34) (crates/liteparse/src/ocr_merge.rs)

New (40)

  • Ambiguous method return type. quarter_turns_to_apply returns a Result but the signature provided does not specify the error type or the success type. Given the context of orientation correction, it likely returns a number of turns, but the lack of explicit type information in the summary makes it hard to distinguish from methods that might return a boolean or status code. More importantly, having angle and page as properties but a method to calculate 'quarter turns' suggests a derived state that might be better exposed as a computed property or a clearer enum.
  • Duplicate intent with different input types. screenshot takes a str (path) while screenshot_input takes PdfInput. This mirrors the inconsistency in the parsing methods and suggests a lack of unified input abstraction for all operations.
  • Duplicated block (10–14 lines × 2) (packages/python/liteparse/parser.py)
  • Duplicated block (11 lines × 2) (packages/python/liteparse/parser.py)
  • Duplicated block (18 lines × 2) (packages/python/liteparse/parser.py)
  • Edited copy of a member (19 corresponding lines) (crates/pdfium/src/bitmap.rs)
  • FileTooLong: src/raw_text.rs (crates/liteparse/src/raw_text.rs)
  • FunctionTooLong: liteparse::ocr_merge::merge_ocr_results (crates/liteparse/src/ocr_merge.rs)
  • FunctionTooLong: liteparse::ocr_merge::render_pages_for_ocr (crates/liteparse/src/ocr_merge.rs)
  • FunctionTooLong: liteparse::raw_text::build_item (crates/liteparse/src/raw_text.rs)
  • High: security finding (details withheld)
  • Hotspot: packages/python/liteparse/parser.py (packages/python/liteparse/parser.py)
  • Inconsistent and potentially incorrect types for numeric data. Font metrics (height, ascent, descent, weight) and text width are represented as String in both TextItem and TextMetadata. This forces consumers to parse strings to get numeric values, which is error-prone and inefficient. TextItem also has font_size: String while RawTextItem has font_size: f32. This inconsistency within the same domain (text metrics) is confusing.
  • Medium vulnerability: RUSTSEC-2026-0285 (Cargo.lock)
  • Multiple entry points for parsing with unclear separation of concerns. parse takes a string (likely a path), while parse_input takes a PdfInput type. parse_from_pages and parse_from_blocks appear to be lower-level internal helpers exposed publicly, taking already-parsed intermediate structures (Page, PositionedBlock) rather than raw input. This exposes implementation details and creates confusion about which method to use for standard parsing vs. incremental or batch processing.
  • Off the main sequence: liteparse
  • Off the main sequence: liteparse-pdfium-sys
  • Outdated: blake3
  • Outdated: clap
  • Outdated: flate2
  • …and 20 more

Changes since last survey

  • 28 commits — 24 feature/other, 4 fixes

By area

  • (repo) — 8 commits
  • crates/liteparse — 7 commits
  • (root) — 3 commits
  • packages/python — 2 commits
  • .github/workflows — 1 commit
  • crates/liteparse-napi — 1 commit
  • crates/pdfium-sys — 1 commit
  • dataset_eval_utils/README.md — 1 commit
  • dataset_eval_utils/src — 1 commit
  • ocr/paddleocr — 1 commit
  • packages/node — 1 commit
  • scripts/build-glibc-cli.sh — 1 commit

Notable commits

  • fix: Merge pull request #470 from SammyTourani/fix/issue-468
  • fix: Merge pull request #471 from SammyTourani/fix/issue-469
  • fix: fix debian packages in builds
  • fix: fix(ocr): keep parentheses around numbers in OCR table cleanup
  • change: Add page_orientation_corrections: counter-rotate pages by a caller-supplied angle
  • change: Expose the parse pipeline as public stage functions
  • change: Merge pull request #460 from run-llama/logan/extract-compat-glyph-outline
  • change: Merge pull request #461 from run-llama/logan/latest-pdfium
  • change: Merge pull request #462 from run-llama/logan/python-import-times
  • change: Merge pull request #463 from run-llama/logan/orientation-passthrough
  • change: Merge pull request #464 from run-llama/logan/update-benches
  • change: Merge pull request #467 from run-llama/logan/stages-module
  • change: chore: update bench provider code
  • change: clean up implementation
  • change: drasticaly reduce python import times
  • change: drasticaly reduce python import times
  • change: napi: stop adding the build-time pdfium dir to the .node rpath
  • change: nit
  • change: nit
  • change: nit
  • …and 8 more

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

run-llama/liteparse was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 29 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit 69f73d95dd54907baaa01d2fabdbf95d147c25ff — the exact code this score is about.
  • Scored under rubric-2026.09.17 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-70910855e4b4.