Skip to content
CAI
Software that uses CAICheck a score

datalab-to/surya

48.6

Weak · 19 September 2026

11.2k

lines of production code

Python

primary language

1

measurement over time

CAI band scale
CAI lens gauges

What this system is

Surya is a high-performance document analysis system that performs layout detection, text recognition, and table structure extraction from images and PDFs. It utilizes a client-server architecture with shared model processes to optimize memory usage and throughput, supporting both GPU and CPU inference via vLLM and llama.cpp backends. The system provides granular capabilities including reading-order detection, OCR error identification, and HTML-based table output, all managed through a unified, type-safe inference interface.

How it got here

2024 — Project initialization and architectural modernization

6 changes.

The project was initialized with core infrastructure, licensing, and a migration from Poetry to Hatch, while upgrading core dependencies to modern versions. Significant architectural changes included replacing legacy OpenCV-based postprocessing and VLM-based detection with a faster, shared-server layout predictor and standardized input utilities.

2025–2026 — Infrastructure consolidation and server architecture

14 changes.

This period focused on consolidating scattered logic into a shared common infrastructure and refactoring core modules to use a client-server architecture for memory efficiency. The project introduced new inference backends like llama.cpp and vLLM, standardized data schemas, and added modular CLI tools and debug utilities to improve usability and performance.

Features

Introduce dual-mode table structure recognition

The \surya/table\_rec\ module now provides a \TableRecPredictor\ with two distinct recognition modes. The default 'simple' mode predicts rows and columns, deriving cells geometrically from their intersections, which is suitable for standard grid structures. The new 'full' mode uses a block prompt to generate complete HTML output, supporting complex table features like colspan, rowspan, and headers, with the resulting HTML accessible via the \TableResult.html\ property for direct consumption by downstream tools.

_surya/table\rec · high confidence

Introduce fast RF-DETR layout and table detection models

Surya now includes fast layout and table detection capabilities powered by a vendored, slimmed copy of Roboflow's RF-DETR. This change adds a pure PyTorch implementation of the RF-DETR architecture—including DINOv2-based backbones, multi-scale projectors, and LW-DETR detection heads—directly into the codebase. By vendoring the model, the system avoids a runtime dependency on the heavy \rfdetr\ package and its transitive dependencies (such as \roboflow\, \rf100vl\, \albumentations\, \supervision\, and \peft\), while still supporting CPU, MPS, and CUDA execution.

surya/common/rfdetr · high confidence

Introduce fast-layout detection with reading-order support and shared server architecture

Surya now includes a new fast-layout predictor (surya.fast\_layout) that uses a lightweight rf-detr object detector instead of the previous VLM-based approach, providing significantly faster page layout analysis. This new component introduces autoregressive reading-order detection, allowing the layout engine to determine the correct reading sequence of detected text blocks. To optimize resource usage, the fast-layout model runs in a single shared server process that coalesces requests from multiple client processes, reducing memory overhead and improving throughput. The system also includes a new screenshot application (surya\_screenshot.html) for visualizing full-page OCR results with layout boxes. Additionally, the old detection and processing modules (surya/detection.py, surya/model/processing.py, surya/model/segformer.py) have been removed as part of this architectural shift.

surya · high confidence

Introduce shared common infrastructure for Surya models

This change establishes a new \surya/common\ package that consolidates shared logic previously scattered across the project. It adds a \BasePredictor\ and \ModelLoader\ to standardize model initialization, device placement, and batch-size handling, while introducing \S3DownloaderMixin\ to enable loading models directly from S3 with parallel downloads and retry logic. A new \RfDetrTorch\ implementation provides a fast, pure-PyTorch layout detector backed by vendored rf-detr weights, and utility modules are added for detecting blank image regions, validating transformer backbone output indices, and cleaning bounding boxes.

surya/common · high confidence

New PDF and image loading utilities

Added \surya/input/load.py\ and \surya/input/processing.py\ to provide standardized functions for loading PDFs and images. The new \load\_from\_file\ and \load\_from\_folder\ functions automatically detect file types (using \filetype\) and handle page ranges and DPI settings for PDFs, while \load\_image\ handles single image files. The underlying \processing.py\ module uses \pypdfium2\ to render PDF pages to PIL images.

surya/input · high confidence

New block-based OCR recognition predictor with full-page fallback

The recognition module now uses a new RecognitionPredictor that performs per-block OCR by cropping layout-defined regions and running a block-specific prompt, while also supporting a full-page OCR mode that includes automatic detection and removal of blank text blocks and handling of repetition loops. This change introduces a more granular approach to text extraction, allowing for better handling of complex layouts and improved accuracy through full-page regeneration strategies when needed.

surya/recognition · high confidence

New debug visualization tools for bounding boxes and HTML rendering

Added a suite of new modules in the \surya/debug\ package to support visual debugging and HTML export. \draw.py\ and \text.py\ provide utilities for drawing bounding boxes and polygons on images with labels, while \fonts.py\ handles automatic font downloading and path resolution. Additionally, \render\_html.py\ and \katex.js\ enable rendering text content as HTML with support for KaTeX math expressions, allowing users to export debug views as styled web pages.

surya/debug · high confidence

New llama.cpp and vLLM inference backends with auto-scaling concurrency

The inference system now supports two new execution backends: llama.cpp (spawning the native \llama-server\ binary) and vLLM (spawning a Docker container). These backends share a common OpenAI-compatible client for batch processing, enabling structured outputs and retry logic. The vLLM backend introduces auto-scaling client concurrency that adjusts parallel requests based on detected GPU VRAM, capping at 96 concurrent requests to avoid GPU saturation.

surya/inference/backends · high confidence

New modular CLI scripts and dedicated screenshot viewer

The \surya/scripts\ package now provides a set of standalone, modular command-line tools for specific tasks: layout detection (\detect\_layout.py\), text detection (\detect\_text.py\), full-page OCR (\ocr\_text.py\), and table recognition (\table\_recognition.py\). These scripts share a common \CLILoader\ configuration class that standardizes input handling, output directories, and page-range selection. Additionally, a new screenshot-friendly web application (\screenshot\_app.py\) has been added, offering a side-by-side viewer for PDF/image pages and their OCR output, designed for easy visual inspection and screenshot capture.

surya/scripts · high confidence

Project initialization with licensing, tooling, and dependency infrastructure

The repository has been initialized with the core project infrastructure, including the Apache 2.0 license for the codebase and a modified OpenRAIL-M license for the machine learning models, which imposes use-based restrictions for entities with significant revenue or funding. A Contributor License Agreement (CLA) and CITATION file have been added to manage contributions and academic attribution. Development tooling is established via a pre-commit configuration for Ruff linting and formatting, a pytest configuration, and a \uv.lock\ file pinning Python 3.10+ dependencies such as \aiohttp\.

(repo-wide) · high confidence

Removals

Removed legacy OpenCV-based postprocessing modules

The \surya/postprocessing\ module has removed the legacy OpenCV-based postprocessing logic, specifically deleting \affinity.py\, \heatmap.py\, and \util.py\. This eliminates the previous Sobel/Canny/Hough-based line detection, connected-component box detection, and manual bounding-box scaling utilities, indicating a shift to a different postprocessing implementation (likely model-driven) that no longer relies on these specific computer vision primitives.

surya/postprocessing · high confidence

Behavioural changes

Introduce SuryaInferenceManager with automatic backend selection

The inference layer now uses a centralized SuryaInferenceManager that automatically selects the optimal backend (vLLM for NVIDIA GPUs, llama.cpp for CPU/MPS) based on hardware detection, exposing a unified interface for layout, table, and block prediction tasks with configurable concurrency capacity.

surya/inference · high confidence

Introduces a new batch-service architecture with unified server lifecycle and typed inference schemas

Surya now uses a dedicated batch-service layer to manage model inference servers. This change adds a generic HTTP server with continuous batching (configurable via \batch\_wait\_ms\ and \max\_batch\) and a robust attach-or-spawn lifecycle manager that handles health checks, file locking, and cleanup. It also introduces a comprehensive set of Pydantic schemas for all inference tasks (layout, table, OCR, detection) and a \PolygonBox\ utility for coordinate handling, replacing previous ad-hoc data structures with a consistent, type-safe interface for model inputs and outputs.

surya-ocr · high confidence

Introduction of EfficientViT detection model with S3 support

The detection module now includes a new EfficientViT-based model implementation, replacing the previous architecture. This change introduces a new configuration class and encoder-decoder structure for the detection model, along with added support for loading models directly from S3 storage via the \from\_pretrained\ method.

surya/detection/model · high confidence

Layout predictor now filters blank text regions and maps model labels to canonical names

The layout module now includes logic to detect and drop text-labeled layout blocks that correspond to blank or uniform-color image regions, reducing false positives for text detection. Additionally, raw labels emitted by the underlying foundation model are mapped to a canonical set of public label names (e.g., 'Image' becomes 'Picture', 'Equation-Block' becomes 'Equation'), ensuring consistent output for downstream consumers.

surya/layout · high confidence

Migrate OCR error model to Hugging Face Transformers 5

The OCR error model configuration has been refactored to inherit from Hugging Face Transformers' \PretrainedConfig\, enabling standard model loading and integration. This change introduces support for downloading model weights from S3 via the \S3DownloaderMixin\ and defines the specific DistilBert configuration parameters (such as hidden size, attention heads, and layers) required for the model's operation.

_surya/ocr\error/model · high confidence

OCR error detection now uses a shared server process

The OCR error detection capability has been refactored to run the DistilBert model in a single shared server process, with client code communicating via HTTP. This change significantly reduces memory usage when multiple worker processes are active, as they no longer each load their own copy of the model. The public API remains unchanged, ensuring existing callers are unaffected.

_surya/ocr\error · high confidence

Text detection now runs via a shared server process

The text detection module has been refactored to use a client-server architecture. By default, the DetectionPredictor now acts as a client that sends requests to a shared background server (surya.detection.server), which manages a single model instance and handles continuous batching. This reduces memory overhead when multiple workers are active. Users can still opt for a local, process-owned model instance via the new DetectionPredictor.local() method, and the server supports configurable ports and checkpoints.

surya/detection · high confidence

Test coverage

Initial test suite for Surya prediction models

Added a comprehensive test suite covering the core prediction components: detection, layout, recognition, table recognition, and OCR error detection. The tests verify that predictors load correctly, process sample images, and return expected result structures (such as bounding boxes, reading orders, and cell grids). A specific regression test ensures that model checkpoint weights are loaded correctly without being overwritten by initialization routines. Additionally, tests for the fast-layout server validate engine correctness, continuous batching behavior, and box-merging logic.

tests · high confidence

Dependencies

Migrate from Poetry to Hatch and upgrade core dependencies

The project has switched its build system from Poetry to Hatch (pyproject.toml), replacing the deleted poetry.lock with a new dependency specification. This update renames the package to surya-ocr, raises the minimum Python version to 3.10, and significantly upgrades core libraries: transformers is bumped to \>=5.12.1, torch to \>=2.7.0, and torchvision to \>=0.20.0. Other dependencies such as pypdfium2 (\>=5.10.1) and opencv-python-headless (==4.11.0.86) have also been updated to newer versions.

(dependencies) · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

Baseline

  • First survey — no prior run to compare against. CAI 49.

Lenses

  • Code Health 93
  • Architecture 97
  • Maturity 48
  • Readiness 39
  • Security 67
  • Accessibility 48

Changes since last survey

  • 300 commits — 248 feature/other, 52 fixes

By area

  • (repo) — 65 commits
  • surya/common — 60 commits
  • surya/foundation — 59 commits
  • (root) — 42 commits
  • signatures/version1 — 24 commits
  • surya/recognition — 12 commits
  • surya/layout — 9 commits
  • surya/scripts — 6 commits
  • surya/settings.py — 6 commits
  • surya/inference — 4 commits
  • surya/table_rec — 4 commits
  • static/images — 2 commits
  • surya/detection — 2 commits
  • .github/ISSUE_TEMPLATE — 1 commit
  • .github/workflows — 1 commit
  • benchmark/table_recognition.py — 1 commit
  • surya/fast_layout — 1 commit
  • tests/assets — 1 commit

Notable commits

  • fix: Accuracy fixes
  • fix: Bugfix in cache
  • fix: Bugfix in decode udpate - Text token counts were wrong
  • fix: CI fix
  • fix: Fix #406 - OCR with blank images
  • fix: Fix bad test - Add real latex image
  • fix: Fix beacon issue
  • fix: Fix cache positions for SDPA
  • fix: Fix decode attention mask update
  • fix: Fix edge case for empty tags
  • fix: Fix embedding with a static scatter
  • fix: Fix encoder chunking
  • fix: Fix generation with beacon tokens
  • fix: Fix issues
  • fix: Fix issues with GPU codepaths
  • fix: Fix latex detokenization bug - Unescaping sequences
  • fix: Fix layout and table rec image bbox
  • fix: Fix license
  • fix: Fix mark steps
  • fix: Fix math mode for layout
  • …and 280 more

Architecture

  • 0 containers · 1 bounded contexts · 0 dependency edges (baseline)

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

datalab-to/surya was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 19 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit a2363d3311773a2b6145a40458043211c50c52f4 — the exact code this score is about.
  • Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-13a154b7f5d1.