PaddlePaddle/PaddleOCR
67.5
Adequate · 26 September 2026
129.2k
lines of production code
Python
with C++
3
measurements over time
What this system is
This system is a comprehensive Optical Character Recognition (OCR) and document analysis framework built on PaddlePaddle. It provides a full lifecycle of capabilities, including training, evaluation, and deployment of text detection, recognition, and complex document understanding models like table and layout analysis. The system supports diverse inference environments, ranging from cloud APIs and Docker-based servers to mobile apps (Android/iOS) and browser-based JavaScript execution.
How it got here
2020 — modular architecture and deployment expansion
26 changes.
This period focused on restructuring PaddleOCR into a modular, factory-based architecture for backbones, heads, and post-processing, while significantly expanding model support with new algorithms and multi-language capabilities. It also introduced comprehensive deployment tools, including C++ and Android inference APIs, Paddle-Lite mobile demos, and PaddleHub serving modules for various OCR tasks.
2021–2022 — PP-Structure expansion and model compression
31 changes.
This period focused on expanding the PP-Structure suite with comprehensive table recognition, layout analysis, and key information extraction capabilities, alongside new tools for PDF-to-Word conversion. Significant effort was also directed toward model optimization, introducing new workflows for pruning and quantization to support efficient deployment on edge devices and mobile platforms.
2023–2026 — modular architecture and multi-platform expansion
21 changes.
This period focused on restructuring PaddleOCR into a modular inference package with specialized pipelines and CLI subcommands, while significantly expanding deployment options across cloud, edge, and browser environments. The work introduced comprehensive SDKs for Python, Go, and JavaScript, alongside extensive Docker configurations for diverse hardware accelerators and new Android and C++ inference modules. Additionally, it added capabilities for document-to-Markdown conversion, formula dataset utilities, and integration with LangChain and MCP servers.
Features
Add Docker deployment configurations for AMD, Intel, Huawei, Hygon, Iluvatar, MetaX, and Kunlunxin accelerators
This change introduces new Docker deployment configurations for PaddleOCR-VL across a wide range of hardware accelerators, including AMD GPU, Intel GPU, Huawei NPU, Hygon DCU, Iluvatar GPU, MetaX GPU, and Kunlunxin XPU. Each accelerator directory now contains a complete set of files: a \.env\ file for image tagging and backend selection (vllm or fastdeploy), a \compose.yaml\ for orchestrating the API and VLM server services, and \pipeline.Dockerfile\ and \vlm.Dockerfile\ for building the specific container images. These additions enable users to deploy the PaddleOCR-VL service on these specific hardware platforms using Docker Compose, with appropriate device mounts, environment variables, and backend-specific optimizations (e.g., ROCm for AMD, Ascend for Huawei, DCU for Hygon, etc.).
_deploy/paddleocr\_vl\docker/accelerators · high confidence
Add HTML-to-Excel conversion module (tablepyxl)
Introduces a new \tablepyxl\ module within \ppstructure\ that converts HTML documents containing tables into Excel workbooks. The module parses HTML tables, applies CSS styles (fonts, alignment, borders, fills) via \openpyxl\, and handles cell merging and column widths, allowing users to export structured table data from HTML to \.xlsx\ files.
ppstructure/table/tablepyxl · high confidence
Add KIE SER and SER\_RE Hub Serving modules
New PaddleHub serving modules for Key Information Extraction (KIE) have been added: \kie\_ser\ (entity recognition) and \kie\_ser\_re\ (relation extraction). These modules are deployed on ports 8871 and 8872 respectively, utilize the LayoutXLM algorithm, and expose serving methods that accept base64-encoded images to perform prediction tasks.
_deploy/hubserving/kie\_ser, deploy/hubserving/kie\_ser\re · high confidence
Add KIE data conversion and evaluation tools
New utility scripts have been added to the KIE tools module to support data preprocessing and model evaluation. \trans\_funsd\_label.py\ converts FUNSD dataset annotations into a standardized JSON format, while \trans\_xfun\_data.py\ performs a similar conversion for the XFUN dataset. Additionally, \eval\_with\_label\_end2end.py\ provides an end-to-end evaluation script that calculates precision, recall, and F-measure by matching ground truth and predicted bounding boxes using IoU and comparing text via edit distance.
ppstructure/kie/tools · high confidence
Add OCR text angle classification HubServing module
A new PaddleHub serving module for OCR text angle classification has been added to the deploy/hubserving/ocr\_cls directory. This module exposes a service (default port 8866) that accepts base64-encoded images and returns the predicted text angle (0 or 180 degrees) along with a confidence score. It utilizes the existing TextClassifier inference logic, supports GPU acceleration via CUDA\_VISIBLE\_DEVICES, and is configured with default model paths and parameters in params.py.
_deploy/hubserving/ocr\cls · high confidence
Add OCR text detection HubServing module
Introduces a new deploy/hubserving/ocr\_det module that exposes PaddleOCR text detection as a serving endpoint. The module defaults to the PP-OCRv3 mobile detection model, supports GPU acceleration and MKL-DNN optimization, and listens on port 8865 for incoming image requests.
_deploy/hubserving/ocr\det · high confidence
Add PP-OCR model pruning tutorial and scripts
Introduces a new model compression workflow for PP-OCR models using PaddleSlim. This location provides the documentation (README.md/README\_en.md) and executable scripts (sensitivity\_anal.py, export\_prune\_model.py) that enable users to perform sensitivity analysis, apply FPGM filter pruning to reduce model redundancy, and export the optimized model for inference deployment.
deploy/slim/prune · high confidence
Add PP-OCRv3 mobile text recognition configurations
This change introduces a new set of configuration files for the PP-OCRv3 mobile text recognition model, located in configs/rec/PP-OCRv3. It includes base configurations for the standard mobile model (PP-OCRv3\_mobile\_rec.yml) and a distillation variant (PP-OCRv3\_mobile\_rec\_distillation.yml) which sets up a teacher-student architecture using SVTR\_LCNet with MobileNetV1Enhance backbone and MultiHead (CTC + SAR) heads. Additionally, it adds specific configurations for multiple languages under the multi\_language subdirectory, including English, Arabic, Chinese Traditional, Cyrillic, Devanagari, Japanese, Georgian (ka), Korean, Latin, Tamil, and Telugu, each pointing to their respective character dictionaries.
configs/rec/PP-OCRv3 · high confidence
Add PP-OCRv6 Android Demo application
Introduces a new Android demo application for PP-OCRv6, providing a complete user interface to test the OCR SDK. The app features a modern Compose UI with an image picker for gallery selection and sample images, an overlay for visualizing OCR bounding boxes, and a results list with confidence scores and a copy-to-clipboard function. It also displays performance timing metrics for detection and recognition phases. The underlying \OCRApplication\ component manages the lifecycle of the OCR engine, handling model loading, initialization of OpenCV, and error states with retry capabilities.
deploy/ppocr-android, deploy/ppocr-android/ppocr-sdk · high confidence
Add PP-Structure layout analysis serving module
Introduces a new Hubserving module for document structure layout analysis, allowing users to deploy a service that detects and classifies layout elements (such as text, images, and tables) within documents. The module includes a configuration file setting the service to run on port 8871 with GPU support enabled by default, and a Python implementation that wraps the underlying PaddlePaddle layout predictor to accept base64-encoded images and return bounding box coordinates for detected layout regions.
_deploy/hubserving/structure\layout · high confidence
Add PP-Structure table structure service to HubServing
A new HubServing module for table structure recognition has been added to the deploy/hubserving/structure\_table directory. This service exposes a table structure prediction capability, allowing users to submit images and receive HTML representations of the detected table structures. The implementation includes configuration for GPU usage, a serving endpoint that accepts base64-encoded images, and specific parameters for the underlying table structure model.
_deploy/hubserving/structure\_system, deploy/hubserving/structure\table · high confidence
Add PPOCRv4 auto-compression example for detection and recognition models
This location introduces a complete example for automatically compressing PPOCRv4 detection and recognition models using quantization-aware training and distillation. It includes configuration files for mobile and server variants of both detection (DB algorithm) and recognition (SVTR algorithm) models, a Python entry point (\run.py\) to orchestrate the compression via PaddleSlim, and shell scripts (\test\_ocr\_det.sh\, \test\_ocr\_rec.sh\) to benchmark performance on GPU (with TensorRT) and CPU (with MKLDNN) before and after compression. A dataset processing script is also provided to handle specific image size constraints for the server detection model.
_deploy/slim/auto\compression · high confidence
Add PSE text detection post-processing
Introduces the PSEPostProcess component for the PSE text detection model. This new module handles the conversion of model output maps into bounding boxes, supporting both quadrilateral and polygonal box types, and includes logic for thresholding, area filtering, and coordinate scaling.
_ppocr/postprocess/pse\postprocess · high confidence
Add Paddle-Lite C++ OCR demo for mobile deployment
Introduces a new C++ demo in deploy/lite for running PaddleOCR ultra-lightweight detection and recognition models on mobile devices (Android armv7/armv8) using Paddle-Lite. The change includes the main ocr\_db\_crnn.cc application, preprocessing modules for detection (db\_post\_process), recognition (crnn\_process), and classification (cls\_process), a Makefile to build the binary with OpenCV and Paddle-Lite libraries, a configuration file (config.txt) for model parameters, and documentation (readme.md/readme\_ch.md) guiding users on preparing the cross-compilation environment, optimizing models with paddle\_lite\_opt, and running the demo on a phone via ADB.
deploy/lite · high confidence
Add PaddleOCR DB/EAST/PSE training benchmark scripts
New benchmark scripts have been added to the \benchmark\ directory to measure and analyze the training performance of PaddleOCR detection models (DB, EAST, and PSE). The \run\_det.sh\ script automates the setup by installing dependencies, downloading the ICDAR2015 dataset and pre-trained ResNet backbones, and executing training runs in both single-GPU and multi-GPU modes with varying batch sizes. The \run\_benchmark\_det.sh\ script handles the actual training execution and logs results, while \analysis.py\ provides tools to parse these logs and extract performance metrics like images per second (IPS).
benchmark · high confidence
Add PaddleOCR iOS demo app
A new iOS demo application (PaddleOCRDemo) has been added to the deploy/ios\_demo directory, providing a complete SwiftUI-based interface for running PaddleOCR on-device. The app includes the full source code for the Xcode project, including an OCR engine that orchestrates text detection and recognition via ONNX Runtime, post-processing logic for DB (Differentiable Binarization) and CTC decoding, and UI components for image selection, result display, and engine configuration. It also bundles necessary third-party dependencies (such as OpenCV and Clipper) via CocoaPods and includes sample images and benchmarking tests.
_deploy/ios\demo · high confidence
Add TIPC supplementary training and testing utilities
This change introduces a new \test\_tipc/supplementary\ module providing a complete set of utilities for training, evaluation, and model optimization workflows. It includes a configuration parser (\config.py\) that safely loads YAML settings, data handling components (\data.py\, \data\_loader.py\, \load\_cifar.py\) for image preprocessing and CIFAR-100 dataset loading, and metric/loss modules (\metric.py\, \loss.py\) supporting standard accuracy, label smoothing, and distillation losses (KL/JS). The package also provides a MobileNetV3 implementation (\mv3.py\) with optional custom ReLU operators (C++/CUDA sources included), learning rate schedulers and optimizers (\optimizer.py\), and slimming tools for pruning (\slim\_fpgm.py\) and quantization (\slim\_quant.py\). A main training script (\train.py\) and shell entry point (\train.sh\) tie these components together to support single-GPU and distributed training with AMP, pruning, and quantization capabilities.
_test\tipc/supplementary · high confidence
Add bare-metal PaddleOCR demo for Arm Cortex-M55 on AVH
This change introduces a new deployment example in the \deploy/avh\ directory that runs a PaddleOCR text recognition model on a bare-metal Arm Cortex-M55 CPU using Arm Virtual Hardware (AVH). The package includes a Makefile and CMake toolchain for building the demo with the \arm-none-eabi-gcc\ toolchain and CMSIS-NN, a shell script (\run\_demo.sh\) to automate downloading the model, compiling it via TVM's AOT executor, and launching the simulation on a Corstone-300 FVP, and C source code for the bare-metal application that handles inference and text output.
deploy/avh · high confidence
Add built-in conversion of Office documents (DOCX, XLSX, PPTX) to Markdown
PaddleOCR now includes a new \paddleocr.\_doc2md\ module that converts Word, Excel, and PowerPoint files into Markdown. This feature adds converters for \.docx\, \.xlsx\, and \.pptx\ formats, handling text, formatting, and embedded math formulas (OMML to LaTeX). The module provides a \convert\ function and a registry system to manage these new file-type capabilities.
_paddleocr/\doc2md · high confidence
Add end-to-end OCR evaluation and visualization tools
New scripts have been added to the \tools/end2end\ directory to support end-to-end evaluation of the text detection and recognition pipeline. \convert\_ppocr\_label.py\ converts prediction outputs and ground truth labels into a standardized format for comparison. \eval\_end2end.py\ calculates key metrics including character accuracy, precision, recall, F-measure, and average edit distance. Additionally, \draw\_html.py\ provides a visualization tool to generate an HTML report displaying original images alongside their ground truth labels for debugging and review.
tools/end2end · high confidence
Add layout analysis prediction script and documentation
The \ppstructure/layout\ module now includes a \predict\_layout.py\ script that enables users to run layout analysis inference using PaddlePaddle or ONNX predictors. This script supports configurable preprocessing (resizing, normalization) and post-processing (NMS, score thresholds) for detecting document regions such as text, tables, and figures. Accompanying English and Chinese README files provide comprehensive guides on installation, data preparation (including PubLayNet and CDLA datasets), model training, and evaluation.
ppstructure/layout · high confidence
Add multi-language recognition training configurations
This change introduces a new set of training configuration files and a generation script in the \configs/rec/multi\_language\ directory, enabling users to train the PaddleOCR CRNN model on a wide variety of scripts. The new files provide ready-to-use configurations for Latin, Arabic, Cyrillic, Devanagari, Hebrew, Japanese, Korean, and Samaritan scripts, along with a multi-language base config. A Python script (\generate\_multi\_language\_configs.py\) is also added to help generate specific language configs from a base template, supporting languages such as Italian, Spanish, Russian, Arabic, Tamil, and others.
_configs/rec/multi\language · high confidence
Add new utility modules and character dictionaries to ppocr/utils
This change introduces a suite of new utility files to the ppocr/utils package, including an Exponential Moving Average (EMA) implementation for model training, a network module for downloading models with progress bars, and an export model utility that generates inference configurations with dynamic shapes for various architectures. It also adds character dictionaries (EN\_symbol\_dict, en\_dict, ic15\_dict, ppocr\_keys\_v1) for recognition tasks, along with supporting scripts for label generation, polygon NMS, IoU calculation, logging, profiling, and visualization.
ppocr/utils · high confidence
Added Deteval evaluation metric implementation
The \ppocr/utils/e2e\_metric\ directory now includes \Deteval.py\ and \polygon\_fast.py\, introducing a new evaluation mode for end-to-end text detection. This addition provides the Deteval metric logic, including polygon area and intersection calculations via Shapely, allowing users to evaluate their models using the Deteval standard alongside existing metrics.
_ppocr/utils/e2e\metric · high confidence
Added PGNet end-to-end post-processing utilities
The \ppocr/utils/e2e\_utils\ module now includes support for the PGNet (Polygon Generation Network) end-to-end text detection and recognition pipeline. This change introduces new utility files to handle the specific post-processing requirements of PGNet, including \pgnet\_pp\_utils.py\ which provides a \PGNet\_PostProcess\ class with both fast and slow decoding paths, and dedicated extraction modules (\extract\_textpoint\_fast.py\ and \extract\_textpoint\_slow.py\) for CTC greedy decoding and pivot list generation. Additionally, \extract\_batchsize.py\ was added to manage batch size adjustments for TCL (Text Line Center) ROIs, and \visual.py\ was updated with geometry helpers like \point\_pair2poly\ and \expand\_poly\_along\_width\ to convert detected text points into polygon coordinates for visualization and evaluation.
_ppocr/utils/e2e\utils · high confidence
Added PSE post-processing with automatic compilation
The PSE (Panoptic Scene Understanding) post-processing module is now included in the package, providing a Cython-accelerated implementation for text segmentation. The module automatically compiles the C++ extension on import if it has not been built yet, ensuring the feature works out-of-the-box without manual build steps. This addition supports the PSENet algorithm for detecting text regions in images.
_ppocr/postprocess/pse\postprocess/pse · high confidence
Added TEDS table structure evaluation metric
The \ppstructure/table/table\_metric\ module now includes the Tree Edit Distance Similarity (TEDS) metric for evaluating table structure recognition. This addition introduces a \TEDS\ class that computes similarity scores between predicted and ground-truth HTML table structures by comparing their tree representations, supporting both structure-only and content-aware comparisons. The module also provides parallel processing capabilities via \parallel\_process\ to accelerate batch evaluations, leveraging \ProcessPoolExecutor\ with configurable worker counts and progress tracking.
_ppstructure/table/table\metric · high confidence
Added build scripts for C++ inference toolchain
New shell scripts have been added to the deployment tools to automate the build process for the C++ inference component. The \build.sh\ script configures and compiles the main inference binary using CMake, with configurable paths for PaddlePaddle libraries, OpenCV, CUDA, and cuDNN, and flags for optional features like MKL and TensorRT. Additionally, \build\_opencv.sh\ provides a dedicated script to build OpenCV 4.7.0 from source with specific dependencies (ZLIB, JPEG, PNG, TIFF) and installation settings, ensuring the required image processing library is available for the inference tool.
_deploy/cpp\infer/tools · high confidence
Added data conversion utilities for UniMERNet and LaTeXOCR datasets
New scripts have been added to the \ppocr/utils/formula\_utils\ directory to facilitate the conversion of external formula datasets into formats compatible with PaddleOCR. \math\_txt2pkl.py\ converts LaTeXOCR image and equation text files into a structured pickle format, handling image dimension filtering and relative path normalization. \unimernet\_data\_convert.py\ provides tools to convert UniMERNet and HME100K dataset files into the standard PaddleOCR training and testing label formats, supporting both training and test data modes.
_ppocr/utils/formula\utils · high confidence
Added supplementary test infrastructure for fleet-based training scenarios
This change introduces new files to the \test\_tipc/supplementary/test\_tipc\ directory to support testing of distributed (fleet) training configurations. It adds a \common\_func.sh\ script providing utility functions for parsing parameters and checking status, a \test\_train\_python.sh\ script that orchestrates training runs with various GPU and AMP settings, and several configuration files (\train\_infer\_python\_fleet.txt\, \train\_infer\_python\_FPGM\_fleet.txt\, \train\_infer\_python\_PACT\_fleet.txt\). These configuration files define parameters for the \ch\_PPOCRv2\_det\ model, specifically enabling fleet-based distributed training via \paddle.distributed.launch\ with multi-node IP lists, as well as specific optimization modes like FPGM pruning and PACT quantization.
_test\_tipc/supplementary/test\tipc · high confidence
Android demo initializes Gradle wrapper
The Android demo project now includes a Gradle wrapper configuration, ensuring consistent build environments by pinning the distribution to Gradle 6.5.
_deploy/android\demo/gradle · high confidence
Establishes standard release process and modernizes developer tooling
The repository now includes a formal Release SOP (RELEASING.md) that defines a structured workflow for patch and minor version bumps, requiring official releases to be tagged and published from dedicated \release/X.Y\ branches rather than \main\. This change is accompanied by the introduction of a comprehensive pre-commit configuration (\.pre-commit-config.yaml\) that enforces code quality via Black, Flake8, and clang-format, alongside new \.gitignore\ and \.clang\_format.hook\ files to standardize the development environment. Additionally, the documentation infrastructure has been upgraded with a new \mkdocs.yml\ configuration supporting multi-language navigation and a \mkdocs-ci.yml\ for strict link validation, while the root \README.md\ has been completely rewritten to highlight the latest PaddleOCR-VL and PP-OCRv6 capabilities.
(repo-wide) · high confidence
Expanded loss function library and rotated ROI alignment support
The losses module now supports a significantly wider range of text recognition and detection algorithms, including new loss implementations for models such as ParseQ, UniMERNet, LaTeXOCR, PPFormulaNet, CPPD, and RF-Learning, alongside existing ones like DB, EAST, SAST, and Center Loss. Additionally, a new custom C++/CUDA extension for rotated RoI alignment has been added to support models requiring rotated bounding box processing, and the loss builder has been updated to register all new loss types for configuration-driven instantiation.
ppocr/losses · high confidence
Expanded multilingual and specialized character dictionaries for PP-OCR
The \ppocr/utils/dict\ directory has been significantly expanded with new character-level vocabulary files to support a broader range of languages and specialized recognition tasks. This update adds dictionaries for numerous scripts including Arabic, Belarusian, Bengali, Burmese, Cyrillic, Devanagari, Gujarati, Hebrew, Hindi, Kannada, Kazakh, Korean, Marathi, Nepali, Occitan, Portuguese, Serbian, Samaritan, Syriac, Tamil, and Vietnamese, alongside Latin and English variants. It also introduces specialized dictionaries for LaTeX OCR (\latex\_ocr\_tokenizer.json\, \latex\_symbol\_dict.txt\) and table structure recognition (\table\_dict.txt\). Furthermore, new dictionaries are provided for PP-OCRv5 and PP-OCRv6 models, including specific variants for Arabic, Cyrillic, Devanagari, Greek, English, Slavic, Korean, Latin, Tamil, Telugu, and Thai, ensuring compatibility with the latest model versions.
ppocr/utils/dict · high confidence
Initial release of PaddleOCR DBNet benchmark
This change introduces the PaddleOCR DBNet benchmark, providing a complete implementation of the Differentiable Binarization (DBNet) text detection model using PaddlePaddle. The release includes the core training infrastructure (base trainers and datasets), data loading pipelines with augmentation modules, and configuration files for training on SynthText, ICDAR 2015, and open datasets using various backbones like ResNet-18/50 and ResNest-50. It also provides scripts for training, evaluation, prediction, and model export, along with documentation and licensing.
_benchmark/PaddleOCR\DBNet · high confidence
Introduce PDF-to-Word conversion tool with Windows executable
Adds a new PDF-to-Word conversion utility in the ppstructure/pdf2word directory, providing both a Python script and a standalone Windows executable (pdf2word.exe) for easy, no-environment setup. The tool leverages PP-StructureV2 layout analysis and recovery models to convert PDF documents into Word format, supporting both English and Chinese content with automatic model downloading for the lite version.
ppstructure/pdf2word · high confidence
Introduce PaddleOCR Android demo app
The deploy/android\_demo directory now contains a complete Android application for PaddleOCR, built on PaddleLite v2.10. The app provides a user interface for text detection, text direction classification, and text recognition, supporting six distinct operational modes (such as detection plus recognition, or classification plus recognition). It includes the necessary native C++ components for model inference, image preprocessing, and result post-processing, along with configuration options for CPU thread count, power mode, and model paths.
_deploy/android\demo · high confidence
Introduce PaddleOCR-VL High-Performance Serving (HPS) deployment
Adds a new high-performance serving deployment for the PaddleOCR-VL series (defaulting to PaddleOCR-VL-1.6) that orchestrates a FastAPI gateway, a Triton inference server, and a vLLM-based VLM server via Docker Compose. The gateway handles concurrent request routing with separate limits for inference and non-inference operations, while the pipeline container runs layout parsing and page restructuring models. Configuration is managed through environment variables (e.g., \HPS\_PIPELINE\_NAME\, \HPS\_PADDLEX\_VERSION\) and a \prepare.sh\ script that downloads the corresponding PaddleX HPS SDK and configures the vLLM backend.
_deploy/paddleocr\_vl\docker/hps · high confidence
Introduce PaddleOCR.js browser SDK and demo workspace
Adds the \paddleocr-js\ workspace containing the \@paddleocr/paddleocr-js\ browser SDK (\packages/core\) and a Vite-based demo app (\apps/demo\). The SDK exposes \PaddleOCR.create()\ for main-thread or worker-backed OCR using ONNX Runtime Web and OpenCV.js, with configurable WASM paths and worker transport. The workspace includes monorepo tooling (npm workspaces, ESLint, TypeScript, Vitest) and bilingual documentation (English/Simplified Chinese) covering architecture, development, and release conventions.
paddleocr-js · high confidence
Introduce PaddleOCR.js browser SDK with demo app and PP-OCRv6 support
This change introduces the \@paddleocr/paddleocr-js\ browser SDK, enabling text detection and recognition directly in the frontend using ONNX Runtime Web. The core package provides a \PaddleOCR\ class that supports model selection via \ocrVersion\ (including the new PP-OCRv6\_small and PP-OCRv6\_tiny presets) and allows users to specify custom model assets. It features a \viz\ subpath for visualizing OCR results, supports both main-thread and dedicated Worker execution modes, and exposes runtime metrics. A new demo application (\apps/demo\) is included, providing a UI to select models, choose between WebGPU and WASM backends, adjust detection/recognition thresholds, and visualize the output.
paddleocr-js/apps/demo, paddleocr-js/packages/core · high confidence
Introduce TIPC benchmark training and comprehensive inference test scripts
The test\_tipc directory now includes a new benchmark\_train.sh script that enables dynamic epoch configuration and parameter parsing for training benchmarks, alongside a suite of updated test scripts (test\_inference\_cpp.sh, test\_inference\_python.sh, test\_lite\_arm\_cpp.sh) that standardize the execution of C++ and Python inference across various hardware configurations (CPU, GPU, ARM) and precision modes (fp32, fp16, int8). These changes, supported by common utility functions in common\_func.sh and result comparison logic in compare\_results.py, provide a unified framework for validating model performance and accuracy across the PaddleOCR pipeline.
_test\tipc · high confidence
Introduce centralized metric evaluation module for PaddleOCR
The \ppocr/metrics\ package now provides a unified, factory-based system for evaluating OCR models. A new \build\metric\ function in \\\init\\_.py\ allows users to instantiate specific metric classes (such as \DetMetric\, \RecMetric\, \ClsMetric\, \TableMetric\, \KIEMetric\, and \SRMetric\) via configuration. This change introduces dedicated evaluation logic for text detection (IoU-based), recognition (accuracy and normalized edit distance), classification, table structure, key information extraction, and super-resolution, along with support for end-to-end and distillation scenarios.
ppocr/metrics · high confidence
Introduce langchain-paddleocr integration with PaddleOCR-VL loader
This change adds the \langchain-paddleocr\ package, providing a new \PaddleOCRVLLoader\ document loader that integrates Baidu's PaddleOCR-VL models (such as PaddleOCR-VL, 1.5, and 1.6) into the LangChain ecosystem. Users can now extract text and layout information from PDF and image files via the PaddleOCR SDK, supporting local files and remote URLs. The implementation includes unit and integration tests, a Makefile for development workflows, and Python 3.10–3.14 support.
langchain-paddleocr · high confidence
Introduce layout recovery module for converting documents to editable Word and Markdown files
The \ppstructure/recovery\ module has been added to restore the layout of images and PDFs into editable Word (.docx) or Markdown files. It supports two methods: a standard PDF parse using the \pdf2docx\ library for better handling of non-paper documents, and an image-format PDF parse that combines layout analysis, table recognition, and OCR to better recover images, tables, and titles for paper-like documents. The module includes documentation, installation instructions, and Python scripts (\recovery\_to\_doc.py\, \recovery\_to\_markdown.py\) to handle the conversion process.
ppstructure/recovery · high confidence
Introduce modular MCP server with multi-provider inference support
The \mcp\_server\ directory now contains a new, structured Python package (\paddleocr\_mcp\) that implements the Model Context Protocol (MCP) server for PaddleOCR. This release introduces a factory-based architecture that supports three distinct inference providers: local execution (using \PaddleOCR\, \PaddleOCRVL\, and \PPStructureV3\), Baidu AI Studio cloud APIs, and Baidu Qianfan cloud APIs. The server exposes command-line arguments and environment variables to configure the model, provider, and connection details (such as API keys and base URLs), and includes specific support for PP-OCRv5, PP-OCRv6, PaddleOCR-VL, and PP-StructureV3 models across these backends.
_mcp\server · high confidence
Introduce modular inference scripts for PaddleOCR components
The \tools/infer\ directory now provides dedicated, standalone Python scripts for each stage of the OCR pipeline: \predict\_det.py\ for text detection, \predict\_rec.py\ for text recognition, \predict\_cls.py\ for text line angle classification, \predict\_sr.py\ for super-resolution, and \predict\_e2e.py\ for end-to-end algorithms like PGNet. These scripts replace the previous monolithic entry point, allowing users to run individual model stages independently. The \predict\_system.py\ script remains as the orchestrator that chains these components together for full pipeline inference.
tools/infer · high confidence
Introduce multi-logger support with WandbLogger integration
The logging infrastructure in PaddleOCR has been refactored to support multiple logging backends simultaneously. A new \Loggers\ class acts as a composite wrapper, allowing users to enable several loggers (such as the newly added \WandbLogger\) at once. The \WandbLogger\ implementation handles metrics logging and model artifact uploads to Weights & Biases, while a \BaseLogger\ abstract class defines the standard interface for future logger integrations.
ppocr/utils/loggers · high confidence
Introduce static graph training support for OCR architectures
The architecture module now supports static graph (PaddlePaddle \to\_static\) training for specific models. A new \apply\_to\static\ function in \\\init\\_.py\ converts supported models (DB, SVTR\_LCNet, TableMaster, LayoutXLM, SLANet, SVTR, SVTR\_HGNet, LaTeXOCR, UniMERNet, PP-FormulaNet-S/L) to static graphs by defining precise \InputSpec\ shapes for each algorithm. The \BaseModel\ and \DistillationModel\ classes have been refactored to support this mode, with \BaseModel\ handling the standard transform-backbone-neck-head pipeline and \DistillationModel\ managing multiple sub-models for knowledge distillation.
ppocr/modeling/architectures · high confidence
Introduce table recognition module with SLANet and TableMaster support
The \ppstructure/table\ directory now contains a complete table recognition capability, allowing users to extract table structures and cell content from images. This feature integrates a pipeline of text detection (DB), text recognition (CRNN), and table structure prediction (SLANet or TableMaster) to generate HTML representations of tables. Users can run inference via \predict\_table.py\ to produce Excel and HTML outputs, evaluate model performance using the TEDS metric via \eval\_table.py\, and convert ground-truth annotations to HTML format using \convert\_label2html.py\. The module supports both Chinese and English models and includes documentation for quick starts, training, and evaluation.
ppstructure/table · high confidence
Introduction of a unified post-processing module for PaddleOCR
The \ppocr/postprocess\ package has been introduced to centralize and standardize the post-processing logic for text detection, recognition, and classification models. This change provides a factory function, \build\_post\_process\, which dynamically instantiates specific processors based on configuration, supporting a wide range of algorithms including DB, EAST, SAST, FCE, PGNet, and various recognition decoders (CTC, SAR, NRTR, ViTSTR, ABINet, ParseQ, etc.). For users, this means a more consistent and extensible API for configuring post-processing steps in OCR pipelines, with support for diverse output formats such as quadrilaterals, polygons, and word-level bounding boxes.
ppocr/postprocess · high confidence
Introduction of new text detection and recognition head modules
The \ppocr/modeling/heads\ package now includes implementations for several new text detection and recognition algorithms. Detection capabilities have been expanded with the addition of DBHead (with PFHeadLocal), EASTHead, SASTHead, FCEHead, PSEHead, CT\_Head, PGHead, and DRRGHead. Recognition support has been extended to include CTCHead, AttentionHead, SRNHead, Transformer (NRTR), SARHead, AsterHead, PRENHead, MultiHead, SPINAttentionHead, ABINetHead, RobustScannerHead, VLHead, RFLHead, CANHead, LaTeXOCRHead, SATRNHead, ParseQHead, CPPDHead, UniMERNetHead, and PPFormulaNet\_Head. Additionally, new classification (ClsHead) and key information extraction (SDMGRHead) heads are available, along with table recognition heads (TableAttentionHead, SLAHead, TableMasterHead). These modules are registered in the central \build\_head\ factory function, allowing users to configure and utilize these specific architectures via their configuration files.
ppocr/modeling/heads · high confidence
New C++ OCR inference API with modular pipeline support
The \deploy/cpp\_infer\ directory now provides a complete C++ inference API for PaddleOCR, replacing previous ad-hoc scripts. It introduces a structured CLI (\cli.cc\) and a modular architecture with distinct model wrappers (text detection, recognition, unwarping, and orientation classification) and pipeline orchestrators (\doc\_preprocessor\, \full OCR\). The build system (\CMakeLists.txt\) is updated to support static linking, GPU acceleration, and MKL/OneDNN optimizations, while ensuring Windows compatibility through corrected runtime library flags (/MT).
_deploy/cpp\infer · high confidence
New C++ inference modules for image classification, unwarping, and text detection
This change introduces three new C++ inference modules in the deploy/cpp\_infer directory. The image classification module adds a predictor that supports configurable pre-processing (resize, crop, normalize) and post-processing (top-k ranking), with results exportable to images and JSON. The image unwarping module adds a predictor using a DocTR-based post-processor to output unwarped images, also savable to image and JSON formats. The text detection module adds a predictor with a DBPostProcess step to detect text polygons and scores, supporting various resize strategies (limit side length, resize long, fixed shape) and configurable thresholds. All modules follow a consistent structure with predictor, processor, and result components, enabling batched inference and standardized output handling.
_deploy/cpp\_infer/src/modules/image\_classification, deploy/cpp\_infer/src/modules/image\_unwarping, deploy/cpp\_infer/src/modules/text\detection · high confidence
New C++ text recognition module for local deployment
A new text recognition module has been added to the C++ inference deployment pipeline (deploy/cpp\_infer/src/modules/text\_recognition). This includes a predictor that handles image preprocessing (resize, normalization, batching) and post-processing (CTC label decoding) to extract text and confidence scores from images. The module supports configurable languages, OCR versions, and visualization font directories, and can output results as JSON or annotated images (when OpenCV FreeType is available).
_deploy/cpp\_infer/src/modules/text\recognition · high confidence
New Go SDK for PaddleOCR official API services
Users can now interact with hosted PaddleOCR services via a new official Go client. This SDK allows submitting OCR and document parsing jobs (supporting models like PP-OCRv6, PP-StructureV3, and PaddleOCR-VL-1.6) using file URLs or local file paths. It provides typed result parsing that preserves raw API fields and metadata, handles asynchronous job polling with configurable timeouts, and includes utilities to download result resources like OCR images and parsed document exports.
_api\_sdk, api\sdk/typescript · high confidence
New Key Information Extraction (KIE) module with SER and RE support
PP-Structure introduces a new Key Information Extraction (KIE) module that supports Semantic Entity Recognition (SER) and Relation Extraction (RE) tasks using multimodal models like VI-LayoutXLM and LayoutXLM. This addition includes Python inference scripts (\predict\_kie\_token\_ser.py\ and \predict\_kie\_token\_ser\_re.py\) for end-to-end prediction, comprehensive documentation (\README.md\, \how\_to\_do\_kie.md\) detailing usage on datasets like XFUND, and specific environment requirements (Paddle \< 2.6, PaddleNLP \< 2.6).
ppstructure/kie · high confidence
New PP-OCR model quantization tools and documentation
Added a new \deploy/slim/quantization\ directory containing Python scripts (\quant.py\, \quant\_kl.py\, \export\_model.py\) and documentation (\README.md\, \README\_en.md\) for compressing PP-OCR models. These tools enable Quantization-Aware Training (QAT) and post-training static quantization (KL-based) using PaddleSlim, allowing users to convert FP32 models to INT8 for faster inference on mobile or edge devices.
deploy/slim/quantization · high confidence
New PaddleHub serving modules for PP-OCRv3 text recognition and full OCR system
Added new \ocr\_rec\ and \ocr\_system\ PaddleHub serving modules in \deploy/hubserving\ that wrap the PP-OCRv3 inference pipelines. The \ocr\_rec\ module exposes a text-recognition service (port 8867) using the CRNN algorithm and the \ch\_PP-OCRv3\_rec\_infer\ model, accepting base64-encoded images and returning recognized text with confidence scores. The \ocr\_system\ module exposes a full OCR service (port 8868) that combines DB text detection, CRNN recognition, and angle classification (using \ch\_ppocr\_mobile\_v2.0\_cls\_infer\) to return detected text regions, recognized text, and confidence scores. Both modules support GPU acceleration via \use\_gpu\, MKL-DNN via \enable\_mkldnn\, and are configured with specific batch sizes, image shapes, and model directories defined in their respective \params.py\ files.
_deploy/hubserving/ocr\system · high confidence
New Python SDK for PaddleOCR Cloud API
A new Python API client has been added to the PaddleOCR package, providing both synchronous and asynchronous interfaces to interact with the PaddleOCR cloud service. Users can now perform OCR and document parsing tasks by submitting jobs via file URL or local file path, polling for completion, and retrieving structured results. The SDK includes a CLI subcommand (\paddleocr api\) for direct command-line usage, supports configurable timeouts and polling strategies, and handles resource downloading for result images. It introduces specific error classes for authentication, rate limiting, and job failures, and defaults to the PP-OCRv6 model for OCR tasks.
_paddleocr/\_api\client · high confidence
New VQA data augmentation and tokenization utilities
This change introduces new modules for Visual Question Answering (VQA) data processing within PaddleOCR. The \augment.py\ module adds a utility function \order\_by\tbyx\ to sort OCR bounding boxes by their vertical and horizontal coordinates, ensuring consistent text ordering. The \\\init\\_.py\ module exposes several tokenization classes (\VQATokenPad\, \VQASerTokenChunk\, \VQAReTokenChunk\, \VQAReTokenRelation\, \TensorizeEntitiesRelations\) that facilitate the conversion of VQA data into tensor formats suitable for model training.
ppocr/data/imaug/vqa · high confidence
New VQA tokenization and data augmentation pipeline
Added a new set of image augmentation and tokenization classes for Visual Question Answering (VQA) tasks within the PP-OCR data processing pipeline. This includes \VQASerTokenChunk\ and \VQAReTokenChunk\ for splitting sequences into chunks, \VQATokenPad\ for padding inputs to a maximum sequence length, \VQAReTokenRelation\ for building entity-relation mappings, and \TensorizeEntitiesRelations\ for converting entity and relation data into fixed-shape tensors. These components enable the training system to handle VQA-specific data structures, including entity extraction and relation classification, by preparing tokenized inputs with proper padding, masking, and relation indexing.
ppocr/data/imaug/vqa/token · high confidence
New backbone registry and detection/recognition model implementations
The \ppocr/modeling/backbones\ module has been restructured to introduce a centralized \build\backbone\ factory in \\\init\\.py\ that dynamically instantiates backbone networks based on the model type (detection, recognition, KIE, table, or end-to-end). This change adds support for a wide range of new and updated backbones, including MobileNetV3, ResNet variants (standard, VD, SAST), PPLCNet, PPLCNetV2, PPHGNet, ViT, SVTR, and specialized models like DonutSwin and ViTParseQ for recognition, as well as UNet-based backbones for KIE. The implementation files for these backbones are now explicitly defined in the \det\\ and \rec\_\ sub-modules, enabling users to select specific architectures via configuration for text detection and recognition tasks.
ppocr/modeling/backbones · high confidence
New build pipelines and configuration for PaddleOCR-VL 1.6 with multi-backend support
The PaddleOCR-VL Docker deployment now includes dedicated build scripts (\build\_pipeline.sh\ and \build\_vlm.sh\) and YAML configuration files for PaddleOCR-VL 1.6. These changes introduce support for multiple hardware accelerators (including Intel, AMD, MetaX, and NVIDIA SM120) and allow users to choose between \vllm\ and \fastdeploy\ backends for the VL recognition server. The new pipeline configurations define the document processing workflow, including layout detection with PP-DocLayoutV3 and document preprocessing, enabling more flexible and optimized deployment options for different infrastructure requirements.
_deploy/paddleocr\_vl\docker · high confidence
New feature recognition neck modules added to PP-OCR
The \ppocr/modeling/necks\ module now includes new neck architectures for feature recognition, specifically \RFAdaptor\ (from \rf\_adaptor.py\) and \PRENFPN\ (from \pren\fpn.py\). These are registered in the \\\init\\_.py\ factory and can be selected via configuration to enhance text recognition capabilities.
ppocr/modeling/necks · high confidence
New image augmentation and preprocessing operators for text detection and recognition
The \ppocr/data/imaug\ module now includes a suite of new data augmentation and preprocessing classes to support advanced training pipelines. This adds a \ColorJitter\ wrapper for Paddle's color transforms, a \CopyPaste\ operator for synthetic data generation by pasting external text instances, and specialized processing modules for detection algorithms including EAST, SAST, PGNet, CT, DRRG, and FCENet. Additionally, new augmentation strategies such as \RandomScaling\, \RandomCropFlip\, and \RandAugment\ are available, along with target generation logic for DRRG and FCENet, expanding the variety of image transformations and ground-truth preparation methods available to users.
ppocr/data/imaug · high confidence
New image transformation modules for scene text recognition
Added a new \ppocr/modeling/transforms\ package providing geometric and super-resolution image transformations. This includes \TPS\ (Thin Plate Spline) and \STN\_ON\ (Spatial Transformer Network) for spatial alignment, \TSRN\ and \TBSRN\ for text super-resolution, and \GA\_SPIN\ (Geometric-Absorbed Structure-Preserving Inner Offset Network) for structure-preserving transformation. A central \build\_transform\ factory function allows users to instantiate these modules by name (e.g., 'TPS', 'STN\_ON', 'GA\_SPIN', 'TSRN', 'TBSRN') via configuration.
ppocr/modeling/transforms · high confidence
New inference and evaluation tools for PaddleOCR
The \tools\ directory now includes a comprehensive set of standalone scripts for running inference and evaluation across all supported PaddleOCR tasks. Users can now directly execute \tools/infer\_det.py\, \tools/infer\_rec.py\, \tools/infer\_cls.py\, \tools/infer\_e2e.py\, \tools/infer\_sr.py\, \tools/infer\_kie.py\, \tools/infer\_kie\_token\_ser.py\, and \tools/infer\_kie\_token\_ser\_re.py\ to perform detection, recognition, classification, end-to-end, super-resolution, and key information extraction. Additionally, \tools/export\_model.py\ and \tools/export\_center.py\ provide dedicated entry points for exporting trained models to inference formats and extracting feature centers, while \tools/eval.py\ handles model evaluation. A new \tools/check\_docs\_github\_links.py\ utility is also added to validate documentation links against forbidden GitHub references.
tools · high confidence
New modular PaddleOCR inference package with dedicated CLI subcommands
The \paddleocr/\_models\ directory has been restructured into a new inference package that exposes individual document-analysis capabilities as distinct, reusable classes and CLI subcommands. Users can now invoke specialized tasks directly via the command line—including chart parsing, document VLM queries, formula recognition, layout detection, seal and text detection, table cell detection, table classification, table structure recognition, text detection, text image unwarping, text recognition, and text line orientation classification—each backed by a specific default model (e.g., PP-Chart2Table, PP-DocBee2-3B, PP-OCRv6\_medium\_det/rec). The implementation introduces a base \PaddleXPredictorWrapper\ that delegates to PaddleX predictors and a \PredictorCLISubcommandExecutor\ that registers CLI subparsers, while mixins and base classes (e.g., \TextDetectionMixin\, \BaseDocVLM\, \ImageClassification\, \ObjectDetection\) standardize argument handling and inference execution across modules.
_paddleocr/\models · high confidence
New modular optimizer and learning-rate scheduling system
The \ppocr/optimizer\ package has been introduced to centralize and standardize training configuration. It provides a \build\_optimizer\ entry point that constructs the optimizer, learning-rate scheduler, and weight-decay scheduler from a unified config. New learning-rate schedules are available, including Linear, Cosine, LinearWarmupCosine, Step, CyclicalCosineDecay, OneCycleDecay, and TwoStepCosineDecay, with support for warmup. Weight decay regularization now includes L1, L2, and a new CosineL2Decay that anneals the decay coefficient over training steps with optional warmup. Gradient clipping is supported via \clip\_norm\ (L2 norm) and \clip\_norm\_global\ (global norm). The Adam optimizer also supports grouped learning rates for specific model architectures.
ppocr/optimizer · high confidence
Behavioural changes
2 commits (0 fixes) modifying doc
A change to existing behaviour in doc — 2 commits, 19 files.
doc · medium confidence · unverified
Deprecation of PP-StructureV2 and introduction of PP-StructureV3 migration guidance
The \ppstructure\ directory now contains a README explicitly stating that the current codebase represents the second generation (PP-StructureV2) and is scheduled for removal as maintenance is discontinued. Users are directed to upgrade to the third generation (PP-StructureV3) via the integrated PaddleOCR wheel package for enhanced document analysis capabilities, and to use the PP-ChatOCRv4 model in PaddleOCR 3.x for key information extraction tasks previously handled by the second-generation KIE capability.
ppstructure · high confidence
Document expiration warnings and Giscus comment integration
The documentation site now warns users when viewing outdated content by displaying a notice if a page has not been updated within a configurable number of days (defaulting to 365), and it enables community discussions via Giscus comments that automatically sync with the site's color theme.
overrides · high confidence
New data loading infrastructure with secure pickle handling and URL prefetching
The \ppocr/data\ module has been restructured to provide a unified data loading pipeline. This includes a new \build\_dataloader\ function that dynamically instantiates datasets (such as \SimpleDataSet\, \LMDBDataSet\, \PGDataSet\, \PubTabDataSet\, and \LaTeXOCRDataSet\) and configures PaddlePaddle \DataLoader\ instances with support for distributed sampling and shared memory. To address security vulnerabilities, \LMDBDataSet\ and \LaTeXOCRDataSet\ now use a restricted pickle unpickler that only allows safe built-in types, preventing arbitrary code execution during dataset loading. Additionally, \SimpleDataSet\ introduces a per-worker URL prefetch cache to asynchronously download remote images, improving I/O performance for datasets hosted on HTTP/HTTPS endpoints.
ppocr/data · high confidence
New modular pipeline architecture with layout parsing fixes
PaddleOCR introduces a new internal pipeline structure under the \_pipelines module, replacing the previous monolithic design with distinct, composable components such as DocPreprocessor, DocUnderstanding, FormulaRecognitionPipeline, PaddleOCR, PaddleOCRVL, PPChatOCRv4Doc, PPDocTranslation, PPStructureV3, SealRecognition, and TableRecognitionPipelineV2. This modularization allows users to leverage specialized pipelines for specific tasks like document understanding, formula recognition, and table processing, while maintaining backward compatibility through the main PaddleOCR class. Additionally, the update includes critical fixes for layout parsing issues, specifically addressing integer overflow errors in bounding box calculations during document unwarping and handling empty bounding box lists that previously caused crashes.
_paddleocr/\pipelines · high confidence
PaddleOCR 3.5 introduces a new inference package with expanded capabilities
The PaddleOCR package has been updated to version 3.5, introducing a new inference package structure that significantly expands functionality. Users now have access to new pipelines and models including PaddleOCRVL, PPChatOCRv4Doc, PPDocTranslation, and PPStructureV3, alongside specialized models for chart parsing, formula recognition, and seal text detection. The update also adds a new \doc2md\_convert\ feature to convert office documents (docx/xlsx/pptx) to Markdown, and introduces new CLI commands for installing HPI and GenAI server dependencies. Additionally, the package now supports the CINN compiler flag and exposes new API clients (PaddleOCRClient, AsyncPaddleOCRClient) with corresponding error handling classes.
paddleocr · high confidence
Test coverage
Added end-to-end tests for the web OCR model; Added test suite for PaddleOCR API client and model inference.
Dependencies
New multi-language API SDKs and updated platform-specific dependencies
This release introduces official API SDKs for Go and TypeScript, alongside updated dependency configurations for Android, iOS, and Python environments. The new Go SDK is defined in \api\_sdk/go/go.mod\ (Go 1.21), and the TypeScript SDK (\@paddleocr/api-sdk\) is available in \api\_sdk/typescript/\ (Node \>=18, TypeScript ^5.3.0). For mobile deployment, the Android demo now uses a modern Gradle Kotlin DSL with AndroidX, Compose, and ONNX Runtime 1.21.1, while the iOS demo specifies ONNX Runtime 1.24.3 and OpenCV 4.3.0 via CocoaPods. Python dependencies have been standardized across \pyproject.toml\ and various \requirements.txt\ files, pinning \paddlex\ to \>=3.7.0, \fastmcp\ to \>=2.0.0 for the MCP server, and \langchain-core\ to \>=1.2.5 for the LangChain integration.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Score
- CAI 48 → 68 (+19.1)
- Rubric changed (rubric-2026.08.15 → rubric-2026.09.15) — scores are not directly comparable.
Lenses
- Code Health 85 → 85 (-0.5)
- Architecture 97 → 95 (-1.4)
- Maturity 63 → 75 (+11.7)
- Readiness 38 → 70 (+32.0)
- Security 41 → 59 (+17.6)
Resolved (99)
- Coverage not measured — test suite did not build
- Critical CVE: [GHSA redacted] (paddleocr-js/package-lock.json)
- Critical CVE: [GHSA redacted] (paddleocr-js/package-lock.json)
- Dimension evaluation failed
- High CVE: [GHSA redacted] (langchain-paddleocr/uv.lock)
- High CVE: [GHSA redacted] (paddleocr-js/package-lock.json)
- High CVE: [GHSA redacted] (langchain-paddleocr/uv.lock)
- High CVE: [GHSA redacted] (paddleocr-js/package-lock.json)
- High CVE: [GHSA redacted] (paddleocr-js/package-lock.json)
- High CVE: [GHSA redacted] (langchain-paddleocr/uv.lock)
- High CVE: [GHSA redacted] (langchain-paddleocr/uv.lock)
- High CVE: [GHSA redacted] (paddleocr-js/package-lock.json)
- High CVE: [GHSA redacted] (paddleocr-js/package-lock.json)
- High CVE: [GHSA redacted] (paddleocr-js/package-lock.json)
- High CVE: [GHSA redacted] (langchain-paddleocr/uv.lock)
- High CVE: [GHSA redacted] (paddleocr-js/package-lock.json)
- High CVE: [GHSA redacted] (langchain-paddleocr/uv.lock)
- High CVE: [GHSA redacted] (langchain-paddleocr/uv.lock)
- High CVE: [GHSA redacted] (paddleocr-js/package-lock.json)
- High CVE: [GHSA redacted] (langchain-paddleocr/uv.lock)
- …and 79 more
New (1330)
- APP_Image2Doc.downloadModels (cognitive 21) (ppstructure/pdf2word/pdf2word.py)
- Adam.call (cognitive 29) (ppocr/optimizer/optimizer.py)
- Attention.forward (cognitive 19) (ppocr/modeling/heads/rec_latexocr_head.py)
- Attention.forward (cyclomatic 17) (ppocr/modeling/heads/rec_latexocr_head.py)
- AttentionLayers.init (cognitive 24) (ppocr/modeling/heads/rec_latexocr_head.py)
- AttentionLayers.init (cyclomatic 23) (ppocr/modeling/heads/rec_latexocr_head.py)
- AttentionLayers.forward (cognitive 18) (ppocr/modeling/heads/rec_latexocr_head.py)
- AttentionLayers.forward (cyclomatic 16) (ppocr/modeling/heads/rec_latexocr_head.py)
- AttnLabelDecode.decode (cognitive 21) (ppocr/postprocess/rec_postprocess.py)
- AugmenterBuilder.build (cognitive 16) (ppocr/data/imaug/iaa_augment.py)
- AugmenterBuilder.map_arguments (cognitive 19) (ppocr/data/imaug/iaa_augment.py)
- BaseModel.forward (cognitive 18) (ppocr/modeling/architectures/base_model.py)
- BaseRecLabelDecode.get_word_info (cognitive 24) (ppocr/postprocess/rec_postprocess.py)
- BaseRecLabelDecode.get_word_info (cyclomatic 16) (ppocr/postprocess/rec_postprocess.py)
- CTCDecoder.decode (cognitive 23) (deploy/ppocr-android/ppocr-sdk/src/main/java/com/paddle/ocr/postprocess/CTCDecoder.kt)
- CTPostProcess.call (cognitive 28) (ppocr/postprocess/ct_postprocess.py)
- CVRandomAffine.init (cognitive 23) (ppocr/data/imaug/abinet_aug.py)
- Change coupling clique: init.py, init.py, init.py, init.py, program.py (ppocr/losses/init.py)
- Change coupling: recovery_to_doc.py ↔ utility.py (ppstructure/recovery/recovery_to_doc.py)
- Client.SaveResource (cognitive 31) (api_sdk/go/resource.go)
- …and 1310 more
Changes since last survey
- 2 commits — 2 feature/other, 0 fixes
By area
- (root) — 1 commit
- docs/version2.x — 1 commit
Notable commits
- change: docs: deduplicate images and streamline repository cloning (#18359)
- change: docs: organize research papers and unify multilingual citations (#18357)
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
PaddlePaddle/PaddleOCR was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 26 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit dab3fe35379033fdcb2d0e9572fac0b36c9a9ebf — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-a15879f6f801.