docling-project/docling
63.5
Weak · 26 September 2026
74.2k
lines of production code
Python
primary language
4
measurements over time
What this system is
Docling is a modular document processing library that converts a wide variety of input formats—including PDFs, Office documents, XML, and audio/video—into structured, machine-readable outputs. It employs a pluggable pipeline architecture that integrates local and remote AI models for layout detection, OCR, table structure recognition, and vision-language model extraction. The system supports both local execution via a CLI and SDK, as well as remote interaction through a dedicated service client for scalable, asynchronous document conversion.
How it got here
2024–2025 — Docling v2 architecture and modular expansion
24 changes.
This period centered on the major v2 overhaul of Docling, replacing the monolithic converter with a modular, plugin-based pipeline architecture supporting diverse media types like video and audio. The project expanded its ecosystem by introducing new document backends for formats such as AFP, JATS, and EBCDIC, while restructuring the package distribution into slim and client variants. Concurrently, the team enhanced processing capabilities with experimental threaded VLM pipelines, hardware acceleration utilities, and a dedicated CLI for model and document management.
2026 — VLM engine unification and modular architecture
41 changes.
This period focused on refactoring the document processing pipeline into a modular, pluggable architecture centered on a unified Vision-Language Model (VLM) inference engine. The work standardized backends for OCR, layout, and extraction tasks, enabling support for diverse hardware accelerators and remote APIs while introducing new Granite Vision 4.1 capabilities. Concurrently, the project expanded format support to include Apple iWork documents and strengthened the service layer with a dedicated client SDK and comprehensive regression testing.
Features
Add Apple iWork document support (Pages, Keynote, Numbers)
Users can now load and extract content from Apple iWork documents, including Pages (.pages), Keynote (.key), and Numbers (.numbers) files. The backend supports both modern iWork archives (2013+) and legacy XML-based formats (iWork '09 and earlier), recovering text, tables, images, headers, footers, footnotes, and comments.
docling/backend/iwork · high confidence
Add Docling JSON document ingestion backend
Users can now ingest documents in the Docling JSON format. A new backend implementation has been added to handle both file paths and byte streams, ensuring compatibility by stripping UTF-8 BOM characters during decoding.
docling/backend/json · high confidence
Experimental TableCropsLayoutModel for table-only images
An experimental layout model, TableCropsLayoutModel, has been added to the experimental models module. This model treats the entire page as a single table cluster, which is useful when processing images that contain only table crops. It includes configuration options to control cell assignment and handling of empty clusters.
docling/experimental/models · high confidence
Introduce Docling service client SDK
This change introduces a new client SDK for interacting with the docling-serve backend, providing both synchronous and asynchronous clients. Users can now submit conversion tasks, monitor their status via polling or WebSocket, and retrieve results. The SDK supports various source types (local files, HTTP URLs, S3, Azure Blob, Google Drive) and output targets (presigned URLs, S3, Azure Blob, Google Drive, In-Body, Zip). It includes robust error handling for service errors, usage limits, timeouts, and task failures, along with configurable concurrency and retry policies.
_docling/service\client · high confidence
Introduce pluggable VLM runtime system for document conversion
The VLM conversion stage now uses a new pluggable runtime architecture that delegates model inference to specific engines (such as vLLM, MinerU, or OpenAI) via a factory pattern. This change enables support for new models like MinerU 2.5 Pro and exposes OpenAI logprobs as generated tokens, while ensuring that runtime generation settings and metadata are correctly propagated from the engine to the final prediction.
_docling/models/stages/vlm\convert · high confidence
Introduce pluggable VLM runtime system for picture descriptions
The picture description stage now supports a new pluggable VLM engine system alongside the existing direct Transformers implementation. Users can now configure picture descriptions to use different backends (such as Transformers, MLX, or remote APIs) via a unified engine interface, allowing for greater flexibility in model selection and deployment. The new \PictureDescriptionVlmEngineModel\ handles engine initialization and batch prediction, while \PictureDescriptionApiModel\ provides a dedicated path for remote API-based descriptions with concurrency control. The legacy \PictureDescriptionVlmModel\ remains available for direct Hugging Face model usage.
_docling/models/stages/picture\description · high confidence
Introduce service-layer data models for document conversion and chunking
This change introduces a new set of Pydantic data models in the \docling.datamodel.service\ package that define the request and response schemas for the Docling service API. Users can now interact with the service using structured models for sources (S3, Azure Blob, Google Cloud Storage, Google Drive, HTTP, and file uploads) and targets (in-body, ZIP, S3, Azure Blob, Google Cloud Storage, Google Drive, PUT, and presigned URLs). The models also expose configuration options for document chunking (hierarchical and hybrid chunkers with tokenization and image handling settings), conversion options (including VLM and picture description settings), and progress callbacks. Response models include detailed document results with confidence scores, artifact references, and task status tracking, enabling more robust and type-safe integration with the Docling service.
docling/datamodel/service · high confidence
Introduce the Docling CLI for document conversion and model management
This change introduces the \docling\ command-line interface, providing users with a local tool to convert documents into various formats (including Markdown, JSON, HTML, LaTeX, and DCLX) and manage AI model downloads. The CLI supports features such as page range selection, OCR language configuration, and remote conversion via a \docling-serve\ service, while ensuring compatibility with lightweight installations by deferring heavy model imports.
docling/cli · high confidence
Introduces declarative configuration models for all pipeline stages and backends
The datamodel package has been refactored to replace the previous monolithic configuration with a structured set of Pydantic models. This change introduces dedicated option classes for every pipeline stage—including AcceleratorOptions for hardware selection, BackendOptions for input formats (HTML, CSV, Markdown, etc.), and stage-specific models for Layout, Table, Picture Classification, Chart Extraction, and ASR. It also establishes a unified engine system with specific runtime options for Transformers, ONNX, and remote KServe v2 APIs, allowing users to configure inference backends and model presets in a consistent, type-safe manner.
docling/datamodel · high confidence
Introduces pluggable VLM inference engine system with auto-selection and API support
Docling now supports a modular architecture for Vision-Language Model (VLM) inference, allowing users to choose between local backends (Transformers, MLX for Apple Silicon, and vLLM for high-throughput CUDA/XPU serving) or remote services via an OpenAI-compatible API. An auto-selecting engine is available to automatically pick the best local backend based on the operating system and hardware (e.g., MLX on macOS with Apple Silicon, vLLM on Linux/Windows with CUDA), while also enforcing minimum version checks for backing libraries to prevent runtime failures. This change provides a unified interface for VLM generation, including support for custom stopping criteria, temperature control, and artifact path management.
_docling/models/inference\engines/vlm · high confidence
Introduces shared KServe v2 and HuggingFace vision inference utilities
This change adds a new common module for inference engines, providing a transport-agnostic KServe v2 client protocol and concrete HTTP and gRPC implementations, alongside shared HuggingFace vision model loading utilities. Users can now interact with remote KServe v2 endpoints (both REST and gRPC) using a unified interface and leverage standardized helpers for loading HuggingFace vision models and processors.
_docling/models/inference\engines/common · high confidence
Introduction of Hybrid Chunker API
The docling/chunking module now exposes the HybridChunker alongside the existing HierarchicalChunker, allowing users to utilize a hybrid approach for document chunking in addition to the hierarchical method previously available.
docling/chunking · high confidence
Introduction of pluggable model and engine defaults
The system now exposes a structured registry of available inference backends through new default configuration modules. Users can now access predefined lists of supported engines for OCR (including AutoOCR, EasyOCR, Tesseract, and remote KServe/Nemotron options), picture description (VLM engines and API models), layout analysis (including the new experimental TableCropsLayoutModel), and table structure extraction (including TableFormer v2 and GraniteVision). This change establishes the foundation for a plugin-based architecture where these specific model implementations are registered and selectable via a factory system.
docling/models/plugins · high confidence
New LaTeX conversion module for DOCX equations
Added a new \docling/backend/docx/latex\ module containing \latex\_dict.py\ and \omml.py\ to convert Office Math Markup Language (OMML) from Word documents into LaTeX format. This includes a comprehensive dictionary of Unicode-to-LaTeX mappings for symbols, accents, and operators, along with an XML parser that processes OMML elements like fractions, limits, and integrals into valid LaTeX strings.
docling/backend/docx/latex · high confidence
New VLM-based and TableFormer v2 table structure models
The table structure detection stage now supports two new model backends alongside the existing TableFormer: a Granite Vision 4.1-based model for VLM-driven table extraction and a TableFormer v2 model. Users can now choose between these newer architectures via pipeline options, which may offer improved accuracy or performance for complex table layouts compared to the legacy TableFormer implementation.
_docling/models/stages/table\structure · high confidence
New XML-based document backends for JATS, USPTO, XBRL, and DocLang
Users can now process Journal Article Tag Suite (JATS) XML, USPTO patent documents, XBRL financial reports, and DocLang archives. The JATS backend parses structured article content including abstracts, footnotes, and inline formulas, with secure XML parsing to prevent XXE attacks. The USPTO backend handles patent grants and applications using defusedxml for security. The XBRL backend leverages Arelle to extract structured financial data and taxonomy relationships. The DocLang backends support both native DocLang XML and compressed DocLang archives (.dclx), enabling round-trip compatibility with DoclingDocument serialization.
docling/backend/xml · high confidence
New document backends for AFP, AsciiDoc, Box Note, CSV, EBCDIC, and Email formats
This change introduces several new document backends to the library, expanding supported input formats. The AFP backend parses MO:DCA structured fields and extracts text from PTOCA control sequences. The AsciiDoc backend converts AsciiDoc markup into structured documents, handling titles, lists, tables, and code blocks. The Box Note backend reads JSON-based Box Notes and maps them to the document model. The CSV backend parses comma-separated values into tables, including dialect detection and BOM handling. The EBCDIC backend decodes mainframe EBCDIC data files using configurable COBOL layouts. The Email backend processes .eml and .msg files, extracting headers, bodies, and attachments. Additionally, the docling-parse PDF backends (v2 and v4) are now deprecated in favor of the threaded docling-parse backend.
docling/backend · high confidence
New generation utilities and repetition-based stopping criteria
Added a new \docling.models.utils\ package containing \generation\_utils.py\, \hf\_model\_download.py\, and \hf\_stopping\_criteria.py\. The \generation\_utils\ module provides a \build\_generation\_config\ helper to manage Hugging Face generation parameters and introduces a \DocTagsRepetitionStopper\ that detects and halts generation when it encounters repetitive, consecutive document tags with identical text and stable or duplicate coordinates. The \hf\_stopping\_criteria\ module includes an \HFStoppingCriteriaWrapper\ to integrate these custom stopping logics with Hugging Face Transformers, while \hf\_model\_download.py\ adds a utility for downloading models from Hugging Face with improved logging to distinguish between cache hits and network downloads.
docling/models/utils · high confidence
New modular OCR engine architecture with NVIDIA Nemotron support
The OCR stage has been reorganized into a dedicated module (\docling/models/stages/ocr\) containing distinct implementations for each supported engine. This update introduces a new NVIDIA Nemotron OCR model (\NemotronOcrModel\) for Linux environments, which supports English and multilingual recognition via the \nvidia/nemotron-ocr-v2\ checkpoint. The existing engines—RapidOCR, EasyOCR, Tesseract, and OCRmac—have been refactored into their own files (\rapid\_ocr\_model.py\, \easyocr\_model.py\, \tesseract\_ocr\_cli\_model.py\, \ocr\_mac\_model.py\) with improved language canonicalization and configuration handling. An \OcrAutoModel\ orchestrates the selection of the best available engine based on the platform (e.g., NVIDIA Nemotron on Linux, OCRmac on macOS) and installed dependencies, while a \KserveV2OcrModel\ enables remote inference via the KServe v2 API.
docling/models/stages/ocr · high confidence
New performance benchmarking scripts for PDF and XLSX processing
Added three new scripts in the perfs directory to measure and visualize processing performance. iterate\_pdf\_pages.py benchmarks PDF parsing throughput and memory usage using the threaded Docling-parse backend, supporting configurable thread counts and outputting metrics for analysis. plot\_memory\_metrics.py visualizes the memory consumption data collected by the PDF benchmark. xlsx\_merged\_cells.py benchmarks XLSX conversion performance specifically for workbooks with many merged cell ranges, measuring generation time, conversion time, and peak memory usage.
perfs · high confidence
New pipeline architecture and expanded media support
The pipeline module has been restructured to introduce a new base pipeline hierarchy (BasePipeline, ConvertPipeline, PaginatedPipeline) and a dedicated base extraction pipeline, providing a more modular foundation for document processing. This change adds native support for new media types: an AsrPipeline for audio transcription and a VideoPipeline that extracts audio and frames to produce a combined transcript and image document. A new NativePdfPipeline is introduced for fast, model-free PDF conversion using native content, while the existing StandardPdfPipeline is refactored into a threaded, production-ready implementation with improved parallelism and back-pressure. Additionally, the VlmPipeline is updated to support a new pluggable runtime system via VlmConvertOptions, and the legacy StandardPdfPipeline is deprecated in favor of the new threaded version.
docling/pipeline · high confidence
New pre-commit and documentation build scripts
Added four new Python scripts to the \scripts/\ directory to support project hygiene and documentation generation. \check\_max\_lines.py\ enforces a 1000-line limit on source files via a pre-commit hook, with configurable ignore patterns. \check\_tach\_module\_coverage.py\ ensures all Python modules in the \docling\ package are explicitly declared in \tach.toml\, preventing stale or uncovered modules. \render\_cli\_reference.py\ generates the CLI reference documentation in Markdown by introspecting the Typer/Click command tree, replacing the previous \mkdocs-click\ dependency. \render\_notebooks.py\ pre-renders Jupyter notebooks and jupytext \.py\ scripts from \docs/examples\ into Markdown for the documentation site, handling syntax highlighting for percent-format Python files.
scripts · high confidence
New threaded two-stage Layout+VLM pipeline
An experimental \ThreadedLayoutVlmPipeline\ is introduced that processes documents in two stages: first, a layout model detects document elements and coordinates, which are then injected into the VLM prompt to enhance structured output. This pipeline uses a factory path for layout inference and includes a post-processing stage, allowing users to leverage layout-aware VLM processing for improved document understanding.
docling/experimental/pipeline · high confidence
New utility modules for hardware acceleration, layout postprocessing, and VLM parsing
The \docling/utils\ package has been expanded with several new modules that enhance document processing capabilities. \accelerator\_utils.py\ introduces a \decide\_device\ function that resolves the best available hardware accelerator (CUDA, MPS, or Intel XPU) based on system availability and user configuration. \layout\_postprocessor.py\ adds a \LayoutPostprocessor\ class that cleans up layout predictions by resolving overlapping clusters and mapping text cells. Additionally, new utility files (\chandra\_utils.py\, \deepseekocr\_utils.py\, \dots\_utils.py\, \mineru\_utils.py\) provide parsers for specific VLM output formats, while \font\_style.py\ enables the extraction of text weight and slant from PDF font names to improve heading detection.
docling/utils · high confidence
Shared Python agent skills added for Claude, Codex, and OpenCode
New shared skills for building Pydantic AI agents and dignified Python development have been added to the project. These skills are now available as symlinks in the .claude, .codex, and .opencode directories, pointing to the central .agents/skills location, allowing these AI coding assistants to access consistent Python agent development guidelines.
.claude, .codex, .opencode · high confidence
Unified inference engines for image classification and object detection
The system now provides a consistent, extensible inference engine architecture for both image classification and object detection tasks. Users can choose between local execution via ONNX Runtime or Hugging Face Transformers, or remote inference through the KServe v2 API (supporting both HTTP and gRPC transports). This unified approach standardizes configuration, input/output handling, and hardware acceleration options across these model families.
_docling/models/inference\_engines/object\detection · high confidence
Removals
Removal of legacy example scripts
The \examples/convert.py\ and \examples/minimal.py\ scripts have been removed from the repository. These files previously demonstrated how to initialize the \DocumentConverter\, download models, and export converted documents to JSON and Markdown formats; their removal indicates a shift in how users are expected to interact with the library, likely aligning with the broader v2 migration.
examples · high confidence
Architecture
Modularize LaTeX backend into a package structure
The LaTeX document backend has been reorganized from a single file into a structured package under \docling/backend/latex/\. This change introduces a modular architecture with dedicated modules for handling specific concerns: \backend.py\ contains the main \LatexDocumentBackend\ class, \constants.py\ defines macro and environment sets, \context.py\ manages parsing state, and \handlers/\ and \utils/\ subpackages handle macros, environments, math, text, and tables. This improves code maintainability and separation of concerns for LaTeX processing.
docling/backend/latex · high confidence
Modularize LaTeX backend into a utils package
The LaTeX backend utilities have been reorganized into a dedicated \docling/backend/latex/utils\ package. This change introduces new modules for handling encoding (\encoding.py\), table parsing (\table.py\), and text processing (\text.py\), which extract core logic for decoding LaTeX content, parsing complex table structures (including multicolumn and multirow cells), and processing text nodes and macros. This structural refactor improves code maintainability and separation of concerns within the LaTeX document processing pipeline.
docling/backend/latex/utils · high confidence
Modularize LaTeX backend into dedicated handler modules
The LaTeX backend processing logic has been reorganized into a structured package under \docling/backend/latex/handlers\, splitting responsibilities into \environments.py\ (handling document structures like tables, figures, and TikZ), \macros.py\ (managing custom macro extraction and inline processing), and \math.py\ (handling formula cleaning and display logic). This change improves code maintainability and separation of concerns within the LaTeX parsing pipeline without altering the external API.
docling/backend/latex/handlers · high confidence
Move reading-order and list marker models into docling
The reading-order and list-marker models have been relocated into the docling package structure (docling/models/stages/reading\_order). This change reorganizes the internal codebase to better align with the new module layout, ensuring that the reading order prediction logic and list item processing are now part of the core docling models rather than external or separate components.
_docling/models/stages/reading\order · high confidence
Refactored model architecture with plugin-based factories and new base classes
The model layer has been restructured to support a plugin-based architecture. Legacy model implementations (such as EasyOcrModel, LayoutModel, and TableStructureModel) have been removed and replaced with new abstract base classes (BaseLayoutModel, BaseOcrModel, BaseTableStructureModel, etc.) that define standardized interfaces. A new factory system (LayoutFactory, OcrFactory, etc.) has been introduced to dynamically register and load model implementations, allowing for extensibility via plugins. This change also introduces new base classes for VLMs and enrichment models, standardizing how models are instantiated and executed within the pipeline.
docling/models · high confidence
Restructure VLM model implementations into a dedicated submodule
The Vision-Language Model (VLM) engine implementations have been organized into a new \docling/models/vlm\_pipeline\_models\ package. This change introduces distinct modules for each backend: \api\_vlm\_model.py\ for remote API-based inference, \hf\_transformers\_model.py\ for local Hugging Face Transformers execution, \mlx\_model.py\ for Apple MLX support, and \vllm\_model.py\ for vLLM acceleration. By consolidating these specific engine implementations into a single location, the codebase improves modularity and makes it easier to maintain or extend individual inference backends without affecting the broader model registry.
_docling/models/vlm\_pipeline\models · high confidence
Behavioural changes
Added docx backend package initialization
The docling/backend/docx module now includes an \_\init\\_.py file, establishing the package structure for the MS Word backend. This file contains standard SPDX copyright and license headers (MIT) but does not expose any public API or logic in this specific diff.
docling/backend/docx · medium confidence
Chart extraction now uses a generic VLM engine with Granite Vision 4.1 support
The chart extraction stage has been refactored to use a unified Vision-Language Model (VLM) engine system, replacing the previous implementation. This change upgrades the default model to Granite Vision 4.1 and introduces configurable engine options, allowing users to select from various backends (such as Transformers, MLX, or API engines) via a preset mechanism. The new implementation supports multiple output formats (CSV, summary, and Python code) and includes natural-language prompt substitution for deployments that lack fine-tuned special tokens, providing greater flexibility in how chart data is extracted and processed.
_docling/models/stages/chart\extraction · high confidence
Docling Actor now uses docling-serve API with smaller Docker image
The Docling Actor on Apify has been updated to version 1.1.0, switching from the full Docling CLI to the docling-serve API. This change reduces the Docker image size from approximately 6GB to 4GB by using the official quay.io/ds4sd/docling-serve-cpu base image and implementing a multi-stage build. The actor now communicates with the docling-serve service via a standard API endpoint (/v1alpha/convert/source) instead of a custom one, requiring updated JSON payload structures. The actor script handles API communication, health checks, and content extraction, while the Dockerfile sets up the necessary environment, including tmpfs volumes for temporary files and specific permissions for OCR models.
.actor · high confidence
Docling v2: complete API overhaul with new pipeline and extraction architecture
Docling has been upgraded to version 2, introducing a major architectural shift in how documents are processed. The previous \DocumentConverter\ class has been replaced by a modular system featuring distinct pipeline classes (such as \StandardPdfPipeline\, \NativePdfPipeline\, \VideoPipeline\, and \AsrPipeline\) and a new \DocumentExtractor\ class for schema-based extraction. This change brings a new configuration model using \FormatOption\ and \BackendOptions\, supports a significantly expanded range of input formats (including CSV, AFP, iWork, JATS, and EPUB), and introduces features like threaded PDF parsing, page-range limits, and structured error handling.
docling · high confidence
Improved PDF text extraction with ligature normalization and hyperlink propagation
The page assembly stage now normalizes Unicode ligatures (such as ff, fi, fl, and Dutch IJ) into their standard ASCII equivalents during text extraction, ensuring that typographic characters are correctly converted to readable text. Additionally, the model now propagates hyperlink annotations to the resulting DoclingDocument text items, allowing users to access clickable links embedded in the original PDF.
_docling/models/stages/page\assemble · high confidence
Introduce pluggable VLM runtime for code and formula extraction
The code and formula extraction stage now uses a new pluggable VLM runtime system instead of a hardcoded Transformers implementation. Users can configure the extraction stage via \CodeFormulaVlmOptions\ and select from different inference engines (e.g., Transformers, MLX, API) through a unified factory interface, allowing for more flexible hardware acceleration and remote service integration while maintaining the same enrichment behavior for code blocks and formulas.
_docling/models/stages/code\formula · high confidence
Layout detection refactored to use object-detection engines with deprecated legacy shim
The layout detection stage now routes through a new \LayoutObjectDetectionModel\ that leverages generic object-detection inference engines (supporting HF Transformers and ONNX) instead of the previous direct model integration. A deprecated \LayoutModel\ shim is retained to maintain backward compatibility with \LayoutOptions\, automatically mapping legacy configurations to the new \LayoutObjectDetectionOptions\ and warning users to migrate. Additionally, a dedicated \LayoutPostprocessingModel\ stage now handles finalizing raw clusters, computing layout scores, and managing cell assignment, separating post-processing logic from the inference stage.
docling/models/stages/layout · high confidence
Modularized LaTeX backend with optional Tectonic TikZ rendering
The LaTeX backend has been restructured into a modular package, introducing abstract base classes for render engines and library handlers to support extensibility. This change adds optional support for rendering TikZ diagrams using the Tectonic engine, which automatically installs or detects the Tectonic binary to process complex LaTeX graphics.
docling/backend/latex/engines · high confidence
New experimental pipeline options for threaded layout and VLM processing
The \docling/experimental/datamodel\ module now exposes configuration classes for advanced layout processing workflows. \TableCropsLayoutOptions\ provides internal settings for the experimental TableCrops layout model, while \ThreadedLayoutVlmPipelineOptions\ defines the parameters for a threaded pipeline that combines layout detection with Vision-Language Models (VLM). This threaded pipeline configuration enforces the DOCTAGS response format, allows tuning of batch sizes and timeouts for both layout and VLM stages, and defaults to the GraniteDocling 2-stage transformers model.
docling/experimental/datamodel · high confidence
New page preprocessing stage with optimized cell extraction and text quality scoring
A new page preprocessing stage has been introduced to handle page image population and text cell extraction. This stage now skips native segmented-page (text-cell) decoding when full-page OCR is enabled, preventing unnecessary memory consumption on vector-dense pages. It also includes a text quality rating mechanism that penalizes garbage patterns (such as glyph placeholders or fragmented words) to help assess parse confidence, and optionally visualizes extracted cells (text, shapes, bitmaps) for debugging purposes.
_docling/models/stages/page\preprocessing · high confidence
Picture classifier now enriches documents with classification metadata via a unified inference engine
The picture classifier stage has been updated to use a unified image-classification inference engine, allowing it to process picture items and attach classification predictions (class names and confidence scores) to documents. This change introduces a new \DocumentPictureClassifier\ implementation that leverages the \BaseImageClassificationEngine\ for model execution, while maintaining backward compatibility by optionally writing results to the deprecated \annotations\ field when the \\_keep\_deprecated\_annotations\ flag is enabled.
_docling/models/stages/picture\classifier · high confidence
Planning documents for DoclingDocument compatibility, modular packaging, layout migration, and MinerU2-Pro VLM
This change introduces several planning documents for the .plans directory. The docling-compatibility.md plan details a strategy to make the DoclingDocument client a tolerant reader, addressing hard crashes from new enum values and optional fields by switching extra='forbid' to extra='ignore' and using BeforeValidator for enum coercion. The docling-slim.md plan outlines the modular docling-slim package, splitting dependencies into fine-grained extras for PDF backends, models, OCR engines, and input formats. The layout-migration.md plan records the implementation of running all layout inference through the object-detection factory path, replacing the legacy LayoutModel with LayoutObjectDetectionModel and deprecating the old options. Finally, the mineru2-pro-vlm.md plan describes the integration of the MinerU2-Pro VLM model, including a two-step stage execution for layout and recognition, and support for Transformers, MLX, and API engines.
.plans · high confidence
Restores document heading hierarchy in PDF conversions
The PDF conversion pipeline no longer flattens all headings to level 1. A new HeadingHierarchyModel stage now infers the correct section depth by prioritizing PDF bookmarks (the document's own outline), falling back to legal numbering patterns (such as PART I, 1.1, (a)), and finally using visual style cues like font size, weight, and case when structural signals are absent.
_docling/models/stages/heading\hierarchy · high confidence
Rule-based postprocessing for reading order and list markers
The \docling/models/postprocessing\ module now includes rule-based processors for determining document reading order and handling list item markers. The new \ReadingOrderPredictor\ sorts page elements (headers, footers, and content) based on spatial layout to establish a logical reading sequence, while the \ListItemMarkerProcessor\ identifies bullet and numbered list patterns, merging marker-only text items with their content to create proper \ListItem\ structures in the output document.
docling/models/postprocessing · high confidence
Unified extraction model with Granite Vision 4.1 support and prompt-style selection
The extraction logic in docling/models/extraction has been reorganized into a new submodule structure, introducing a unified TransformersExtractionModel that supports multiple prompt styles (NuExtract and Granite Vision). This change adds native support for the Granite Vision 4.1-4b model, automatically disabling trust\_remote\_code and adjusting attention implementation for transformers versions \>= 5.13 to ensure compatibility. Users can now select between the NuExtract-specific input formatting and the standard Granite Vision chat format via the ExtractionPromptStyle option, with the system handling the necessary input construction and generation configuration differences automatically.
docling/models/extraction · high confidence
Fixes
Hardened LibreOffice conversion with timeout and security safeguards
The DOCX backend now uses a hardened configuration for LibreOffice conversions to improve stability and security. A 60-second timeout prevents hung processes from blocking the application, and the LibreOffice instance is launched in an isolated, throwaway profile with strict flags (headless, no restore, no logo) to avoid side effects. Additionally, a custom registry configuration disables macro execution and prevents external link updates, mitigating risks from malicious documents. The module also safely handles missing optional dependencies like pypdfium2 to prevent import errors on minimal installations.
docling/backend/docx/drawingml · high confidence
Test coverage
Added WebVTT test regression data; Added XLSX ground truth test fixtures; Added ground truth data for USPTO patent IPA20180000016; Added ground truth data for legacy PPT format testing; Added regression test fixtures for AFP, AsciiDoc, BoxNote, and EBCDIC backends; Added test ground truth for legacy Office formats and localized headings; Expanded test coverage for the LaTeX backend; Removal of PyPDFium2 backend unit tests; Updated CSV test ground truth data; Updated HTML parsing ground truth for regression tests; Updated JATS test ground truth for elife-56337; Updated LaTeX test ground truth for 'Attention Is All You Need' and 'Optimized Table Tokenization'; Updated PowerPoint test ground truth data; Updated test ground truth for EPUB and XBRL formats.
Dependencies
Docling restructures into modular packages with docling-slim and docling-client
The project has been reorganized into a modular architecture to reduce installation size and improve dependency management. The main \docling\ package is now a meta-package that depends on \docling-slim\ (the core SDK and CLI) and re-exports optional extras (like OCR, VLM, and format support) for backward compatibility. A new \docling-client\ package has been introduced to provide a dedicated client SDK for interacting with remote Docling Serve endpoints. The build system has also been migrated from Poetry to Hatchling, and the project now supports Python versions 3.10 through 3.14.
(dependencies) · high confidence
Housekeeping
Initial repository scaffolding and configuration
The repository is initialized with foundational configuration files, including a Makefile for build and test automation, a uv.lock dependency manifest, a tach.toml module-layering configuration, and an AGENTS.md guide for AI coding assistants. Documentation is structured via mkdocs.yml, and repository metadata is defined in CITATION.cff and CLAUDE.md. Version control hygiene is established through .gitattributes for line-ending normalization and .git-blame-ignore-revs to ignore bulk license-header commits.
(repo-wide) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Score
- CAI 48 → 63 (+15.4)
- Rubric changed (rubric-2026.08.15 → rubric-2026.09.15) — scores are not directly comparable.
Lenses
- Code Health 84 → 76 (-8.1)
- Architecture 96 → 99 (+3.2)
- Maturity 73 → 76 (+2.4)
- Readiness 27 → 47 (+19.5)
- Security 53 → 80 (+27.1)
Resolved (45)
- Coverage not measured — test suite did not build
- Dimension evaluation failed
- Duplicated block (11 lines × 2) (docling/backend/html_backend.py)
- Duplicated block (11 lines × 2) (docling/backend/html_backend.py)
- Duplicated block (12 lines × 2) (docling/backend/msexcel_backend.py)
- Duplicated block (12 lines × 3) (docling/models/stages/reading_order/readingorder_model.py)
- Duplicated block (13 lines × 2) (docling/backend/latex/utils/table.py)
- Duplicated block (13 lines × 2) (docling/backend/mets_gbs_backend.py)
- Duplicated block (13 lines × 2) (docling/pipeline/asr_transcriber.py)
- Duplicated block (14 lines × 2) (docling/pipeline/standard_pdf_pipeline.py)
- Duplicated block (14 lines × 2) (docling/service_client/client.py)
- Duplicated block (15 lines × 2) (docling/backend/mets_gbs_backend.py)
- Duplicated block (15 lines × 2) (docling/models/stages/ocr/tesseract_ocr_cli_model.py)
- Duplicated block (15 lines × 2) (tests/test_asr_whisper_s2t.py)
- Duplicated block (16 lines × 2) (docling/models/stages/table_structure/table_structure_model_v2.py)
- Duplicated block (16 lines × 2) (docling/pipeline/asr_transcriber.py)
- Duplicated block (17 lines × 2) (docling/backend/msword_backend.py)
- Duplicated block (18 lines × 2) (docling/models/stages/ocr/auto_ocr_model.py)
- Duplicated block (19 lines × 2) (tests/test_asr_mlx_whisper.py)
- Duplicated block (19 lines × 3) (docling/backend/xml/uspto_backend.py)
- …and 25 more
New (619)
- ApiVlmModel.call (cognitive 16) (docling/models/vlm_pipeline_models/api_vlm_model.py)
- AsciiDocBackend._parse (cognitive 67) (docling/backend/asciidoc_backend.py)
- AsciiDocBackend._parse (cyclomatic 45) (docling/backend/asciidoc_backend.py)
- AutoInlineVlmEngine._select_engine (cognitive 27) (docling/models/inference_engines/vlm/auto_inline_engine.py)
- BaseOcrModel._find_pdf_aware_layout_ocr_rects (cognitive 22) (docling/models/base_ocr_model.py)
- BaseOcrModel.post_process_cells (cognitive 16) (docling/models/base_ocr_model.py)
- BoxNoteDocumentBackend._add_table (cognitive 24) (docling/backend/boxnote_backend.py)
- Change coupling: api_vlm_model.py ↔ mlx_model.py (docling/models/vlm_pipeline_models/api_vlm_model.py)
- Change coupling: docling_parse_backend.py ↔ pypdfium2_backend.py (docling/backend/docling_parse_backend.py)
- Change coupling: layout_model.py ↔ standard_pdf_pipeline.py (docling/models/stages/layout/layout_model.py)
- ChartExtractionVlmEngineModel.call (cognitive 27) (docling/models/stages/chart_extraction/granite_vision.py)
- ChartExtractionVlmEngineModel.call (cyclomatic 17) (docling/models/stages/chart_extraction/granite_vision.py)
- ConversionAssets.load (cognitive 20) (docling/datamodel/document.py)
- ConversionAssets.load (cyclomatic 18) (docling/datamodel/document.py)
- CsvDocumentBackend.convert (cognitive 22) (docling/backend/csv_backend.py)
- Documentation: no project overview (README.md)
- Duplicated block (10 lines × 2) (docling/backend/html_backend.py)
- Duplicated block (10 lines × 2) (docling/backend/html_backend.py)
- Duplicated block (10 lines × 2) (docling/backend/iwork/archives.py)
- Duplicated block (10 lines × 2) (docling/backend/latex/handlers/macros.py)
- …and 599 more
Changes since last survey
- 202 commits — 85 feature/other, 117 fixes
By area
- docling/backend — 52 commits
- tests/data — 35 commits
- (root) — 26 commits
- docling/models — 25 commits
- docling/datamodel — 24 commits
- docling/cli — 6 commits
- docling/utils — 6 commits
- docs/concepts — 4 commits
- docling/service_client — 3 commits
- docs/usage — 3 commits
- .github/codecov.yml — 2 commits
- docling/.agents — 2 commits
- docling/pipeline — 2 commits
- docs/examples — 2 commits
- tests/test_verify_utils.py — 2 commits
- .github/max-lines-ignore — 1 commit
- .github/workflows — 1 commit
- docling/document_converter.py — 1 commit
- docs/integrations — 1 commit
- docs/reference — 1 commit
Notable commits
- fix: fix(afp): follow PTOCA control sequence chaining (#4297)
- fix: fix(asciidoc): block titles, picture parenting, skipped headings, and list dedent (#4171)
- fix: fix(asciidoc): gate image loading (#4156)
- fix: fix(asciidoc): preserve ordered lists and literal blocks (#4118)
- fix: fix(asciidoc): recognise a rowspan-only cell specifier (#4290)
- fix: fix(asciidoc): remove dead max_image_data_base64_bytes option (#4173)
- fix: fix(asciidoc): skip empty tables after incomplete rows (#4300)
- fix: fix(asciidoc): stop a dedented list from crashing the backend (#3826)
- fix: fix(asr): handle DocumentStream input in MLX Whisper (#3885)
- fix: fix(asr): require whisper-s2t-reborn>=1.7.1 for WhisperS2T correctness fixes (#3941)
- fix: fix(backend): decode text documents that are not UTF-8 (#4202)
- fix: fix(backend): defer pypdfium2 import in office and LaTeX backends (#4285)
- fix: fix(backend): defer the docling-parse import in the PDF backend (#4286)
- fix: fix(backend): translate line endings when decoding text from a stream (#4354)
- fix: fix(cache): fall back safely when serialize_as_any detects false circular reference (#4240)
- fix: fix(cli): defer heavy imports so CLI works on lightweight installs (#4100)
- fix: fix(cli): propagate service flags to AsrPipelineOptions (#4003)
- fix: fix(cli): write the --show-layout HTML export as UTF-8 (#4048)
- fix: fix(csv): detect the dialect when a quoted field spans several lines (#3985)
- fix: fix(csv): drop blank lines instead of turning them into empty rows (#4316)
- …and 182 more
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
docling-project/docling was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 26 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 2d5c590c34b6378fd8a47c65b534b280aa40c93c — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-09659c52afae.