opendataloader-project/opendataloader-pdf
63.4
Adequate · 24 September 2026
23.9k
lines of production code
Java
with Python
5
measurements over time
What this system is
OpenDataLoader PDF is a multi-language (Java, Python, Node.js) library and CLI for extracting structured content from PDF documents. It supports hybrid processing that combines fast local Java parsing with an external AI backend for complex layouts, enabling the generation of JSON, HTML, Markdown, and text outputs. The system includes features for auto-tagging PDFs for accessibility, sensitive data sanitization, and integration with RAG pipelines via an MCP server.
How it got here
2025 — Monorepo restructuring and multi-language SDKs
30 changes.
The project was rebranded to OpenDataLoader PDF and migrated to a monorepo structure supporting Java, Python, and Node.js SDKs. This period focused on establishing a unified build system, removing legacy CLI and Markdown conversion logic, and introducing a comprehensive API with hybrid AI processing capabilities.
2026 — AI hybrid processing and MCP integration
16 changes.
This period focused on integrating an AI-driven hybrid PDF backend using docling-fast, introducing semantic entity classes for enriched content like alt-text and formulas, and launching an MCP server for AI agent integration. Concurrently, the team expanded test coverage for CLI options, Markdown generation, and specific PDF parsing edge cases while refining error handling and auto-tagging accuracy.
Features
Add XY-Cut++ algorithm for improved multi-column reading order detection
Introduces the XY-Cut++ algorithm (XYCutPlusPlusSorter) to enhance reading order detection in PDFs with complex layouts. This new capability handles cross-layout elements like headers and footers, uses adaptive axis selection based on density ratios, and includes logic to filter narrow outlier elements that might bridge column gaps, resulting in more accurate text extraction for multi-column documents.
java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/processors/readingorder · high confidence
Introduce OpenDataLoader PDF CLI with robust error handling and quiet mode
The new CLIMain entry point provides a command-line interface for processing PDF files and folders, featuring a --quiet option to suppress logging while still reporting final results, and specific error handling for non-PDF inputs, missing files, and unwritable temporary directories. It ensures non-zero exit codes on failures, prevents JVM hangs by reusing the hybrid client, and displays a clear summary when a specified folder contains no processable PDFs.
java/opendataloader-pdf-cli/src/main · high confidence
Introduce hybrid PDF processing with docling-fast backend
This change adds the core infrastructure for hybrid PDF extraction, routing pages to an external AI backend (docling-fast) while keeping a Java fallback. It introduces HybridConfig (with defaults like 0ms timeout and 50-page chunks to prevent backend hangs), a HybridClientFactory that caches and reuses client instances to avoid JVM hangs, and the DoclingSchemaTransformer which converts Docling JSON into the OpenDataLoader IObject hierarchy (paragraphs, headings, tables, pictures). A new ElementMetadata sidecar class exposes per-element details such as the backend's object\_id, AI scores, and text source (stream/OCR). Triage decisions are now logged to JSON for benchmark evaluation, and the system supports chunking large documents to improve stability.
java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/hybrid · high confidence
Introduces StaticLayoutContainers for thread-local state management
A new \StaticLayoutContainers\ class has been added to centralize thread-local storage for PDF processing state, including content IDs, semantic headings, image metadata (directory, format, embedding status), and per-page replacement character ratios. This component also manages a cache for embedded image bytes, ensuring that image data is correctly shared and accessed across different processing stages and threads, which supports the new capabilities for relative image paths and Base64 embedding in output formats.
java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/containers · high confidence
Introduces hybrid PDF processing mode with AI backend integration and in-memory tagging
This release adds a new hybrid processing capability that routes complex PDF pages to an external AI backend (docling-fast) for advanced parsing, while keeping simple pages in the fast Java path. The core changes include the new \TaggingResult\ class for holding in-memory tagged PDFs and processing metadata, the \HybridClient\ interface and \DoclingFastServerClient\ implementation for communicating with the backend, and the \TriageProcessor\ which uses vector graphics and layout signals to decide which pages need AI assistance. Additionally, the \TextGenerator\ now supports configurable page separators and header/footer inclusion, and the system exposes raw responses and timings for downstream tooling.
opendataloader-pdf-core · high confidence
New Agent Skill for correct opendataloader-pdf usage
Added the \odl-pdf\ Agent Skill, which provides AI coding assistants with a durable, runtime-discovery procedure for using opendataloader-pdf correctly. Instead of relying on a static list of flags that may become outdated, the skill instructs agents to read the installed tool's \--help\ at runtime to discover options, build minimal commands, and verify that extractions actually succeeded (guarding against silent failures like empty outputs on clean exits). The skill also includes a maintenance kit (\odl-pdf-maintenance\) with a version-coupling lint to prevent baking specific option names or versions into the instructions, ensuring the skill remains valid across tool updates.
skills · high confidence
New CI verification harness for CLI options and output assertions
Added a new \verification/ci-verify.py\ script that provides a three-level verification framework (smoke, content assertion, and comparison) for the opendataloader-pdf CLI. This tool validates that CLI options are covered, ensures output files are generated correctly, checks for specific content strings, and prevents stack-trace leakage in user-facing error messages, thereby strengthening the reliability of PR and release workflows.
verification · high confidence
New CLI options for image resolution, space ratio, and hybrid chunking
The CLI now exposes three new configuration knobs: --image-resolution allows setting the rendering DPI for extracted images (default 144.0), --space-ratio controls the threshold for automatic space insertion in text (default 0.17), and --hybrid-chunk-size lets users split large documents into batches sent to the hybrid backend (default 50 pages). These options are defined in the newly created CLIOptions adapter class, which centralizes the mapping of Apache Commons CLI arguments to the core Config and HybridConfig objects.
java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/api/cli · high confidence
New HTML output generator with accessibility and formatting features
The PDF conversion tool now includes a dedicated HTML generator that produces semantic HTML5 output from PDF documents. This new capability supports embedding images as Base64 data URIs or relative file paths, configurable page separators, and optional inclusion of page headers and footers. The generator handles complex content structures including tables, lists, headings, paragraphs, and mathematical formulas, while also supporting strikethrough and underline text decorations. Security is addressed through HTML attribute and text escaping to prevent XSS vulnerabilities, and accessibility is improved via sanitized alt text for image descriptions. The implementation consists of the HtmlGenerator class for core logic, HtmlGeneratorFactory for instantiation, and HtmlSyntax constants for HTML structure definitions.
java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/html · high confidence
New Java API for PDF extraction, tagging, and output generation
The library now exposes a public Java API in the \org.opendataloader.pdf.api\ package, providing programmatic access to PDF processing capabilities. Users can perform a simple extraction-and-output workflow via \OpenDataLoaderPDF.processFile\, or use a two-phase approach with \DocumentProcessor.extractContents\ followed by \OutputWriter.writeOutputs\ to generate JSON, Markdown, HTML, text, and images without re-parsing. The new \AutoTagger\ class allows generating in-memory tagged PDFs from extraction results, supporting PDF/UA compliance and structural tree generation. Configuration is managed through the \Config\ class, which exposes options for hybrid backends (e.g., \docling-fast\), reading order algorithms (XY-Cut++), table detection methods, image handling (Base64 embedding vs. external files), and sensitive data filtering. A dedicated \FilterConfig\ class allows fine-grained control over content sanitization, including filtering hidden text, tiny text, out-of-page content, hidden OCGs, background vector graphics, and sensitive data patterns (emails, credit cards, IPs, etc.).
java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/api · high confidence
New MCP server for AI-driven PDF conversion
Introduces a new Model Context Protocol (MCP) server for OpenDataLoader PDF, enabling AI agents to convert PDF files into Markdown, JSON, HTML, and text formats. The server exposes a \convert\_pdf\ tool that supports options for password-protected files, page selection, content sanitization, and hybrid processing modes. It is designed to integrate with MCP-compatible clients such as Claude Desktop, Cursor, and OpenAI Codex via the \uvx\ command.
python/opendataloader-pdf-mcp · high confidence
New Markdown output generator with enhanced formatting and configuration options
The Markdown generator has been rewritten to support richer content extraction and better formatting control. Users can now include headers and footers, use custom page separators, and embed images as Base64 data or external files. The output quality is improved with proper handling of table headers, list labels, and image link destinations (wrapped in angle brackets to handle special characters). Additionally, the generator now supports semantic formulas and picture descriptions, and allows output to stdout for piping.
java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/markdown · high confidence
New batch processing and RAG integration examples
Added new Python examples in the \examples/python/batch\ and \examples/python/rag\ directories. The batch processing example demonstrates how to convert multiple PDFs in a single JVM invocation to improve performance, supporting both file lists and recursive directory scanning. The RAG examples provide working implementations for Retrieval-Augmented Generation pipelines, including a basic chunking script with three strategies (by element, by section, and minimum size) and a LangChain integration example using the \OpenDataLoaderPDFLoader\.
examples · high confidence
New build, test, and release automation scripts
The repository now includes a comprehensive suite of shell and Node.js scripts to streamline development and release workflows. Developers can use \build-all.sh\ to build and test the Java, Python, and Node.js packages in a single step, or run individual scripts like \build-java.sh\, \build-python.sh\, and \build-node.sh\ for specific languages. Local development is supported via \test-java.sh\, \test-python.sh\, and \test-node.sh\. The \bench.sh\ script allows running benchmarks against the locally built JAR, with an optional \--skip-build\ flag. For releases, \preflight.sh\ verifies deploy credentials (npm, Maven, GPG, GitHub, PyPI) before the build starts, and \set-dev-version.sh\ updates the version across all three ecosystems. Additionally, \generate-options.mjs\ and \generate-schema.mjs\ auto-generate CLI option definitions and JSON Schema documentation from central \options.json\ and \schema.json\ files, ensuring consistency across Node.js, Python, and documentation.
scripts · high confidence
New semantic entity classes for enriched PDF content
The PDF processing core now includes new entity classes to support enriched semantic content: EnrichedImageChunk carries AI-generated image descriptions (alt text) and tracks their source (original vs. AI-generated) for accessibility and JSON output; SemanticFootnote maps to the PDF FENote structure element; SemanticFormula stores mathematical formulas in LaTeX format; and SemanticPicture holds picture metadata with optional AI-generated descriptions. These classes enable the hybrid backend to bridge AI-generated content (like SmolVLM picture descriptions and LaTeX formula extraction) into the PDF struct tree and downstream output formats (JSON, Markdown, HTML) with proper sanitization and source attribution.
java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/entities · high confidence
New structured PDF processing pipeline with auto-tagging and hybrid AI integration
The \org.opendataloader.pdf.processors\ package has been restructured to introduce a multi-phase extraction and tagging pipeline. A new \AutoTaggingProcessor\ now generates PDF structure trees (marked content, parent trees, and semantic elements like headings, lists, and tables) to support PDF/UA compliance and structured outputs. Table detection is handled by a new \AbstractTableProcessor\ base class and a \ClusterTableProcessor\ implementation that uses spatial clustering to identify table borders. Content is pre-processed by a \ContentFilterProcessor\ that removes backgrounds, hidden text, and out-of-page artifacts. The \DocumentProcessor\ now coordinates a two-phase flow (extraction then output generation) and supports parallel per-page processing. Additionally, a \HybridDocumentProcessor\ enables AI-assisted processing by routing pages to an external backend (e.g., Docling) while handling Java-side fallback, OCR enrichment, and result merging.
java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/processors · high confidence
Node.js wrapper for OpenDataLoader PDF CLI released
The \node/opendataloader-pdf\ package is now available, providing a Node.js interface to the OpenDataLoader PDF conversion engine. It exposes a \convert()\ library API for programmatic use and a bundled CLI for command-line execution. The wrapper supports a comprehensive set of options including output format selection (JSON, HTML, Markdown, etc.), image handling (embedded Base64 or external files), hybrid backend integration, and structural extraction features like strikethrough detection and reading-order algorithms. It also includes sensitive data sanitization, per-page parallelism, and selective page extraction.
node/opendataloader-pdf · high confidence
Removals
Removal of PDF-to-Markdown conversion and CLI entry point
The command-line interface entry point (CLIMain) and its option definitions (CLIOptions) have been removed, eliminating the ability to run the tool from the command line. Additionally, the core PDF-to-Markdown conversion logic has been deleted, including the MarkdownGenerator and MarkdownGeneratorFactory classes, as well as supporting utilities for handling images, line art, and text serialization. This change removes the product's capability to generate Markdown output from PDF files.
src/main · high confidence
Behavioural changes
Automated cross-language versioning and JAR artifact management
Build scripts now synchronize version numbers across the Java and Python components of the monorepo and manage the required runtime dependency. The new set\_version.py script reads a single VERSION file and updates both the Maven pom.xml and the Python pyproject.toml to ensure consistent release versions. Additionally, fetch\_shaded\_jar.py automates the retrieval of the latest shaded JAR from the Java build output, selecting the highest semantic version and copying it to the Python package source tree as runtime.jar for use by the Python library.
build-scripts · high confidence
Improved PDF structure tree accuracy for charts and nested XObjects
The auto-tagging engine now correctly associates path-drawing operators (used for charts, graphs, and shapes) with their marked-content sequences, ensuring they appear in the PDF structure tree rather than being orphaned. It also preserves nested XObjects (Do operators) within their parent streams and guards against null references during processing, resulting in more reliable semantic tagging for complex documents.
java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/autotagging · high confidence
JSON output schema expanded with TOC, hybrid metadata, and AI enrichment fields
The JSON writer now supports a richer output schema: it emits a dedicated hybrid root block for AI-derived data (such as picture descriptions and formula enrichment), includes per-node metadata like AI scores and PDF/UA tags, and adds new element types including Table of Contents (TOC) items, captions, and formulas. The package namespace has been updated from com.hancom to org.opendataloader, and the license header has been migrated from MPL-2.0 to Apache-2.0. Additionally, LineArtChunk entries are now suppressed from the JSON output, and the --include-header-footer flag is correctly applied to control whether header/footer elements are written.
java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/json · high confidence
JSON serialization overhaul: new serializers, metadata, and accessibility enhancements
The JSON output format has been significantly updated with new serializers for formulas, images, pictures, text lines, and tables of contents, alongside a refactored utility layer. Users will now see element-level metadata (such as AI confidence scores, source labels, and text source origins) included in the JSON for semantic elements. Image handling has been enhanced to support Base64 embedding or relative file paths, with a unified alt-text schema that distinguishes between original and AI-generated descriptions. Additionally, table header semantics are now explicitly preserved, and LineArtChunks are suppressed from JSON output to reduce noise.
java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/json/serializers · high confidence
License migration, package rename, and TOC support in PDF annotation engine
The PDF annotation core library has migrated its license from MPL-2.0 to Apache-2.0 and moved from the \com.hancom.opendataloader\ to the \org.opendataloader\ package namespace. Functionally, the PDF writer now supports generating Table of Contents (TOC) annotations, handling multi-page bounding boxes via \MultiBoundingBox\, and correctly identifying table header cells. The implementation also improves robustness by using structured logging instead of console output and fixing bounding box calculations relative to page coordinates.
java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/pdf · high confidence
PDF processing utilities refactored and relocated to org.opendataloader.pdf.utils
The PDF processing utility classes have been moved from the com.hancom.opendataloader package to org.opendataloader.pdf.utils and updated to the Apache-2.0 license. This refactoring introduces new capabilities including Base64 image embedding for self-contained outputs, enhanced bulleted paragraph detection with extensive Unicode support, and configurable text node statistics for heading detection. The changes also include a content sanitizer for text replacement, improved image handling with memory-safe BufferedImage management, and various bug fixes for file name parsing and null-safety.
java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/utils · high confidence
Project rebranded to OpenDataLoader PDF with Apache 2.0 license and monorepo structure
The project has been renamed from 'open-pdf-dataloader' to 'OpenDataLoader PDF' and migrated to a monorepo structure supporting Python, Node.js, and Java SDKs. The license has changed from MPL-2.0 to Apache 2.0, and the repository now includes a comprehensive CLI option schema (\options.json\) and JSON output schema (\schema.json\). The README has been updated to highlight AI-ready data extraction and PDF accessibility automation features, including auto-tagging to Tagged PDF.
(repo-wide) · high confidence
Python package build system migrated to uv and hatchling with hybrid server support
The Python package in python/opendataloader-pdf has been rebuilt using the uv package manager and hatchling build system, replacing the previous setuptools configuration. This change introduces a custom hatch build hook (hatch\_build.py) that automatically copies the Java CLI JAR, license, and notice files into the Python package distribution, ensuring the package is self-contained for pip installs. The package now exposes a new Python API with a modern \convert()\ function and a deprecated \run()\ function for backward compatibility, alongside a new \opendataloader\_pdf.hybrid\_server\ module that provides a FastAPI-based server for hybrid PDF processing using the docling backend. The minimum supported Python version is raised to 3.10, and the dependency lockfile is now managed via uv.lock.
python/opendataloader-pdf · high confidence
Fixes
New specific exception types for PDF processing errors
The PDF core library now introduces three distinct exception classes to provide clearer error reporting: TempDirectoryNotWritableException is thrown when the JVM's temporary directory cannot be written to, causing processing to fail fast; InvalidPdfFileException distinguishes between missing PDF headers (e.g., non-PDF files with .pdf extensions) and corrupted/truncated content; and EncryptedTaggedPdfNotSupportedException prevents tagged-PDF generation on encrypted documents with a friendly message. These changes allow users to handle specific failure modes—environment issues, invalid inputs, and unsupported document states—more precisely than generic IOExceptions.
java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/exceptions · high confidence
Test coverage
Added CLI exit-code and logging regression tests; Added Python test suite for CLI options, hybrid server, and runner behavior; Added integration tests for PDF core features; Added regression tests for veraPDF ToUnicode byte overflow fix; Added synthetic CID font PDF test fixture and generator; Added test coverage for Node.js PDF conversion library and CLI; Added test coverage for PDF processing processors; Added tests for AutoTagger, Config, and FilterConfig APIs; Added tests for CLI content safety and image output options; Added tests for HTML text escaping and style generation; Added tests for JSON serialization of element metadata, images, and line art; Added tests for Markdown heading normalization, link destination escaping, and merged table cells; Added tests for StaticLayoutContainers configuration and error logging; Added tests for picture description sanitization and alt-text generation; Added tests for the PDF conversion MCP tool; Added unit tests for PDF utility classes; Added unit tests for XY-Cut++ reading order sorter; Added unit tests for hybrid PDF processing components.
Dependencies
Initial dependency manifests for OpenDataLoader PDF
This change introduces the initial dependency configuration files for the OpenDataLoader PDF project, establishing the build and runtime requirements for the Java, Python, and Node.js components. The Java build is defined by a new parent POM and module POMs (opendataloader-pdf-core, opendataloader-pdf-cli) that declare dependencies on veraPDF, PDFBox, Jackson, and OkHttp, alongside build plugins for shading and publishing. Python dependencies are specified in pyproject.toml files for the core package (including optional hybrid dependencies like docling and fastapi) and the MCP server, while Node.js dependencies are managed via package.json and pnpm-lock.yaml, requiring Node \>= 22.13 and using packages like commander, vitest, and eslint. Example requirements.txt files are also added for Python batch and RAG examples.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Score
- CAI 61 → 63 (+2.4)
- Rubric changed (rubric-2026.08.19 → rubric-2026.09.15) — scores are not directly comparable.
Lenses
- Code Health 84 → 81 (-2.8)
- Architecture 98 → 96 (-2.6)
- Maturity 66 → 60 (-6.2)
- Readiness 57 → 56 (-0.6)
- Security 55 → 74 (+19.5)
Resolved (108)
- Change coupling: HancomAIClient.java ↔ HancomAISchemaTransformer.java (java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/hybrid/HancomAIClient.java)
- Coverage not included — suite not readable by the collector
- Dependency hygiene not measured — dependency manifest found but not parsed for hygiene
- Duplicated block (10 lines × 2) (java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/entities/EnrichedImageChunk.java)
- Duplicated block (10 lines × 2) (java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/hybrid/DoclingSchemaTransformer.java)
- Duplicated block (10 lines × 2) (java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/hybrid/HancomAIClient.java)
- Duplicated block (10 lines × 2) (java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/hybrid/HancomAISchemaTransformer.java)
- Duplicated block (10 lines × 2) (java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/json/serializers/ImageSerializer.java)
- Duplicated block (10 lines × 2) (java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/markdown/MarkdownGenerator.java)
- Duplicated block (10 lines × 2) (java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/processors/TaggedDocumentProcessor.java)
- Duplicated block (11 lines × 2) (java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/html/HtmlGenerator.java)
- Duplicated block (11 lines × 2) (java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/processors/ParagraphProcessor.java)
- Duplicated block (11 lines × 2) (java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/processors/TableStructureNormalizer.java)
- Duplicated block (11 lines × 2) (java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/processors/readingorder/XYCutPlusPlusSorter.java)
- Duplicated block (12 lines × 2) (java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/hybrid/DoclingSchemaTransformer.java)
- Duplicated block (12 lines × 2) (java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/hybrid/DoclingSchemaTransformer.java)
- Duplicated block (5 lines × 2) (java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/html/HtmlGenerator.java)
- Duplicated block (5 lines × 2) (java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/hybrid/HancomAIClient.java)
- Duplicated block (5 lines × 2) (java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/hybrid/HancomAISchemaTransformer.java)
- Duplicated block (5 lines × 2) (java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/hybrid/HancomAISchemaTransformer.java)
- …and 88 more
New (139)
- ClassTooLong: AutoTaggingProcessor (java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/processors/AutoTaggingProcessor.java)
- ClassTooLong: CLIOptions (java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/api/cli/CLIOptions.java)
- ClassTooLong: HybridDocumentProcessor (java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/processors/HybridDocumentProcessor.java)
- ClassTooLong: ListProcessor (java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/processors/ListProcessor.java)
- ClassTooLong: TriageProcessor (java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/hybrid/TriageProcessor.java)
- Coverage not measured — JavaScript/TypeScript suite
- Dependency hygiene PARTLY measured — Maven/Gradle declarations read, no dependency graph resolved
- DocumentProcessor.setIDsInComplexObject (cognitive 26) (java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/processors/DocumentProcessor.java)
- Documentation: no architecture or design documentation (README.md)
- Documentation: no installation or build instructions (README.md)
- Documentation: no usage examples (README.md)
- Duplicated block (10 lines × 2) (java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/html/HtmlGenerator.java)
- Duplicated block (10 lines × 2) (java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/json/serializers/ListItemSerializer.java)
- Duplicated block (10 lines × 2) (java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/processors/ParagraphProcessor.java)
- Duplicated block (11 lines × 2) (java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/html/HtmlGenerator.java)
- Duplicated block (11 lines × 2) (java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/markdown/MarkdownGenerator.java)
- Duplicated block (11 lines × 2) (java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/processors/ListProcessor.java)
- Duplicated block (11 lines × 2) (java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/processors/TaggedDocumentProcessor.java)
- Duplicated block (12 lines × 2) (java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/html/HtmlGenerator.java)
- Duplicated block (12 lines × 2) (java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/json/serializers/TextChunkSerializer.java)
- …and 119 more
Changes since last survey
- 65 commits — 46 feature/other, 19 fixes
By area
- java/opendataloader-pdf-core — 31 commits
- python/opendataloader-pdf — 10 commits
- .github/workflows — 5 commits
- java/opendataloader-pdf-cli — 3 commits
- docs/residual-review-findings — 2 commits
- java/pom.xml — 2 commits
- node/opendataloader-pdf — 2 commits
- python/opendataloader-pdf-mcp — 2 commits
- scripts/dla-corpus-scan — 2 commits
- (root) — 1 commit
- .github/dependabot.yml — 1 commit
- docs/hybrid — 1 commit
- examples/python — 1 commit
- samples/json — 1 commit
- scripts/build-all.sh — 1 commit
Notable commits
- fix: Fix NPE in needToAddAnnotationToStructTree() (#670)
- fix: Fix issue with duplicated content ids (#691)
- fix: chore(deps): unify Node 24 / pnpm 11 and fix postcss override
- fix: ci: push the version bump with the PAT, fix the retry loop
- fix: fix(autotag): drop the source /XRefStm from the tagged trailer
- fix: fix(cli): reject blank hybrid chunk size and preserve exception cause
- fix: fix(cli): restore PDFBox to compile scope so it ships in the shaded jar
- fix: fix(cli): trim hybrid chunk size once and strengthen validation tests
- fix: fix(cli): use valid PDF fixture for hybrid chunk size validation test
- fix: fix(core): fail fast when the temporary directory is not writable
- fix: fix(core): fail when the probe file cannot be deleted
- fix: fix(core): wrap path construction in its marked-content sequence
- fix: fix(header-footer): skip text nodes with no first non-space line
- fix: fix(hybrid): give list lines the same chunk matching paragraphs get
- fix: fix(hybrid): keep the heading depth docling infers
- fix: fix(release): let pnpm version run on the version-injected tree
- fix: fix(review): apply review findings
- fix: fix(tagging): write /ParentTree keys in ascending order
- fix: fix(test): sync lorem.json sample with pdfua_tag output
- change: Add --hybrid-chunk-size CLI option and bindings
- …and 45 more
Architecture
- Containers 0 added · 0 removed · contexts 1 added · 2 removed · edges 0 added · 0 removed
Added bounded contexts (1)
- opendataloader-pdf-core
Removed bounded contexts (2)
- opendataloader-pdf-parent
- python
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
opendataloader-project/opendataloader-pdf was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 24 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 323578baf821a851aabc99f6cfc49acc1d175de4 — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-ae95d6cad036.