xberg-io/xberg
48.1
Weak · 29 September 2026
772.9k
lines of production code
Rust
primary language
2
measurements over time
What this system is
Xberg is a high-performance, multi-language document intelligence framework that extracts text, tables, images, and structured metadata from over 100 file formats. It provides a comprehensive suite of AI-powered capabilities, including OCR, layout detection, named entity recognition, and semantic chunking, with support for both cloud LLMs and self-hosted on-device models. The system is designed for extensibility and broad deployment, offering native bindings for Rust, Python, Node.js, Java, PHP, and Go, alongside a Kubernetes-ready Helm chart and a built-in HTTP API server.
Features
Add ColBERT late-interaction multi-vector embedding support
Users can now perform retrieval using ColBERT-style late-interaction models, which generate per-token multi-vector embeddings instead of single pooled vectors. This crate introduces the \LateInteractionEngine\ to handle ONNX inference with specific tokenization tricks (marker insertion and query augmentation) and exposes a \MultiVectorEmbedding\ type for MaxSim scoring. The feature includes self-hosted presets for \colbert-small-v1\ and \gte-moderncolbert-v1\ from the \xberg-io/late-interaction-models\ repository, with model files pinned via SHA-256 checksums to ensure integrity.
_crates/xberg/src/late\interaction · high confidence
Add GLiNER2 schema-prompt inference engine
The \crates/xberg-gliner/src/v2\ module introduces a new inference engine for the GLiNER2 model architecture, which uses a schema-prompt approach rather than the original span-mode. This includes a dedicated tokenizer (\V2Tokenizer\) for pre-tokenized word-level encoding, a splitter (\V2Splitter\) that lowercases input to match training data, and a preprocessing pipeline (\encode\_v2\) that constructs the required schema prompt and position tensors. The engine (\Gliner2\) runs the ONNX model and decodes the resulting \span\_scores\ into entity spans, supporting configuration options for thresholds, maximum width, and NER overlap policies.
crates/xberg-gliner/src/v2 · high confidence
Add PaddleOCR-VL vision-language model support
Introduces a new PaddleOCR-VL backend for the Candle OCR engine, enabling multi-task document understanding including text, table, formula, and chart recognition. This change adds the complete inference pipeline for this model, including JSON-driven configuration structures for the SigLIP vision encoder and ERNIE text decoder, a processor for image preprocessing and tensor manipulation, and the core model implementation with spatial merging and rotary embeddings.
_crates/xberg-candle-ocr/src/models/deepseek\_ocr, crates/xberg-candle-ocr/src/models/glm\_ocr, crates/xberg-candle-ocr/src/models/paddleocr\vl · high confidence
Add SPLADE sparse embedding support with self-hosted presets
Users can now generate sparse (SPLADE) embeddings for hybrid retrieval, producing high-dimensional vocabulary vectors that capture exact lexical term importance. This change introduces the \SparseEmbeddingEngine\ which runs a \BertForMaskedLM\ ONNX model to compute sparse vectors via SPLADE pooling (log, relu, max-pool, L2-normalize, threshold). Two self-hosted presets are included: \splade\ (SPLADE++ EN v1) and \opensearch-v3-distill\ (OpenSearch neural-sparse v3 distill), both pinned via SHA-256 manifests for security and integrity. The implementation supports concurrent inference via \Arc\ and integrates with the existing ONNX acceleration cache.
_crates/xberg/src/sparse\embeddings · high confidence
Add WordPerfect text extraction crate
Introduces the \xberg-libwpd\ crate, enabling extraction of text and Markdown from WordPerfect documents (\.wpd\, \.wp5\, \.wp6\) on Linux, macOS, and Windows. The crate compiles \libwpd\ and \librevenge\ from source at build time, providing safe Rust functions to extract plain text, Markdown, and metadata like footnotes and tables, while stubbing out support for unsupported platforms like WASM.
crates/xberg-libwpd · high confidence
Add optional diff comparison for extraction results
Users can now compare two \ExtractedDocument\ values to identify structural differences in content, tables, metadata, and embedded children. This capability is available behind the \diff\ Cargo feature (enabled via \xberg = { features = \["diff"\] }\) and exposes a \compare\ function along with types like \ExtractionDiff\, \DiffOptions\, and \DiffHunk\. The diff engine supports optional truncation of content, conditional inclusion of metadata and embedded changes, and provides structured outputs for added/removed/changed tables and cell-level modifications.
crates/xberg/src/diff · high confidence
Add static assets for the docs site
Added CNAME, robots.txt, and demo.html files to the docs-site/public directory to support the new documentation site deployment and live demo experience.
docs-site/public · high confidence
Add text chunking library vendored from text-splitter v0.30.1
The text splitting logic in \crates/xberg/src/chunking/text\_splitter\ has been replaced with a vendored copy of the \text-splitter\ crate (v0.30.1). This introduces \TextSplitter\ for plain text and \MarkdownSplitter\ for Markdown documents, both of which split content into chunks based on configurable semantic levels (such as headings, paragraphs, or sentences) while respecting a target chunk size. The implementation uses \pulldown-cmark\ for Markdown parsing and \icu\_segmenter\ for Unicode-aware fallback segmentation (words, sentences, graphemes). Token-based sizing is supported via the optional \chunking-tokenizers\ feature, which integrates with \tokenizers\ v0.23 to count tokens using Hugging Face models.
_crates/xberg/src/chunking/text\splitter · high confidence
Added PDF corpus oracle example for regression detection
A new \pdf\_oracle\ example has been added to the \xberg-native-pdf\ crate to serve as a deterministic baseline for comparing PDF processing behavior. This tool scans the \test\_documents/\ corpus and outputs a TSV snapshot of document page counts, extracted text hashes, and error statuses for every PDF. It supports \--timing\ to measure wall-clock performance separately and \--headers\ to analyze raw byte-level header reading. Users can run this tool before and after code changes to verify that no regressions in text extraction or error handling have occurred.
crates/xberg-native-pdf/examples · high confidence
Added YAKE-based keyword extraction module
The \crates/xberg/src/keywords/yake\ module introduces a new keyword extraction capability based on the YAKE algorithm. This addition includes text preprocessing (sentence and word splitting), term tagging (identifying digits, punctuation, acronyms, and capitalization), and statistical scoring (using online statistics and Levenshtein distance for deduplication) to identify and rank important key phrases from input text.
crates/xberg/src/keywords/yake · high confidence
Benchmark harness introduces structured adapter system with OCR language and batch execution support
The benchmark harness now uses a modular adapter system to standardize how extraction frameworks are benchmarked. This change introduces explicit support for OCR language policies, allowing frameworks to specify how they handle language requests (e.g., per-document vs. batch-global), and adds native batch execution capabilities to measure performance on multiple files simultaneously. The harness now includes adapters for external tools like Docling, Unstructured, and Mineru, as well as a native Rust adapter for Xberg, ensuring consistent timeout handling, format validation, and resource monitoring across all tested frameworks.
tools/benchmark-harness · high confidence
CLI commands reorganized into submodules with new diagnostic and format-reporting capabilities
The CLI command implementations have been split into dedicated submodules (cache, chunk, config, doctor, embed, formats, ner, server, tree\_sitter) to improve maintainability. This change introduces a new \doctor\ command that probes configured backends and reports execution readiness, and a \formats\ command that accurately lists only the extraction formats supported by the currently compiled binary features, preventing the previous mismatch where unsupported formats were advertised. Additionally, the \server\ module now supports configurable allowed hosts for the MCP HTTP transport, and the \cache\ module provides new statistics and clear commands.
crates/xberg-cli/src/commands · high confidence
CLI extraction command restructured with batch support and enhanced output envelopes
The \xberg-cli\ extraction command has been refactored into a modular structure (\batch\, \images\, \manifest\, \runtime\, \timing\) to support new capabilities and improved reliability. Users can now process multiple documents in parallel via the new \batch\ command, which reads inputs from paths or manifest files (JSON/JSONL) and outputs results in JSON, Text, or TOON formats. The JSON and TOON output envelopes now consistently include timing metadata (\extraction\_time\_ms\, \stage\_timings\) and peak memory usage (\peak\_memory\bytes\), enabling better performance monitoring. Additionally, extracted images are now written to disk with predictable naming (\image\{index}.{format}\) when using text or TOON output, and the CLI runtime stack size has been increased to 16 MB to prevent crashes during deep extraction pipelines.
crates/xberg-cli/src · high confidence
Configurable multi-label classification for pages and chunks
The text classification module now supports configurable multi-label classification for both pages and document chunks. Pages are classified using a new \page\_classifier\ that respects a \multi\_label\ configuration flag, allowing multiple labels per page and aggregating results into a document-level label set. Additionally, a new \chunk\_classifier\ applies multi-label classification to individual document chunks, grouping them into batches for efficient LLM processing. Both classifiers use structured JSON responses from the LLM and support custom prompt templates.
crates/xberg/src/text/classification · high confidence
DOCX extraction now parses drawing objects and converts OMML math to LaTeX
The DOCX extractor now handles complex content that was previously ignored or lost. It parses drawing objects (inline and anchored images, shapes, and text boxes), exposing their properties and embedded text. Additionally, it converts Office Math Markup Language (OMML) equations into LaTeX, supporting fractions, radicals, matrices, and other mathematical structures. These changes ensure that visual and mathematical content in Word documents is preserved in the extracted output.
crates/xberg/src/extraction/docx · high confidence
Deterministic diagram recovery from vector sources
The diagram extraction module now uses a deterministic, geometry-based approach to recover diagram graphs from vector sources (SVG and PDF) and flat ODF drawings (.fodg). Instead of relying on probabilistic detection, it parses outlines, connectors, and labels directly from the source geometry, enabling exact graph recovery with preserved styling. The system includes specialized front-ends for SVG, PDF, and ODF formats, with shared matching logic for assembling nodes and edges from geometric primitives. This change improves accuracy for diagram extraction by eliminating heuristic-based detection and handling complex cases like clusters, arrowheads, and edge labels through precise geometric analysis.
crates/xberg/src/extraction/diagram · high confidence
Hermes plugin for xberg document extraction added
The xberg Hermes plugin package (version 1.3.0) is now available, providing the integration layer for the xberg document extraction engine. This update introduces the plugin structure and comprehensive skill documentation covering batch extraction, text chunking, keyword and language detection, table extraction, OCR capabilities, and output format selection. The plugin acts as a no-op adapter by default, allowing users to register custom Hermes tools, hooks, and commands via a local \register\ implementation.
plugin/.hermes/package, plugin/.hermes/plugins/xberg · high confidence
Initial JNI bindings for Android integration
The \crates/xberg-jni\ crate has been added, providing a Java Native Interface (JNI) shim for Android applications. This new component exposes native Rust functions (such as \nativeClassifyChunksOwned\) that allow Android apps to interact with the core \xberg\ extraction engine, handling data serialization and runtime management across the Java-Rust boundary.
crates/xberg-jni · high confidence
Initial PHP bindings release for xberg
This change introduces the first version of the PHP bindings for the xberg library, generated by the Alef codegen tool. It provides a native PHP extension (via ext-php-rs) and a comprehensive set of PHP stubs and interfaces, enabling PHP applications to perform document extraction, OCR, embedding, and chunk classification. The release includes the core \Xberg\ facade class with methods like \extract\, \extractBatch\, and \doctor\, as well as plugin interfaces for custom \DocumentExtractor\, \OcrBackend\, \EmbeddingBackend\, \PostProcessor\, \Renderer\, \RerankerBackend\, \TokenizerBackend\, and \Validator\ implementations. Additionally, it exposes configuration classes such as \AccelerationConfig\ and data structures like \ArchiveEntry\ and \Registry\ to manage presets and backends.
crates/xberg-php · high confidence
Initial release of the Dart/Flutter binding package
The \packages/dart\ directory now contains the complete, source-only \xberg\ package for Dart and Flutter, version 1.1.0. This package provides document extraction capabilities via \flutter\_rust\_bridge\, including a native library loader that resolves platform-specific binaries from the pub.dev package or downloads them to a user cache, and exposes the full extraction API (text, tables, images, metadata) through \XbergBridge\. The release also includes the necessary package infrastructure such as \pubspec.yaml\, \analysis\_options.yaml\, and generated bridge code.
packages/dart · high confidence
Initial release of the Node.js binding package
The \crates/xberg-node\ crate is introduced as the official TypeScript/Node.js binding for the Xberg document intelligence engine. This new package provides native NAPI-RS bindings that expose the core extraction capabilities—including text, tables, images, and metadata extraction from over 100 file formats—directly to Node.js applications. The release includes the necessary build infrastructure (\build.rs\), auto-generated TypeScript type definitions (\index.d.ts\), and the JavaScript loader (\index.js\) that handles platform-specific native library loading for macOS, Linux, and Windows. It also establishes the MIT license and initial documentation for the Node.js ecosystem.
crates/xberg-node · high confidence
Initial release of the xberg-ffi crate with comprehensive error handling
The new \xberg-ffi\ crate provides the auto-generated Foreign Function Interface bindings for the xberg library, enabling integration with other languages via the Alef code generator. This entry point establishes a robust error-handling layer that maps internal Rust errors (such as OCR, parsing, validation, and I/O failures) to specific, distinct error codes for consumers, while also safely catching Rust panics at the FFI boundary to prevent crashes in host applications.
crates/xberg-ffi/src · high confidence
Initial release of the xberg-libheif Rust library
This change introduces the \xberg-libheif\ crate, a new Rust wrapper for the \libheif\ library. It provides a comprehensive API for HEIF image processing, including context management for reading and writing HEIF files, image decoding and encoding with support for various compression formats (HEVC, AV1, JPEG, etc.), and detailed handling of color profiles, chroma subsampling, and image orientation. The library exposes types for error handling, encoder/decoder descriptors, and decoding options, enabling users to integrate HEIF support into their Rust applications.
crates/xberg-libheif/src · high confidence
Introduce Alef-generated documentation and language-specific README templates
The project now uses the Alef binding generator to produce documentation and per-language READMEs. New Jinja templates in \templates/docs/\ generate LLM-consumable reference files (\llms-body.md.jinja\, \llms.txt.jinja\) and AI agent skills for the API, CLI, and MCP server. In \templates/readme/\, a generic \language\_package.md.jinja\ and language-specific templates (e.g., \python.md\, \go.md\) provide installation, quick-start, and feature documentation for all 15 bindings, supported by partials for badges, features, and installation instructions.
templates · high confidence
Introduce GLiNER span-mode NER inference engine
Added a new Rust crate for zero-shot Named Entity Recognition using the GLiNER span-mode model. The implementation provides an ONNX Runtime backend (enabled via the \ort-backend\ feature) that handles preprocessing, tokenization, and inference, exposing a \Gliner\ struct for loading models and running inference. It also includes a \Gliner2\ engine for the V2 pipeline and a \candle\ backend feature for pure-Rust inference without ONNX. The crate defines configuration structs for inference parameters (threshold, max width, overlap policies) and runtime settings, along with robust input validation and error handling for model schema mismatches and invalid inputs.
crates/xberg-gliner/src · high confidence
Introduce PaddleOCR backend with PP-OCRv6 support and dual inference engines
Users can now use PaddleOCR for text extraction, featuring the new PP-OCRv6 model generation (defaulting to the 'small' tier) alongside legacy PP-OCRv5 support. The backend offers a choice between the native ONNX Runtime engine and a pure-Rust Tract engine, allowing operation on targets where ORT cannot link (e.g., WASM). Configuration includes explicit control over model tiers, language-specific recognition models (including Korean and Japanese), and table detection (disabled by default). The implementation handles auto-rotation, preserves vertical CJK reading order, and validates backend options before processing.
_crates/xberg/src/paddle\ocr · high confidence
Introduce Python v4 package with async API and OCR progress reporting
The \packages/python\ directory now contains the new Python v4 package (\xberg\), replacing previous iterations. This release provides a native Python binding with async/await support for document extraction, featuring a structured \ExtractionResult\ envelope and support for multiple OCR backends (Tesseract, PaddleOCR, Candle). A key behavioral addition is the ability to report per-page OCR progress via optional \on\_progress\ callbacks in \extract\ and \extract\_batch\ functions, allowing users to track extraction status in real-time. The package is auto-generated by the Alef toolchain and includes comprehensive type stubs and a MIT license.
packages/python · high confidence
Introduce Sceptre OCR backend with paragraph grouping and metadata support
A new Sceptre-based OCR backend is added to the system, supporting multiple inference engines (ONNX Runtime, tract, and candle) with lazy reader initialization and caching. The backend implements document assembly by grouping recognized lines into paragraph blocks based on geometry, merges word fragments sharing a line's y-center, and splits indent-marked paragraphs. It also forwards block, font, language, and table metadata to callers, declares backend-specific confidence meanings, and handles rotated pages. Resource usage is bounded via reader cache limits and execution semaphores.
_crates/xberg/src/sceptre\ocr · high confidence
Introduce TrOCR engine with image dimension validation
The candle-OCR backend now includes a TrOCR engine supporting four variants (base/large for printed and handwritten text) with pinned model revisions and SHA-256 checksums for secure loading. To prevent hallucinated output from aspect-ratio mismatches, the new image processor exposes a \dimensions\ function that allows callers to inspect raw raster shape before resizing, and the engine rejects whole-page inputs that do not match the expected text-line granularity.
crates/xberg-candle-ocr/src/models · high confidence
Introduce Xberg Ruby gem with native bindings and static type checking
The Ruby package now provides a native gem (xberg v1.3.0) that exposes the document extraction engine via a Ruby extension built with Magnus and rb\_sys. The binding includes idiomatic Ruby classes for configuration and results, supports batch processing, and integrates with the broader cross-language parity. To ensure code quality and type safety, the package ships with a RuboCop configuration for linting, RBS type signatures for static analysis, and a Steepfile to configure the Steep type checker, alongside an RSpec test suite that validates the generated class constructors.
packages/ruby · high confidence
Introduce Xberg document loader for LangChain
Added a new \XbergLoader\ in \integrations/python/langchain\ that integrates the \xberg\ library with LangChain. The loader supports extracting documents with configurable chunking and per-page splitting, mapping extraction results (including text, tables, and metadata) into LangChain \Document\ objects, and handling errors from the underlying \xberg\ extraction process.
integrations/python/langchain · high confidence
Introduce Zig language bindings for document extraction
This change adds the initial Zig binding package, providing idiomatic Zig wrappers over the existing C FFI surface. It includes the core generated API (\src/xberg.zig\) with explicit memory ownership and typed error sets, a \build.zig\ configuration that links against the \xberg\_ffi\ library, and a hand-written test suite (\test/xberg\_test.zig\) to verify FFI round-trips and error handling. The package is versioned at 1.3.0 and requires Zig 0.16.0+.
packages/zig · high confidence
Introduce engine-neutral inference seam with ONNX Runtime and tract backends
The inference layer now uses a unified abstraction (\InferenceBackend\ and \InferenceSession\) to run ONNX models, decoupling model execution from specific engine implementations. ONNX Runtime remains the default backend for native builds, while the pure-Rust \tract\ engine is used as the default on no-ORT targets (such as WASM and Android x86\_64). This change enables cross-engine parity testing and allows layout detection and auto-rotation features to run on tract when ORT is unavailable, while maintaining a consistent tensor interface (\InferenceTensor\) and error handling across both engines.
crates/xberg/src/inference · high confidence
Introduce native PDF engine with annotation, color, and extraction support
The \xberg-native-pdf\ crate is introduced as the new native PDF engine, providing core capabilities for reading and processing PDF documents. This includes a comprehensive annotation system supporting types like highlights, links, and form widgets, alongside a robust color management layer using the qcms backend for ICC profile handling. The engine also features configurable text extraction profiles optimized for different document types (e.g., academic, policy, scanned) and a content stream parser for graphics state and operator execution.
crates/xberg-native-pdf/src · high confidence
Introduce new layout detection engine with multi-backend support
The layout detection subsystem has been restructured into a new \LayoutEngine\ that supports multiple model backends (RT-DETR, YOLO, PP-DocLayout-V3, and custom ONNX models) and two inference engines (ONNX Runtime and pure-Rust Tract). This change introduces a unified configuration interface (\LayoutEngineConfig\) for selecting backends, setting confidence thresholds, and managing hardware acceleration. It also adds granular inference timing data (\DetectTimings\) to help users monitor preprocessing, inference, and postprocessing performance. The engine includes robust model management with atomic downloads and caching from Hugging Face Hub, and provides a WASM-compatible path for injecting model bytes directly.
crates/xberg/src/layout · high confidence
Introduce official Helm chart for Kubernetes deployment
Users can now deploy the Xberg document-intelligence framework on Kubernetes using the new official Helm chart (version 1.3.0). The chart provides a complete set of templates including Deployment, Service, Ingress, HorizontalPodAutoscaler, PodDisruptionBudget, and PersistentVolumeClaim for model caching. It supports configurable replica counts, resource limits, and an optional init container to handle cache directory ownership. The chart is published as an OCI artifact to GHCR and includes a JSON schema for values validation, allowing users to install via \helm install xberg oci://ghcr.io/xberg-io/charts/xberg --version 1.3.0\.
charts/xberg · high confidence
Introduce post-processor plugin system with staged execution
The \crates/xberg/src/plugins/processor\ module now provides a registry and trait-based system for custom post-processing of extraction results. Users can register plugins that transform or enrich extracted documents (e.g., adding metadata, cleaning text) via three execution stages: Early, Middle, and Late. Processors are executed in stage order, allowing for structured pipelines where foundational tasks like language detection happen before final enrichment steps. The system includes functions to register, unregister, and list processors, with built-in recovery for automatic built-in processors.
crates/xberg/src/plugins/processor · high confidence
Introduce safe Rust bindings for Tesseract OCR and Leptonica image processing
The \xberg-tesseract\ crate now provides a safe Rust interface to the Tesseract OCR engine and Leptonica image library. This includes a \TesseractAPI\ for initializing the engine, setting images, and extracting text, along with iterators (\PageIterator\, \ResultIterator\) to access detailed layout information such as bounding boxes, block/paragraph metadata, and symbol-level data. The crate also introduces a safe \Pix\ wrapper for Leptonica, enabling image preprocessing steps like grayscale conversion, thresholding, and deskewing before OCR. Memory safety is enforced through RAII patterns, and image dimension validation uses \i64\ arithmetic to prevent buffer overflows.
crates/xberg-tesseract/src · high confidence
Introduce self-hosted reranking with Qwen3 generative head and SHA-256 pinned presets
The reranking module now supports a new Qwen3 generative-reranker scoring head alongside the existing cross-encoder path, enabling users to select this model for relevance scoring. To ensure integrity, the self-hosted preset fleet is now pinned with a SHA-256 manifest, and the engine enforces thread-safe concurrent inference for ONNX models.
crates/xberg/src/reranking · high confidence
Introduce stateful WASM engine with injected JS bridges and browser-side layout/orientation detection
The \xberg-wasm\ crate now exposes a stateful \XbergEngine\ handle that accepts injected JavaScript objects for OCR and Named Entity Recognition (NER) backends, allowing the WASM module to delegate these tasks to the host environment with configurable timeouts. This engine also provides direct extraction capabilities. Additionally, the crate adds standalone WASM exports for document layout detection (RT-DETR) and page orientation detection (PP-LCNet), which accept raw ONNX model bytes from the JavaScript host to perform inference entirely within the browser without embedding large model weights in the binary.
crates/xberg-wasm/src · high confidence
Introduce structured-extraction preset system with embedded library and registry
The \xberg\ crate now includes a complete preset system for structured document extraction. This adds a JSON-based preset format (validated against a Draft 2020-12 meta-schema) and an in-memory registry that embeds a synthetic \generic\_document\ preset by default. Users can load this registry, resolve presets with custom schema overrides and context templates, and extend the registry with additional preset files from a directory at runtime. The system also provides a lightweight \PresetSummary\ projection for registry listing endpoints.
crates/xberg/src/presets · high confidence
Introduce validator plugin system for extraction quality gates
The \crates/xberg/src/plugins/validator\ module now provides a plugin-based system for validating extraction results. This introduces a \Validator\ trait that allows implementing custom validation logic (e.g., quality checks, compliance, or security scans) which, if they fail, cause the extraction to fail immediately. The module exposes functions to register, unregister, list, and clear validators, along with a global registry for managing these plugins.
crates/xberg/src/plugins/validator · high confidence
Introduce xberg doctor to probe backends and validate configuration
The \xberg doctor\ command now probes configured backends and reports what will actually execute on the host, answering whether issues stem from the document or the environment. It performs static configuration linting (detecting conflicts like \force\_ocr\ with \disable\_ocr\, or missing VLM configs), validates OCR and layout model cache integrity via checksums, and checks for stray files in xberg-owned cache directories. The report provides pass, warn, fail, or skip verdicts for each check, allowing users to identify misconfigurations or missing dependencies before processing documents.
crates/xberg/src/doctor · high confidence
Introduce xberg\_cli Python package for native binary management
Added the \xberg\_cli\ package, which provides a CLI entry point that automatically resolves and executes the native \xberg\ binary. The package supports pre-bundled binaries within wheels and falls back to a self-healing runtime downloader that fetches the correct asset from GitHub releases based on the platform target triple. The downloader includes security measures such as HTTPS-only redirects, SHA256 checksum verification, and protection against zip-slip/tar-slip archive extraction vulnerabilities.
_cli-proxy/pypi/xberg\cli · high confidence
Introduces DeepSeek-OCR and GLM-OCR candle backends with checksum-pinned weights
The candle-based OCR pipeline now includes two new backends: DeepSeek-OCR and GLM-OCR. DeepSeek-OCR automatically downloads weights from the pinned \deepseek-ai/DeepSeek-OCR\ repository, verifying them against a checked-in SHA-256 manifest, and intelligently selects floating-point precision (BF16, F16, or F32) based on the compute device to optimize memory usage. GLM-OCR supports multiple tasks (OCR, Table, Formula, Chart, Caption) and a layout-aware 'Paired' mode that detects regions before inference. Both backends use process-wide engine pools to avoid redundant weight loading and resolve compute devices once per process to eliminate repetitive logging.
_crates/xberg/src/candle\ocr · high confidence
Introduces OpenTelemetry-based telemetry and Prometheus metrics support
The telemetry module now provides structured observability for document extraction operations. It defines semantic conventions for span attributes and metrics under the \xberg.\*\ namespace, including counters for extraction totals and cache hits, histograms for duration and input/output sizes, and gauges for concurrent extractions. Span helpers are added to instrument extraction, pipeline stages, batch processing, OCR, and model inference, with file paths sanitized to prevent PII leakage. Additionally, an opt-in \init\_prometheus\ function installs a Prometheus exporter as the global OpenTelemetry meter provider, enabling scraping of these metrics via a \/metrics\ endpoint when the \prometheus\ feature is enabled.
crates/xberg/src/telemetry · high confidence
Introduces Rust-only extension seams for engine extensibility
The engine now exposes six new extension points (seams) in Rust, allowing advanced users to inject custom implementations for caching, LLM calls, model resolution, preset resolution, structured extraction policies, and progress reporting. Each seam provides a trait and a default implementation that preserves current behavior: the cache defaults to a no-op (no caching), the LLM client delegates to the existing liter-llm path, the model provider uses the existing on-demand download/cache logic, the preset resolver uses the embedded registry, the structured policy uses the existing heuristics, and the progress sink discards events by default. These seams are gated behind specific features (liter-llm, presets, heuristics, layout-detection, tokio-runtime) and are intentionally excluded from the language-binding surface, meaning they are available only to Rust consumers via the EngineBuilder.
crates/xberg/src/engine/seams · high confidence
Introduces a robust, filesystem-backed cache with automatic cleanup and namespace validation
The \crates/xberg/src/cache\ module now provides a complete, thread-safe caching system for extraction results. It includes automatic LRU-style eviction based on configurable age and size limits, implemented via a new \cleanup.rs\ module that scans and prunes stale \.msgpack\ files. To prevent path-traversal attacks, \namespace.rs\ enforces a strict allowlist on cache namespaces, rejecting any input containing path separators or leading dots. Cache keys are now versioned using a tag that incorporates the crate version, a manually bumped schema version, and a build identifier, ensuring that stale entries from previous extraction behaviors are invalidated. The core \GenericCache\ structure manages these entries with lock-poisoning-resistant synchronization.
crates/xberg/src/cache · high confidence
Introduction of xberg-cli with platform-specific binary linking
The xberg-cli crate has been added, providing a command-line interface for document extraction, MIME type detection, batch processing, embeddings, and cache management. The build script configures runtime library paths (rpath) for macOS and Linux to ensure the binary can locate shared dependencies correctly on these platforms.
crates/xberg-cli · high confidence
Java package license and documentation added
The Java binding package now includes an explicit MIT License file and a comprehensive README. The license clarifies that the software is provided by Kreuzberg, Inc. under the MIT terms, and the README provides installation instructions for Maven and Gradle, along with usage examples and configuration details for the Java binding.
packages/java · high confidence
Keyword extraction support via YAKE and RAKE algorithms
The \xberg\ crate now includes a keyword extraction module that allows users to automatically identify relevant keywords from extracted document content. This feature supports two algorithms—YAKE (statistical) and RAKE (co-occurrence based)—selectable via configuration. Users can tune extraction behavior through \KeywordConfig\, specifying parameters such as the algorithm, maximum number of keywords, minimum score thresholds, n-gram ranges, and language-specific stopword filtering. The extracted keywords are stored in the document's metadata and are accessible after processing. The implementation includes a post-processor that integrates into the extraction pipeline, handling initialization and ensuring keywords are only processed when explicitly configured.
crates/xberg/src/keywords · high confidence
LaTeX formula recognition via RapidLaTeXOCR
Added a new module for recognizing LaTeX from rasterized formula images using the RapidLaTeXOCR model set (image resizer, ViT encoder, and autoregressive transformer decoder). The implementation downloads and verifies four model files (resizer, encoder, decoder, tokenizer) by SHA256, applies grayscale normalization and padding constraints, and runs the decode loop in Rust with argmax sampling. It exposes utilities to check model cache status, probe model health, and list models in the manifest.
_crates/xberg/src/formula\recognition · high confidence
Language detection now reports per-language confidence, proportion, and script
The language detection module now exposes detailed per-language metrics in addition to the list of detected language codes. When \detect\_multiple\ is enabled, the system aggregates results across 200-character chunks to determine the proportion of text belonging to each language, along with the specific script (e.g., Latin, Cyrillic) and confidence score for each. This structured data is stored in \detected\_language\_confidences\, allowing consumers to assess the reliability of the detection beyond a simple language code match.
_crates/xberg/src/language\detection · high confidence
Native Rust LaTeX extractor added to xberg
A new native Rust-based LaTeX text extractor has been added to the xberg crate, replacing or supplementing previous extraction methods. This implementation parses LaTeX documents to extract structured content including metadata (title, author, date), section hierarchies (chapter, section, subsection), inline formatting (bold, italic, code), lists (itemize, enumerate, description), and tables (tabular, longtable). It also preserves inline and display math modes and handles various LaTeX commands like citations, references, and hyperlinks, outputting the result as Markdown-compatible text with associated table structures.
crates/xberg/src/extractors/latex · high confidence
Native extraction of legacy .doc files with structured paragraph data
The \xberg\ extractor now natively parses legacy Microsoft Word 97–2003 (.doc) files without requiring external tools like LibreOffice. This change introduces structured extraction that preserves paragraph-level metadata, including automatic list membership (with nesting depth and ordered/bullet distinction) and heading levels derived from document styles. It also correctly extracts OLE summary metadata (title, author, subject, etc.) and fixes previous issues where field instructions were emitted as text, non-breaking hyphens were dropped, and the piece table was not walked due to an incorrect FIB index.
crates/xberg/src/extraction/doc · high confidence
New Apple iWork extractors for Pages, Numbers, and Keynote
Added extractors for Apple iWork files (.pages, .numbers, .key) that parse the modern IWA container format (ZIP → Snappy → protobuf). The Pages extractor pulls document text and annotations; the Numbers extractor preserves table structure, cell values, sheet names, and formulas; and the Keynote extractor extracts slide text and speaker notes. All three enforce security limits, including a new max\_pages cap on slide counts and bounds on container reads to prevent decompression bombs.
crates/xberg/src/extractors/iwork · high confidence
New C FFI crate with generated bindings and build infrastructure
The \crates/xberg-ffi\ crate has been introduced to provide a stable C ABI for native integration, shared library distribution, and cross-language interop. This location contributes the C header generation logic (\build.rs\), the \cbindgen\ configuration (\cbindgen.toml\), and the public documentation (\README.md\). The generated header (\include/xberg.h\) exposes the core extraction, embedding, MIME detection, and plugin lifecycle functions, along with JSON marshaling helpers and handle lifecycle management, enabling Go, Java, and C\# bindings to interact with the Rust core without per-field accessor overhead.
crates/xberg-ffi · high confidence
New CI scripts for benchmarking, dependency setup, and artifact validation
This change introduces a suite of new shell and PowerShell scripts under \scripts/benchmarks/\ and \scripts/ci/actions/\ to support the CI pipeline. The benchmark scripts (\ensure-benchmark-harness-exists.sh\, \publish-corpus-cache.sh\, \restore-corpus-cache.sh\, \run-benchmark.sh\, \restore-binary-permissions.sh\) manage the reference corpus cache, validate the harness existence, and execute benchmark runs with quality and OCR support. The CI action scripts handle environment setup for ONNX Runtime across Linux, macOS, and Windows, download and stage PDFium binaries, prefetch Hugging Face embeddings, and run Node.js smoke tests. Additionally, Python test files are added to validate the benchmark workflow matrix and the PyMuPDF4LLM extraction adapter.
scripts · high confidence
New Candle-based GLiNER2 inference backend with LoRA adapter support
This change introduces a new inference backend for GLiNER2 using the Candle library, located in \crates/xberg-gliner/src/candle\. It provides a complete pipeline including a DeBERTa-v2 encoder, inference heads (span representation, count prediction, and scoring), and a decode loop that extracts entities based on thresholds and overlap policies. The backend supports loading models from local directories or in-memory buffers, with specific optimizations for WebAssembly (wasm32) such as F16 precision to reduce memory usage and a streaming loader to prevent out-of-memory crashes. It also includes runtime support for PEFT-format LoRA adapters, allowing users to load and merge adapters into the base model weights on non-wasm32 targets.
crates/xberg-gliner/src/candle · high confidence
New Djot format extraction and round-trip support
The xberg extractor now supports parsing Djot markup documents. It extracts metadata from YAML frontmatter, plain text, tables, and document structure (headings, links, code blocks). Additionally, it provides APIs to convert extracted Djot content back into Djot markup, preserving block structure, inline formatting, and attributes, and to render Djot source to HTML.
_crates/xberg/src/extractors/djot\format · high confidence
New Dockerfile variants for CLI, core, full, and native bindings
The project introduces a comprehensive set of new Dockerfiles to build and distribute the Xberg document intelligence platform. Dockerfile.binstall-musl produces a fully static, self-contained CLI binary for \cargo binstall\, while Dockerfile.cli, Dockerfile.core, and Dockerfile.full provide Alpine and Debian-based images for the CLI, core extraction, and full-featured environments respectively. Additionally, Dockerfile.manylinux-node and Dockerfile.manylinux-rustler build Node.js and Elixir native bindings with strict glibc floor enforcement, and Dockerfile.musl-build, Dockerfile.musl-ffi, and Dockerfile.musl-node deliver musl-compatible binaries and FFI libraries with bundled runtime dependencies. These images standardize the build process, enforce strict dependency vendoring (e.g., ONNX Runtime, PDFium, libheif), and ensure compatibility across Linux distributions and architectures.
docker · high confidence
New Elixir binding with Rustler NIF support
The Elixir package is now available, providing a native BEAM binding for document extraction via Rustler NIFs. This release includes the core \Xberg\ module for extracting text, tables, images, and metadata from over 100 file formats, along with comprehensive configuration structs for OCR, browser fallbacks, and hardware acceleration. The package is licensed under MIT and requires Elixir 1.14+ and Erlang/OTP 26+.
packages/elixir · high confidence
New HTML extraction engine with structured document support
The HTML extraction module has been replaced with a new implementation that converts HTML to Markdown or Djot while simultaneously building a structured document tree. This new engine captures inline images, extracts metadata, and handles complex structures like nested lists, tables, and MathML (converted to LaTeX). It also includes stack management to prevent crashes on large files in WASM environments and on native platforms.
crates/xberg/src/extraction/html · high confidence
New Kotlin Android binding package
The \packages/kotlin-android\ directory now contains a complete, generated Kotlin Android binding package (AAR) for the Xberg document extraction engine. This new package provides a JNI-backed library for mobile extraction workloads, including Gradle wrapper scripts, ProGuard rules, and a full set of generated Kotlin data classes and enums (such as \AccelerationConfig\, \AnnotationKind\, and \ArchiveEntry\) that mirror the core extraction API. It is published under the \io.xberg\ Maven namespace and supports both ARM64 and x86\_64 native architectures.
packages/kotlin-android · high confidence
New LangChain.js loader for Xberg document extraction
Added the \@xberg-io/langchain-xberg\ integration, providing an \XbergLoader\ that allows LangChain.js applications to load and extract content from documents using the Xberg engine. The loader supports loading from file paths, directories (via glob patterns), or raw byte buffers, and exposes Xberg's extraction capabilities—including chunking, page-level extraction, and metadata mapping (such as title, authors, keywords, and quality scores)—directly into LangChain's \Document\ objects.
integrations/node/langchain-xberg · high confidence
New LlamaIndex integration for xberg document extraction
This location introduces the \XbergReader\ and \XbergNodeParser\ components for the LlamaIndex Node.js SDK. The reader leverages the \@xberg-io/xberg\ Rust engine to extract content from 107 document formats, supporting file paths and raw bytes. The node parser then structures the extracted data into LlamaIndex \TextNode\ objects, prioritizing native chunks or structural elements to enable precise retrieval.
integrations/node/llamaindex-xberg · high confidence
New PDF extraction modules for bookmarks, embedded files, and adaptive layout gating
The PDF extraction backend now includes dedicated modules for extracting document bookmarks (outlines) and embedded file attachments, both with bounded traversal to prevent infinite recursion on malformed files. An adaptive layout gate has been added to pre-screen pages using cheap geometry signals, allowing the system to skip expensive layout detection on plain single-column text pages. Additionally, comprehensive error types and metadata structures have been introduced to expose scan confidence, fabricated text provenance, and layout gate reasons to consumers.
crates/xberg/src/pdf · high confidence
New REST API server for document extraction
The \xberg\ crate now includes a built-in HTTP API server (powered by Axum) that exposes endpoints for single and batch document extraction, MIME type detection, cache management, and OpenWebUI/Docling compatibility. The server supports configuration via file, environment variables, and per-request overrides, includes an in-memory job store for asynchronous extraction with cancellation, and provides an OpenAPI 3.1 schema for client generation.
crates/xberg/src/api · high confidence
New WASM bridge for injected OCR and local NER inference
The \crates/xberg-wasm/src/bridge\ module introduces the plumbing for two distinct recognition capabilities in the browser. First, it adds an injected OCR bridge that calls an external JavaScript \ocr\ backend, returning extracted text along with per-line geometry (bounding boxes) and confidence scores, with a 30-second timeout to prevent hanging. Second, it introduces a local NER (Named Entity Recognition) path via the \NerModel\ class, which loads a GLiNER2 model (weights, tokenizer, config) directly into WASM memory for synchronous, in-binary entity detection, offering a fallback when no external NER backend is injected.
crates/xberg-wasm/src/bridge · high confidence
New WordPerfect structured extraction crate
Added the \xberg-libwpd\ crate, which extracts structured content from WordPerfect documents on Linux, macOS, and Windows. It uses a C++ shim over libwpd/librevenge to serialize a versioned binary document model (wire format v1) containing text, formatting, tables, lists, links, fields, and metadata, which the Rust side decodes into typed \WpdDocument\, \WpdEvent\, and \WpdMetadata\ structures. The crate exposes \extract\_document\ and \is\_supported\ entry points, handles platform-specific compilation (including MSVC compatibility macros), and returns specific errors for unsupported formats, encryption, or platform constraints.
crates/xberg-libwpd/src · high confidence
New built-in post-processor plugins for classification, captioning, and more
This location introduces the registration and implementation of several new built-in post-processors that extend the extraction pipeline. The \captioning\ processor now generates VLM-based captions for extracted images, with bounded concurrency and proper LLM usage accounting. The \chunk\_classification\ and \page\_classification\ processors enable LLM-driven multi-label classification of document chunks and pages respectively. The \ner\ processor provides named-entity recognition with support for both ONNX and LLM backends. The \qr\ processor decodes QR codes from images and routes URL payloads into the document's URI collection. The \redaction\ processor applies pattern-based and NER-based PII redaction with an audit trail. The \summarization\ processor supports both extractive (TextRank) and abstractive (LLM) summarization strategies. The \translation\ processor enables LLM-based content translation. All processors are registered via a robust \register\_builtin\ function that attempts all registrations and aggregates failures, ensuring one broken processor doesn't silently prevent others from loading.
crates/xberg/src/plugins/processor/builtin · high confidence
New chunking module with structural boundary detection and page provenance
The \crates/xberg/src/chunking\ module has been introduced to provide text chunking capabilities, including structural boundary detection for plain text (identifying headers via ALL-CAPS, numbered sections, and title heuristics), page boundary validation and range calculation, and a heuristic semantic classifier for chunk types. It also adds per-page bounding-box aggregation and node-ID backreferences to chunks, allowing users to map chunks back to their source document structure and page locations.
crates/xberg/src/chunking · high confidence
New configuration types for acceleration, concurrency, classification, and output formats
The core configuration module now exposes explicit settings for hardware acceleration (ONNX execution providers like CUDA and CoreML), concurrency limits (thread budgets and OCR session caps), and post-processing (page and chunk classification with LLM-backed definitions). It also introduces configuration for output formatting (including DocTags and Jupyter cell rendering), content filtering (headers, footers, watermarks, and repeating text), and specialized extractors (CSV delimiters, GeoJSON coordinate bounds, HTML theming, and email codepages). These changes allow users to fine-tune performance, control extraction detail, and customize output structure without modifying code.
crates/xberg/src/core/config · high confidence
New derivation pipeline for internal documents
The \crates/xberg/src/extraction/derive\ module has been introduced to bridge the internal flat document representation and public-facing types. This new pipeline handles relationship resolution (converting key-based targets to indices), tree reconstruction from flat elements into a hierarchical \DocumentStructure\, and content string derivation. It also includes per-page content building and formatting logic, ensuring that page-level content matches the selected output format (Markdown, HTML, etc.) while correctly handling container markers and geometry. The module also includes comprehensive tests for formula projection, deduplication, and custom output format handling.
crates/xberg/src/extraction/derive · high confidence
New deterministic test fixture generator for integration tests
A new CLI tool (\\generate\_test\_fixtures\\) has been added to scaffold deterministic test fixtures for integration tests. It generates binary files and corresponding JSON ground-truth sidecars for DOCX track-changes, ODT tracked changes, XLSX revision headers, PPTX comments, PDF incremental updates, paired diff inputs, and security edge cases (such as DDE formulas and oversized embedded files).
_tools/generate\_test\fixtures · high confidence
New document orientation detection using PP-LCNet model
Added a new document orientation detection module that uses the PP-LCNet\_x1\_0\_doc\_ori model to determine page rotation (0°, 90°, 180°, or 270°) with confidence scoring. This replaces the previous reliance on Tesseract's DetectOrientationScript, which was prone to crashes on raw images without DPI metadata. The implementation supports both native targets (downloading models from HuggingFace) and WASM (accepting model bytes from the host), and is integrated into the OCR pipeline when auto-rotation is enabled.
_crates/xberg/src/doc\orientation · high confidence
New document processing heuristics for chunking, confidence, and structured extraction
The \xberg\ crate now includes a new \heuristics\ module that provides configurable logic for document processing decisions. Users can now control chunking strategies via \HeuristicsConfig\ (e.g., file size, page count, and text-layer thresholds) and \UserChunkConfig\ overrides. The module also introduces a confidence scoring system (\ConfidenceSignals\) that combines text coverage, OCR aggregate confidence, and schema compliance to evaluate extraction quality. Additionally, it provides multi-document boundary detection for splitting PDFs into separate documents and a structured extraction call-mode heuristic (\choose\_call\_mode\) to optimize LLM usage by selecting between text-only, vision-only, or fallback modes based on document type and content density.
crates/xberg/src/heuristics · high confidence
New document summarization module with extractive and LLM backends
Added a new \summarization\ module to the \xberg\ crate that provides two methods for generating document summaries. The default backend uses a pure-Rust, deterministic TextRank algorithm (TF-IDF cosine similarity and PageRank) to extract key sentences, supporting multiple languages via stopword lists and capping output length by token count. An optional abstractive summarization backend is available behind the \summarization-llm\ feature flag, which sends document text to an LLM via the shared text-completion helper to generate a concise prose summary, while also capturing LLM usage metrics for tracking.
crates/xberg/src/text/summarization · high confidence
New extraction modules for AsciiMath, blank detection, capacity estimation, DocTags, and email formats
The extraction engine now includes dedicated modules for converting AsciiMath to LaTeX (via MathML), detecting blank pages based on non-whitespace character counts, and estimating string buffer capacities to optimize memory allocation for various document formats. It also adds support for parsing Docling DocTags into internal documents and introduces comprehensive extraction for email formats (.eml and .msg), including handling of nested messages, embedded MSG attachments, and RTF content with security budgeting to prevent excessive memory usage.
crates/xberg/src/extraction · high confidence
New first-party integrations for LangChain.js, LlamaIndex.TS, n8n, and Spring AI
This release adds four new first-party integrations that connect Xberg document extraction to popular AI and workflow frameworks. For Node.js, the \@xberg-io/langchain-xberg\ package provides an \XbergLoader\ for LangChain.js, while \@xberg-io/llamaindex-xberg\ offers an \XbergReader\ and \XbergNodeParser\ for LlamaIndex.TS, enabling structure-aware node splitting. For workflow automation, the \@xberg-io/n8n-nodes-xberg\ community node adds a Document Extract operation to n8n. For Java, the \spring-ai-xberg\ package implements the Spring AI \DocumentReader\ interface, supporting batch extraction and rich metadata mapping. All integrations run locally in-process via the Xberg native binding, require no API keys, and are versioned in lockstep with the core Xberg release.
integrations · high confidence
New generic structured-extraction mechanism for PDF and image inputs
The \crates/xberg/src/engine/structured\ module introduces a reusable, policy-free engine for extracting structured data from documents. It handles PDF and image rasterization, token-aware batching of pages for vision-LLM calls, schema validation and merging, and citation fusion that enriches LLM output with OCR provenance. The system also includes a prompt-assembly layer that builds system and user prompts with nonce-fenced untrusted content and supports text-only, vision-only, and fallback modes. All configuration (DPI, token budgets, merge strategies, citation thresholds) is passed as caller-supplied parameters, ensuring the mechanism remains generic and free of hardcoded defaults.
crates/xberg/src/engine/structured · high confidence
New local Whisper transcription engine with audio decoding and metadata extraction
This change introduces the internal transcription pipeline for the \transcription\ feature, adding modules to decode arbitrary audio files into 16 kHz mono PCM, run local Whisper ONNX inference for speech-to-text, and extract audio metadata. The engine supports Tiny, Base, Small, Medium, and LargeV3 models, downloading them from Hugging Face Hub (using \onnx-community\ for smaller models and \Xenova\ for Medium/LargeV3) and caching them locally. It handles multi-channel audio down-mixing, sample rate resampling, and timestamped segment generation. Additionally, it extracts standard audio tags (title, artist, duration, etc.) from various container formats.
crates/xberg/src/transcription · high confidence
New markdown footnote and citation parsing module
The \xberg\ crate now includes a new \markdown\_footnotes\ module that provides utilities for parsing standard markdown footnotes and structured citation conventions. Users can now find footnote anchor references and definitions, detect inference markers, identify unmarked claims, and parse structured citation blocks (extracting source, locator, and excerpt). The module also exposes a \FootnoteConfig\ to control whether structured citation parsing is enabled.
_crates/xberg/src/text/markdown\footnotes · high confidence
New native JATS and OPML document extractors
The JATS extractor now handles Journal Article Tag Suite XML documents, extracting rich metadata (title, authors, DOI, PII, dates, journal info), article abstracts, section hierarchies, tables, figures, citations, and inline annotations (bold, italic, links, etc.). The OPML extractor provides native Rust-based parsing for outline structures, extracting metadata from the head section and preserving the outline hierarchy with indentation in the content.
crates/xberg/src/extractors/jats · high confidence
New native PDF extraction modules for annotations, forms, hierarchy, images, and metadata
The \crates/xberg/src/pdf/native\ directory now contains dedicated modules for extracting PDF annotations, form fields (AcroForm and XFA), heading hierarchy, embedded images, and document metadata. These modules map the \xberg\_native\_pdf\ backend types to Xberg's public models, handling specific extraction logic such as visible free-text annotation filtering, hybrid form field deduplication, font-metric-based heading detection, parallel image decoding with row-padding safety, and structured page boundary reporting.
crates/xberg/src/pdf/native · high confidence
New npm proxy package for installing the xberg CLI
The \cli-proxy/npm\ directory now contains a new npm package (\xberg-cli\) that acts as a proxy to install the \xberg\ CLI binary. The \install.js\ script detects the user's platform (Windows, Linux, macOS) and architecture, fetches the appropriate pre-compiled binary from the \xberg-io/xberg\ GitHub releases, verifies its SHA256 checksum against the release's \SHA256SUMS\ asset, and installs it locally. This allows users to run the CLI via \npx xberg-cli\ without needing to manually download or configure the binary. Tests in \test.mjs\ verify the logic for filtering out non-CLI artifacts (like FFI bindings or native libraries) and selecting the correct archive for the target platform.
cli-proxy/npm · high confidence
New office metadata extraction module for Office and OpenDocument formats
The \xberg\ crate now includes a new \office\_metadata\ module that extracts comprehensive metadata from Office Open XML documents (DOCX, XLSX, PPTX) and OpenDocument files (ODT). This module parses \docProps/core.xml\, \docProps/app.xml\, and \docProps/custom.xml\ to expose Dublin Core fields (title, creator, dates), application-specific statistics (word/page counts, editing time), and user-defined custom properties. It also extracts ODT metadata from \meta.xml\, including document statistics and generator information. Security flags (password protection, read-only restrictions) are decoded from the raw \DocSecurity\ bit field into named boolean flags for easier consumption.
_crates/xberg/src/extraction/office\metadata · high confidence
New plugin traits for embedding, reranking, tokenization, and rendering
The plugin system now exposes dedicated traits and global registries for in-process embedding backends (\EmbeddingBackend\), reranker backends (\RerankerBackend\), tokenizers for chunking (\TokenizerBackend\), and document renderers (\Renderer\). These additions allow users to register custom implementations for vector generation, result re-ranking, token-budgeted chunking, and output formatting via the new \register\\\ and \list\\\ functions, while the \mod.rs\ module re-exports these capabilities for public use.
crates/xberg/src/plugins · high confidence
New rendering infrastructure and output formats
The rendering subsystem has been restructured to use a shared \InternalDocument\-based architecture, introducing a new \common.rs\ module for nesting state, annotated text, and footnote handling. This change adds support for several new output formats: Docling DocTags (with bounding box location tokens for PDFs), Djot markup, and Graphviz DOT for recovered diagrams. It also introduces a new styled HTML renderer (\html\styled.rs\) that uses direct \kb-\\ class hooks and CSS custom properties instead of the previous comrak-based approach, and adds a dedicated HTML renderer for PDF annotations. Existing renderers (Markdown, plain text, JSON, etc.) have been refactored to use the new common infrastructure, improving consistency and fixing issues like panic on mid-codepoint text annotations.
crates/xberg/src/rendering · high confidence
New static (model2vec) embedding backend and SHA-256 pinned presets
Users can now generate dense text embeddings without ONNX Runtime via a new pure-Rust static backend (gated by the \static-embeddings\ feature), which is the only dense-embedding option on WASM and Android x86\_64 emulator targets. The module also introduces a self-hosted preset fleet (e.g., all-MiniLM-L6-v2, bge-base-en-v1.5, gte-modernbert-base, arctic-embed-m-v2.0, qwen3-embedding-0.6b, potion-base-8m) whose model files are pinned with SHA-256 checksums in \presets.sha256sum\ to ensure integrity. Additionally, the embedding engine now supports thread-safe concurrent inference and provides cache management APIs to evict specific models or clear all resident engines, with configurable limits on the number of resident engines.
crates/xberg/src/embeddings · high confidence
New structured type definitions for PDF annotations, diagrams, and document structure
The \crates/xberg/src/types\ module now includes comprehensive type definitions for several new extraction capabilities. \annotations.rs\ introduces \PdfAnnotation\ and \PdfAnnotationType\ to expose PDF annotation details (such as highlights, links, and comments) including bounding boxes and marked text. \diagram.rs\ adds \DiagramGraph\, \DiagramNode\, and \DiagramEdge\ to represent vector diagrams recovered from sources like SVG. \document\_structure.rs\ and \builder.rs\ define the \DocumentStructure\ tree and \DocumentStructureBuilder\ for hierarchical document representation. Additionally, \classification.rs\ and \entity.rs\ provide types for page classification and named-entity recognition results, while \djot.rs\ defines the \DjotContent\ structure for Djot document extraction.
crates/xberg/src/types · high confidence
New text quality scoring and OCR confidence capping
The \xberg\ text module now includes a \QualityProcessor\ that calculates a text cleanliness and readability score (0.0–1.0) based on OCR artifacts, script/style noise, navigation chrome, and structural cues. When OCR is used, this quality score is capped at the word-count-weighted mean OCR recognition confidence to prevent high shape-clarity scores on low-confidence OCR results. The module also provides SIMD-accelerated UTF-8 validation and Windows codepage mapping for RTF and email extraction.
crates/xberg/src/text · high confidence
New text-processing utilities and improved encoding provenance in xberg
The \xberg\ utils module now includes several new capabilities and behavioral improvements. It introduces JSON key-casing conversion utilities (\snake\_to\_camel\ and \camel\_to\_snake\) for transforming nested JSON structures, and lightweight markdown header detection for chunking and quality features. A new string interning pool (\InternedString\) is available to reduce memory allocation for repeated strings like MIME types. Text quality processing has been expanded with a \clean\_extracted\_text\ function that normalizes whitespace, removes OCR artifacts, and strips navigation boilerplate. Crucially, the encoding decode logic now reports provenance via \DecodeOutcome\, allowing callers to detect lossy decodes even when the \quality\ feature's mojibake cleanup would otherwise strip replacement characters. XML extraction now uses an \EntityReader\ to coalesce text nodes and resolve entity references, preventing data loss from split text events.
crates/xberg/src/utils · high confidence
New token reduction pipeline for text compression
The \crates/xberg/src/text/token\_reduction\ module introduces a configurable text reduction pipeline that compresses input text to lower token counts while preserving meaning and structure. Users can select from five intensity levels (Off, Light, Moderate, Aggressive, Maximum) via \TokenReductionConfig\. The pipeline applies stopword removal, whitespace normalization, and HTML comment stripping, while offering options to preserve Markdown formatting, code blocks, and specific regex patterns. It includes a CJK tokenizer for Chinese and Japanese characters, semantic analysis for importance-based filtering and hypernym compression, and SIMD-optimized text processing for performance.
_crates/xberg/src/text/token\reduction · high confidence
New token reduction pipeline with configurable text compression
The \crates/xberg/src/text/token\_reduction/core\ module introduces a new \TokenReducer\ that compresses text into fewer tokens using four reduction levels (Light, Moderate, Aggressive, Maximum). The pipeline normalizes punctuation, filters common words based on frequency and length, and selects the most important sentences using a scoring system that considers position, word count, and content density. It also supports semantic analysis for aggressive/maximum modes and handles CJK text via a universal tokenizer. Users can configure the reduction level, language hint, and whether to preserve important words (e.g., acronyms, numbers).
_crates/xberg/src/text/token\reduction/core · high confidence
PDF extraction engine restructured with layout-aware reading order and OCR fallback
The PDF extractor module has been reorganized into dedicated components for extraction logic, layout hint conversion, layout execution, and page content management. This change introduces layout-guided reading order reconstruction that projects text spans onto detected regions to fix multi-column and rotated text issues, and implements a chunked layout detection runner that bounds memory usage and retries on CPU when GPU acceleration fails. The engine now supports a new Pdfium backend for native text extraction, enforces per-document page limits for security, and improves OCR fallback handling by preserving document structure and reporting per-page confidence.
crates/xberg/src/extractors/pdf · high confidence
PHP 8.2+ document extraction bindings with plugin architecture
The PHP package now provides a modern, type-safe API for PHP 8.2+ that allows users to extract text, tables, images, metadata, and code intelligence from over 100 file formats. The package includes a comprehensive plugin system, exposing interfaces for custom DocumentExtractors, OCR backends, embedding models, post-processors, renderers, rerankers, tokenizers, and validators, enabling deep integration of custom logic into the extraction pipeline.
packages/php · high confidence
PaddleOCR engine now supports the Tract inference backend
The OCR engine now runs on the pure-Rust Tract backend in addition to the existing ONNX Runtime (ORT) backend. This enables text detection and recognition on platforms where native ORT cannot link, such as wasm32 and the Android x86\_64 emulator. The change introduces a runtime-neutral inference seam, allowing users to select the engine at load time and ensuring parity between ORT and Tract for detection, angle classification, and recognition.
crates/xberg-paddle-ocr/src · high confidence
Python bindings now report per-page OCR progress
The Python bindings for xberg-py now expose a progress callback mechanism for document extraction. Users can pass an \on\_progress\ listener to \extract\_with\_progress\ and \extract\_batch\_with\_progress\, which receives \ocr\_page\ events containing the current page number, total pages, completed count, backend name, and input index. This allows applications to track OCR processing status in real-time without blocking.
crates/xberg-py · high confidence
Recover graph structure from SVG diagrams
The SVG extraction module now reconstructs diagram graphs (nodes, edges, and labels) from vector SVG files. It uses \usvg\ to parse shapes and connectors, while a separate XML pass recovers text labels and titles that \usvg\ drops. The implementation includes safeguards against stack overflows from deeply nested SVGs and handles various label sources, including standard text elements and foreign objects.
crates/xberg/src/extraction/diagram/svg · high confidence
Restored xberg-pdfium-render crate with Pdfium 7678 bindings
The xberg-pdfium-render crate has been re-added to the workspace, providing PDF rendering capabilities via the Pdfium library (API version 7678). This update includes generated Rust bindings for both dynamic and static linking, a memory-based font provider for loading fonts without system installation, and a comprehensive error handling system. Users can now load PDF documents, render pages to images, and manage annotations and form fields using the restored library interface.
crates/xberg-pdfium-render/src · high confidence
Reversible redaction with per-person erasure
The redaction engine now supports a \TokenReplace\ strategy that substitutes PII with structured tokens (e.g., \\[EMAIL\_1\]\) instead of static masks. This enables reversible redaction: by providing a passphrase, the original text can be recovered from an encrypted rehydration map. Additionally, the system now supports per-person erasure, allowing specific individuals' data to be completely removed from the rehydration map even after redaction, ensuring that sensitive information can be fully forgotten upon request.
crates/xberg/src/text/redaction · high confidence
Semantic chunking with topic-aware merging and page provenance
The semantic chunking module in \crates/xberg/src/chunking/semantic\ now splits text into fine-grained segments, detects topic boundaries using embedding similarity (when the \embeddings\ feature is enabled), and merges segments into coherent chunks that respect those boundaries and a configurable size budget. Merged chunks include overlap from the previous group to maintain context, and each chunk now carries precise page provenance (\page\_spans\, \first\_page\, \last\_page\) and heading context, ensuring that structural and pagination information is preserved through the semantic chunking pipeline.
crates/xberg/src/chunking/semantic · high confidence
Spring AI integration now supports batch extraction and granular document splitting
The Spring AI integration has been refactored to support batch processing of multiple resources via a single Xberg call, significantly improving performance for multi-file workflows. The reader now splits extracted content into Spring AI Document instances using a priority-based strategy: it prefers chunks, then elements, then pages, and finally the whole document, ensuring the highest granularity available. Additionally, the integration now maps rich metadata from the extraction engine—including chunk indices, element types, bounding boxes, detected languages, and quality scores—into the resulting documents, and allows users to attach custom metadata via the builder.
integrations/java/spring-ai · high confidence
Swift package scaffolding and plugin trait bridges
The Swift package now includes the foundational scaffolding required for integration, including an \.editorconfig\, \.gitignore\, \.swiftformat\, and a \LICENSE\ file. It provides a \RustBridge\ module with core FFI marshalling types (e.g., \RustString\, \RustVec\) and generated protocol definitions for plugin traits—\SwiftPluginBridge\, \SwiftDocumentExtractorBridge\, \SwiftEmbeddingBackendBridge\, \SwiftOcrBackendBridge\, and \SwiftPostProcessorBridge\—along with their corresponding Box classes that expose \alef\\\ shim methods to Rust. A minimal \Demo\ target is also added to verify that the \Xberg\ module loads successfully on macOS and iOS.
packages/swift · high confidence
Unified NER backend trait with ONNX, Candle, and LLM implementations
The NER module now exposes a shared \NerBackend\ trait that allows the redaction engine and NER post-processor to swap detection backends without code changes. Three backends are provided: an ONNX-based GLiNER backend (\gline\) that downloads pinned models from \xberg-io/gliner-models\ with SHA-256 verification; a pure-Rust Candle backend (\candle\) for GLiNER2 inference without ONNX Runtime, supporting local loading with optional LoRA adapters and in-memory loading for WebAssembly; and an LLM-based backend (\llm\) that uses structured JSON prompts for zero-shot detection. All backends handle long documents by splitting input into overlapping windows to prevent silent truncation and ensure every occurrence of a detected entity is reported for complete redaction.
crates/xberg/src/text/ner · high confidence
Unified extraction API and new URL discovery capability
The extraction module now exposes a unified public API (\extract\, \extract\_batch\) that delegates to a process-global default engine, simplifying the internal structure while maintaining stable signatures for bindings. Additionally, a new \map\_url\ function is available (when the \url-ingestion\ feature is enabled) to discover URLs and sitemaps from a given URI without extracting document content, returning a \MapResult\ for use in building crawl queues or validating scope.
crates/xberg/src/core/extract · high confidence
Unified plugin registry with global singletons and built-in renderers
The plugin system has been restructured into a centralized registry module that manages all plugin types (OCR, embedding, extraction, post-processing, rendering, reranking, tokenization, and validation) through dedicated, thread-safe registries exposed as global singletons. This change introduces a new \RendererRegistry\ that registers built-in renderers (Markdown, HTML, Djot, DocTags, Graphviz DOT, and plain text) and ensures public renderers receive full document structure for layout-aware output. It also adds registries for embedding, reranker, and tokenizer backends that allow host-language bridges to register in-process implementations, removing the need to download external models for these capabilities. The \PostProcessorRegistry\ now includes a generation counter to help consumers detect and invalidate stale caches when plugins are added or removed.
crates/xberg/src/plugins/registry · high confidence
Vendored shared inference infrastructure for VLM-OCR backends
The \crates/xberg-candle-ocr/src/vendor\ directory now contains a vendored subset of the \jhqxxx/aha\ library (Apache-2.0), providing the foundational image processing, tensor utilities, and model architecture primitives required by the DeepSeek-OCR and PaddleOCR-VL 1.5 backends. This includes image loading and resizing helpers (\aha::image\), core neural network modules like attention mechanisms and MLPs (\aha::modules\), RoPE position embedding implementations (\aha::rope\), and the Qwen2 decoder structure (\aha::qwen2\). These components establish the shared inference interface (\InferenceModel\) and multimodal data handling (\MultiModalData\) that the specific OCR model implementations will consume.
crates/xberg-candle-ocr/src/vendor · high confidence
Xberg plugin v1.3.0: new opencode integration and updated skill documentation
The plugin has been updated to version 1.3.0, introducing a new Opencode integration that exposes \xberg\_extract\, \xberg\_detect\, and \xberg\_formats\ tools for local document processing. The plugin's documentation and skills have been refreshed to reflect the current 107 supported formats and 141 extensions, with updated guides for batch extraction, chunking, table extraction, OCR, and keyword extraction.
plugin · high confidence
Removals
Removal of core file extraction and parsing modules
The \src\ directory has removed the entire file extraction and parsing implementation, including the public API (\extraction.py\), internal extraction logic (\\_extraction.py\), helper utilities (\\_string.py\, \\_sync.py\), exception definitions (\exceptions.py\), and MIME type mappings (\mime\_types.py\). This eliminates the library's ability to extract text from PDFs, images, and various document formats via Pandoc and Tesseract.
src · high confidence
Security
Hardened Tesseract source archive extraction and caching
The build system now includes strict security and stability measures when downloading and unpacking the Tesseract and Leptonica source archives. It enforces limits on the number of archive entries and total uncompressed size to prevent resource exhaustion, rejects symbolic links to avoid path traversal attacks, and validates that all files remain within the expected root directory. Additionally, the cache mechanism verifies the integrity of downloaded artifacts using SHA-256 hashes and ensures the source tree is complete before use, while also normalizing Windows paths to ensure compatibility with CMake and native toolchains.
_crates/xberg-tesseract/build\support · high confidence
Secure archive extraction with path confinement and decompression limits
The archive extraction module now enforces strict security limits on ZIP, TAR, 7Z, and GZIP formats to prevent decompression bombs and excessive memory usage. All member reads are bounded by \max\_content\_size\ and \max\_archive\_size\ limits, rejecting archives that exceed these thresholds. Additionally, the \ArchiveEntry\ type now includes a \confined\_path\ method that normalizes and validates file paths, rejecting dangerous patterns such as absolute paths, directory traversal (\..\), NUL bytes, and Windows drive/UNC prefixes to prevent path traversal vulnerabilities.
crates/xberg/src/extraction/archive · high confidence
Architecture
OCR processor refactored into modular submodules with concurrency and caching fixes
The OCR processor implementation has been reorganized into focused submodules (\api\_pool\, \config\, \execution\, \validation\) to improve code structure and maintainability. This change introduces a bounded resource pool for Tesseract API handles to cap concurrent recognition sessions, preventing resource exhaustion. It also fixes critical caching issues by ensuring the cache key includes resolved tessdata paths, security limits, and all Tesseract engine variables, preventing stale or incorrect results from being served. Additionally, the processor now correctly handles image preprocessing, DPI normalization, and language validation, ensuring consistent OCR output across different configurations and environments.
crates/xberg/src/ocr/processor · high confidence
PDF structure pipeline restructured into modular components
The PDF structure extraction logic has been reorganized into distinct modules (\adapters\, \assembly\, \classify\, \constants\, \geometry\, \layout\_classify\, \layout\_debug\, \lines\) to improve maintainability and isolate concerns. This change introduces a new \OcrFontSizeScale\ type to correctly handle mixed OCR coordinate systems, adds diagnostic environment variables (e.g., \XBERG\_LAYOUT\_NO\_DEMOTE\) for debugging layout overrides, and implements spatial geometry primitives for bounding box operations. It also refines text assembly logic with new constants for glyph fragmentation repair and segment spacing, ensuring more accurate paragraph and heading classification.
crates/xberg/src/pdf/structure · high confidence
Behavioural changes
Benchmark harness introduces vendor/reference corpus split and structured aggregation schema
The benchmark harness now separates its PDF test corpus into two redistribution classes: 73 vendor fixtures (permissive/PD sources served from a public GCS bucket) and 92 reference fixtures (license-restricted sources materialized on demand into a private GCS cache). This split is documented in CORPUS.md and enforced via a strict ground-truth validation gate in CI. The harness also adopts Aggregation Schema v2.9.0 for its consolidated results, introducing a failure\_summary and format\_support matrix, while tightening percentile reporting with v2.10.0-style nulling of p95/p99 when sample counts are too low. To support these workflows, the harness ships with a set of predefined cohorts (e.g., layout-pdf-fast, native-office-fast, ocr-images-fast) that select fixtures in a fixed order with deterministic batch sizes, and includes an initial baseline JSON capturing per-document quality and performance metrics for regression tracking.
tools · high confidence
CLI overrides refactored into domain-specific submodules with validation and application logic
The CLI extraction override handling has been restructured from a single large file into a modular set of submodules (analysis, chunking, general, html, layout, llm, ocr, output, pdf) under the overrides directory. This change introduces explicit validation for CLI flags (e.g., rejecting invalid DPI ranges, zero concurrency, or contradictory layout settings) and ensures that field-specific flags (like --chunk-size or --ocr-backend) correctly materialize their respective configuration sections even when the main feature flag (like --chunk or --ocr) is not explicitly set. It also adds a warning when layout detection is enabled but the output format is plain, and standardizes LLM API key resolution across all LLM-backed features.
crates/xberg-cli/src/commands/overrides · high confidence
Configurable allowed hosts for MCP HTTP transport
The MCP HTTP transport now supports an allowlist of additional hosts to extend the default loopback-only restriction, enabling operation behind reverse proxies or ingress controllers. Users can configure these extra hosts via the \--allowed-hosts\ CLI flag, the \XBERG\_MCP\_ALLOWED\_HOSTS\ environment variable, or the \\[mcp\] allowed\_hosts\ key in their configuration file (TOML, YAML, or JSON), with CLI taking precedence over environment, which takes precedence over the config file. This change resolves the inability to run the MCP server behind proxies that forward different \Host\ headers, while maintaining the default security posture against DNS-rebinding attacks when no override is provided.
crates/xberg/src/mcp · high confidence
Configurable async job timeout with unified server configuration
The server configuration module now exposes a \job\_timeout\_secs\ setting (defaulting to 600 seconds) that acts as a fallback cap for \POST /extract-async\ jobs, ensuring that requests omitting or explicitly nullifying a per-request timeout still terminate after a bounded duration to prevent denial-of-service risks. This setting is part of a broader \ServerConfig\ system that supports loading from TOML, YAML, or JSON files, allows all settings to be overridden via environment variables (such as \XBERG\_JOB\_TIMEOUT\_SECS\), and enforces strict validation on host, port, CORS origins, upload limits, and the new job timeout to reject invalid configurations before the server binds.
_crates/xberg/src/core/server\config · high confidence
Core extraction engine restructured with new batch mode, diagnostics, and multi-document PDF splitting
The core extraction module has been reorganized into dedicated sub-modules to improve reliability and performance. A new batch processing mode (batch\_mode.rs) uses task-local flags to enable parallelism for CPU-intensive work, while a centralized diagnostics system (diagnostics.rs) deduplicates processing warnings to prevent log flooding. Multi-document PDFs can now be split and extracted in a single pass (split.rs) using either heuristic detection or explicit page ranges. The engine also introduces a centralized image re-encoding helper (image\_encode.rs) for format conversion, optimized file I/O with memory-mapped reading for large files (io.rs), and a robust path resolver (path\_resolver.rs) that safely handles relative image references and symlinks. Additionally, a compile-time registry of 55 standardized format fields (formats.rs) now serves as the single source of truth for metadata validation across all language bindings.
crates/xberg/src/core · high confidence
Documentation site migrated to Astro and Starlight
The documentation site has been rebuilt using Astro with the Starlight theme and the @xberg-io/docs-theme package. This migration introduces a new site structure with specific sidebar navigation for Getting Started, Guides, Concepts, Integrations, and Reference sections. It also includes configuration for LLM-friendly content indexing via starlight-llms-txt and adds redirects for index-less section roots to prevent 404 errors.
docs-site · high confidence
Domain rebrand from kreuzberg.dev to xberg.io
The project has officially rebranded its domain from kreuzberg.dev to xberg.io. This change updates the domain references across the codebase, including documentation, configuration files, and potentially package metadata, to reflect the new identity.
(repo-wide) · high confidence
Dynamic API snippet tabs based on target platform
The documentation site now uses a new ApiSnippetGroup component to display generated code examples. This component dynamically renders tabs only for the programming targets that have available snippets for a specific topic, rather than showing a static list of all languages. It correctly distinguishes between different bindings (such as Node.js and WebAssembly) that share the same language by keying tabs on the target platform, ensuring users see distinct code samples for each supported environment.
docs-site/src/components · high confidence
EPUB extraction rewritten with MathML-to-LaTeX conversion and hardened parsing
The EPUB extractor has been refactored into dedicated modules (content, metadata, parsing) to improve reliability and security. MathML formulas in EPUB content are now converted to LaTeX, ensuring mathematical content is preserved in extracted text. The extraction process now includes stricter security bounds on XML parsing depth and ZIP member sizes to prevent resource exhaustion, and handles encrypted/DRM-protected EPUBs by skipping encrypted spine items and issuing warnings. Metadata extraction supports both EPUB 2 and EPUB 3 standards, and the content pipeline better preserves structural elements like headings and lists while stripping scripts and styles.
crates/xberg/src/extractors/epub · high confidence
Engine refactoring with crawl memoization and structured extraction seams
The extraction engine internals have been reorganized into a Rust-only \Engine\ wrapper, introducing a fingerprinted crawl-engine memo (\crawl\_handle\) that reuses a single \CrawlEngine\ across multi-URL batches to share middleware, cache, and rate-limiters. The engine now supports injected \CacheBackend\ and \ProgressSink\ seams, allowing users to enable content-addressed caching for byte inputs and receive coarse progress events. Additionally, a new \ParsedDocument\ memo caches MIME detection, page counts, and lazily rendered pages to avoid redundant rasterization within a single extraction operation, while a generic structured-extraction mechanism is exposed via the \heuristics\ feature.
crates/xberg/src/engine · high confidence
Excel extraction refactored with OLE bypass and ODS metadata support
The Excel extraction module has been restructured into dedicated files for dispatch, metadata, and tests. A key behavioral change is that legacy OLE-based \.xls\ and \.xla\ files are now explicitly bypassed from ZIP validation, preventing false positives when these files contain embedded ZIP structures. Additionally, ODS files now correctly extract document metadata (such as title, author, and dates) from \meta.xml\, resolving a previous issue where ODS metadata was empty. The module also includes improved handling for XLSX revisions and comments, and more robust error handling for add-in files.
crates/xberg/src/extraction/excel · high confidence
Go package v4.0.0-rc.12 release
The Go bindings package has been updated to version 4.0.0-rc.12. This release includes the updated version number across the Go module and documentation, ensuring compatibility with the latest release candidate of the core library.
packages/go · high confidence
HWP extraction now recovers body text, metadata, and converts equations to LaTeX
The HWP extractor in \crates/xberg/src/extraction/hwp\ has been rewritten to correctly parse HWP 5.0 Compound File Binary documents. This change fixes a critical bug where body text was silently dropped due to incorrect record tag IDs and stream path mismatches, ensuring that paragraph content is now extracted. It also adds support for reading document metadata (title, author, dates) from the OLE SummaryInformation stream and converts HWP-specific equation editor scripts into LaTeX format. Additionally, table grids are now bounded to prevent excessive memory allocation, and the module includes a new error type and model structures to handle parsing failures and document structure explicitly.
crates/xberg/src/extraction/hwp · high confidence
Improved OCR reliability with device fallback warnings and repeat-decode protection
The OCR engine now prevents silent performance degradation by logging a warning when automatic device selection falls back to the CPU due to missing GPU support, ensuring users are aware of potential speed impacts. Additionally, a new repeat guard detects and truncates degenerate token loops caused by visually repetitive document regions (such as identical table rows), preventing the model from generating infinite or repetitive output during decoding.
crates/xberg-candle-ocr/src · high confidence
Improved PDF hierarchy extraction with rotated text support and scale-invariant heading detection
The PDF hierarchy module now correctly handles rotated text segments by computing upright reading frames, ensuring that reading order and spatial analysis remain accurate for non-standard orientations. Additionally, the logic for promoting headings to H1–H6 levels has been changed to rely exclusively on a scale-invariant font-size ratio relative to the body text cluster, removing the previous absolute-gap threshold that caused incorrect classification in OCR-derived documents where pixel-based measurements distorted the results.
crates/xberg/src/pdf/hierarchy · high confidence
Improved semantic element classification for headings, lists, and images
The extraction transform now more accurately classifies document structure: markdown-style headings (e.g., \\#\# Title\) are detected and preserved with their level in metadata, isolated numbered lines (like \1. Introduction\) are promoted to headings instead of narrative text, and \\[Image: ...\]\ placeholders are recognized as image elements with descriptions. List detection has been expanded to support bullet (\-\, \\*\, \•\), numbered, lettered, and indented items, with proper handling of line endings (CRLF/CR) and byte-offset tracking. Element IDs are now deterministically generated based on type, text, and page number, ensuring stable references across transformations.
crates/xberg/src/extraction/transform · high confidence
Integration of generated API examples via Astro content collections
The documentation site now automatically ingests and validates generated code snippets from the \src/snippets-generated\ directory using Astro's content collections. This change introduces a new \content.config.ts\ that defines a schema for these snippets, ensuring that fields like \id\, \language\, \target\, and \side\_effect\ are present and correctly typed, while allowing the \level\ field to be optional to accommodate fixtures that do not specify a validation depth. This setup enables the site to dynamically reflect changes in generated bindings without manual index maintenance, failing the build with clear errors if snippet metadata is malformed.
docs-site/src · high confidence
Introduce xberg CLI launcher with on-demand native binary installation
The CLI entry point has been renamed to xberg.js and now acts as a launcher that executes the native xberg binary. It includes logic to verify the binary's health (checking file size and executable permissions) and automatically triggers an on-demand download via install.js if the binary is missing or corrupt. If the binary cannot be obtained or is unsupported on the current platform, the CLI provides specific installation instructions (Homebrew or plugin marketplace) instead of crashing.
cli-proxy/npm/bin · high confidence
LLM client refactored to unify request parameters and expose new configuration options
The LLM client module has been restructured to centralize request-time parameter handling via a new \apply\_request\_time\_params\ function, ensuring that \top\_p\, \stop\, \seed\, \presence\_penalty\, and \frequency\_penalty\ are consistently applied across all LLM call sites (text completion, structured extraction, VLM OCR, and NER) rather than being silently dropped. The \LlmConfig\ now supports additional fields including \reasoning\_effort\, \extra\_body\, and \response\_body\_cap\, and custom providers are explicitly registered in the process-global registry to prevent silent failures. VLM OCR now defaults to a 300-second timeout to accommodate full-page transcription, and embeddings are requested in base64 format with validation to reject misaligned provider responses.
crates/xberg/src/llm · high confidence
Node.js bindings now support Buffer and Uint8Array inputs
The Node.js bindings for xberg-node now accept binary data (such as PDFs or images) via JavaScript Buffer, Uint8Array, or Array\<number\> types. Previously, the default NAPI v3 behavior expected a plain Array\<number\>, which caused issues when passing standard Node.js Buffers. This change introduces a custom deserialization wrapper that transparently handles these common JS binary types, ensuring that users can pass binary payloads directly without manual conversion.
crates/xberg-node/src · high confidence
PPTX extraction rewritten with security hardening and new content support
The PPTX extraction module has been completely rewritten to improve security, robustness, and content fidelity. The new implementation enforces strict security limits, including bounding individual file reads to prevent decompression bombs and validating archive entry counts. It now correctly extracts slide comments with author attribution, recovers text from math formulas, charts, SmartArt diagrams, and alternate content wrappers, and properly handles container path resolution to prevent panics on crafted files. Metadata extraction now surfaces document protection flags, and the system provides better error reporting for parsing failures instead of silently discarding data.
crates/xberg/src/extraction/pptx · high confidence
Post-processing pipeline refactored into modular components with caching and improved page boundary handling
The post-processing pipeline has been restructured into distinct modules (cache, execution, features, format, page\_markers) to improve maintainability and performance. A new processor cache reduces lock contention by snapshotting post-processors per stage and rebuilding only when the registry generation changes. Page marker injection is now normalized to verbatim RawBlock elements, ensuring consistent rendering in Markdown and Djot outputs. Page boundary recomputation logic has been hardened to handle rendering divergences by interpolating best-effort boundaries, preventing skipped pages during chunking. Output format application now correctly swaps pre-rendered content and handles custom format fallbacks to plain text.
crates/xberg/src/core/pipeline · high confidence
PyPI package now bundles native binaries for offline use
The \xberg-cli\ PyPI package has been updated to include platform-specific native binaries directly in its wheels, enabling offline execution without requiring an internet connection to download assets at runtime. A custom Hatch build hook now extracts the correct binary and its native dependencies (such as \libheif\ on Linux or \\*.dylib\ on macOS) into the wheel based on the target triple, ensuring the command works immediately after installation on supported platforms.
cli-proxy/pypi · high confidence
RTF extractor rewritten with improved encoding, formatting, and binary payload handling
The RTF extractor has been refactored into a modular structure (encoding, formatting, images, metadata, tables) to improve correctness and maintainability. Key behavioral changes include: correctly consuming \\bin payloads by source byte count rather than character count to prevent parser desynchronization and data loss; decoding hex escapes using the declared ANSI codepage or font fcharset for accurate character representation; and normalizing whitespace while preserving byte-offset mappings for precise text alignment. The extractor now also provides more robust metadata extraction (author, dates, counts) and image metadata parsing.
crates/xberg/src/extractors/rtf · high confidence
RTF parser refactored into modular components
The RTF parser implementation in \crates/xberg/src/extractors/rtf/parser\ has been restructured from a single large file into four distinct modules: \control\_word\, \formatting\_extract\, \text\_extract\, and \mod\. This change consolidates the 27-argument \handle\_control\_word\ function into a structured \ControlWordCtx\ context object and splits the control-word dispatch logic into three category-specific handlers (scope/state, text emission, and table/formatting) to maintain file size limits. The \formatting\_extract\ module now handles a dedicated pass for extracting formatting metadata (bold, italic, color, hyperlinks), while \text\_extract\ manages the primary text extraction and table/image processing, improving code maintainability without altering the external API.
crates/xberg/src/extractors/rtf/parser · high confidence
Refactored MathML and OOXML embedded object extraction for improved reliability and security
The MathML conversion logic in \crates/xberg/src/extraction/mathml\ has been split into separate collection and rendering phases, introducing dedicated handling for complex structures like \mfenced\ (suppressing default comma separators for infix operators), \munder\/\mover\ (mapping accents to LaTeX macros like \\\hat\ or \\\underline\), and \mtable\ (rendering as matrix environments). This refactoring also ensures private use characters and prefixed OpenOffice MathML DOCTypes are handled correctly. In \crates/xberg/src/extraction/ooxml\_embedded\, the extraction of embedded objects from OOXML archives (DOCX/PPTX) now includes robust security measures against memory exhaustion attacks by clamping untrusted ZIP central-directory size declarations to configured limits, and enforces file count and depth caps with appropriate warnings.
_crates/xberg/src/extraction/mathml, crates/xberg/src/extraction/ooxml\embedded · high confidence
Refactored extraction core into modular byte and file handlers
The internal extraction logic in \crates/xberg/src/core/extractor\ has been reorganized into distinct modules (\bytes.rs\, \file.rs\, \helpers.rs\) to separate in-memory byte array processing from filesystem-based extraction. This change introduces dedicated helpers for resolving MIME types (including sniffing for unknown extensions) and retrieving extractors from the registry, while ensuring that timeout enforcement and cancellation tokens are correctly applied to the extraction futures. The refactoring also adds comprehensive unit tests for these core extraction paths, covering scenarios like empty files, long paths, and invalid MIME types.
crates/xberg/src/core/extractor · high confidence
Refactored scan detection tests and corrected OCR scoring logic
The scan detection test suite has been reorganized into a dedicated file, with stale and duplicate test cases removed. The OCR scan-detection logic was updated to align with image-based OCR processing: it now correctly counts mapped artifact glyphs while ignoring fallback glyphs, and gates the scan-density policy by its consumers. Additionally, the test suite was expanded to verify that page mapping and provenance detection behave correctly in parallel execution environments, ensuring results match sequential runs and are dispatched across thread pools.
_crates/xberg/src/pdf/scan\detect · high confidence
Regenerated C header with expanded feature flags and new type definitions
The \xberg.h\ C header has been regenerated to reflect the current state of the \xberg\ crate. This update introduces a comprehensive set of \XBERG\FEATURE\\*\ macros (such as \XBERG\_FEATURE\_PADDLE\_OCR\, \XBERG\_FEATURE\_LITER\_LLM\, and \XBERG\_FEATURE\_MCP\) that allow consumers to detect available capabilities at compile time. Additionally, the header now includes forward declarations for numerous new opaque types, including \XBERGAccelerationConfig\ for hardware acceleration, \XBERGBedrockConfig\ for AWS Bedrock integration, and \XBERGArchiveEntry\ for archive handling, ensuring the C interface remains synchronized with the underlying Rust implementation.
crates/xberg-ffi/include · high confidence
Restored xberg-pdfium-render crate with explicit runtime library loading behavior
The xberg-pdfium-render crate has been restored as a workspace member, providing a Rust wrapper around the Pdfium library. For users, this introduces a specific runtime behavior: if neither PDFIUM\_STATIC\_LIB\_PATH nor PDFIUM\_DYNAMIC\_LIB\_PATH is set during the build, the crate will compile successfully but will not link against a PDFium binary. Instead, it relies on dynamic loading at runtime, which means the system must have libpdfium available in the working directory or system library search path; otherwise, PDF operations will fail with a load-library error when first attempted. The crate is now licensed under MIT.
crates/xberg-pdfium-render · high confidence
Shared ONNX model-loading helpers with SHA-256 verification and companion file resolution
A new shared ONNX helper module consolidates model downloading, caching, and session building for embeddings and reranking capabilities. It enforces SHA-256 integrity checks on downloaded files using a pinned manifest to prevent tampering, resolves companion files (like tokenizers) from both the model's subdirectory and the repository root to support various Hugging Face repo structures, and respects the user's download progress settings across all ONNX-backed features.
crates/xberg/src/onnx · high confidence
Standardized global cache directory and cooperative cancellation support
Xberg now uses platform-appropriate global cache directories (e.g., macOS \~/Library/Caches/xberg, Linux $XDG\_CACHE\_HOME/xberg) instead of per-CWD .xberg folders, with an XBERG\_CACHE\_DIR environment variable to override the location. Additionally, a cooperative CancellationToken is available to request extraction stop at the next checkpoint, improving responsiveness for long-running or interactive workloads.
crates/xberg/src · high confidence
Translation now covers all document text fields, not just main content
The translation feature has been expanded to translate every text-bearing field in an extracted document, including tables, pages, metadata, semantic elements, and the structured document tree. Previously, only the main content, formatted content, and chunk content were translated, leaving other fields in the source language. This change ensures that all user-visible text is translated into the target language while preserving call-volume efficiency through batched LLM requests.
crates/xberg/src/text/translation · high confidence
Unified extraction configuration with per-file overrides and environment variable support
The extraction configuration system has been consolidated into a single \ExtractionConfig\ struct that aggregates all processing options (OCR, chunking, content filtering, image handling, etc.) and introduces \FileExtractionConfig\ for per-file overrides within batch processing. Configuration can now be loaded from TOML, YAML, or JSON files, with automatic discovery in the current directory and the XDG platform config directory. Environment variables (e.g., \XBERG\_OCR\_LANGUAGE\, \XBERG\_CHUNKING\_MAX\_CHARS\) can override settings at runtime, and the system now validates loaded configs and rejects unknown nested fields to prevent misconfiguration.
crates/xberg/src/core/config/extraction · high confidence
Unified extraction input envelope and automatic telemetry for document extractors
The document extractor plugin system now uses a unified \ExtractInput\ envelope (supporting both bytes and paths) and returns a standardized \ExtractedDocument\ result, simplifying how custom extractors handle input. Additionally, when the \otel\ feature is enabled, all extraction operations are automatically wrapped in an \InstrumentedExtractor\ that records tracing spans and metrics (such as duration, input/output bytes, and status) without requiring individual extractor annotations.
crates/xberg/src/plugins/extractor · high confidence
Unified layout detection models with engine-neutral inference seam
The layout detection models (PP-DocLayout-V3, RT-DETR, SLANeXT, TATR, YOLO variants, and table classifier) now run through a common \LayoutModel\ trait backed by the \crate::inference\ seam, allowing them to execute on either the ORT or Tract inference engine where supported. This change standardizes how models are loaded, preprocessed, and invoked, ensuring consistent behavior across different backend configurations while maintaining engine-specific optimizations and fallbacks.
crates/xberg/src/layout/models · high confidence
Xberg 1.3.0 release with Tract backend integration and cache build-id stability
The Xberg Rust library has been updated to version 1.3.0, introducing a new Tract backend for inference capabilities including OCR (PaddleOCR, Sceptre) and layout detection (RT-DETR) for targets without ONNX Runtime. The release also includes a critical fix to the caching mechanism: the cache key now incorporates a unique build identifier derived from the git commit SHA and working tree state (or an explicit \XBERG\_BUILD\_ID\ environment variable), ensuring that separately built binaries with the same crate version do not silently serve each other's cached results. Additionally, the \TesseractConfig.psm\ option is now optional, and the library ships with a pure-Rust PDF backend (\xberg-native-pdf\) requiring no system libraries.
crates/xberg · high confidence
xberg-tesseract: rebranded crate with hardened builds and WASM support
The \xberg-tesseract\ crate has been rebranded (relicensing under MIT) and upgraded to vendored Tesseract 5.5.3 and Leptonica 1.87.0. The build process now includes security hardening by verifying source archive identities and pinning native inputs, alongside improved source cache recovery. It adds support for dynamic linking to system libraries and introduces WebAssembly (WASM) compilation via specific patches that disable incompatible features like OpenCL and fix stack allocation limits. The crate also normalizes Windows CMake paths and resolves symlinked build roots.
crates/xberg-tesseract · high confidence
Fixes
C\# NuGet package updated to version 4.0.0-rc.12
The C\# binding package (Kreuzberg) has been updated from version 4.0.0-rc.11 to 4.0.0-rc.12. This release includes a fix for the MSBuild target to prevent overwriting CI-downloaded native assets during the build process, ensuring that cross-platform runtime files are preserved correctly across Windows, macOS, and Linux environments.
packages/csharp · high confidence
Centralized configuration validation for server and OCR settings
A new \config\_validation\ module has been added to centralize and enforce validation rules for configuration values, eliminating duplication across language bindings. This change introduces server-boundary validators for listen host, port, CORS origins, and upload size limits, ensuring that network-facing fields are checked before the server starts. It also adds specific validators for OCR configuration, including support for the new Sceptre backend, acceptance of all registered Candle backend names, and validation of language codes for Tesseract and other backends. Users will now receive clear, actionable error messages if their configuration contains invalid hosts, ports, CORS entries, or unsupported OCR backends, rather than encountering silent failures or cryptic errors later in the process.
_crates/xberg/src/core/config\validation · high confidence
Fix PDF OCR DPI handling and image normalization
PDF pages are now rendered at a DPI that respects the configured target, minimum, and maximum DPI settings, as well as dimension and memory constraints, rather than ignoring these settings. Scanned pages are rendered at their native density (capped at 300 DPI) to preserve detail, ensuring consistent OCR results between PDF and image inputs. Image preprocessing now correctly applies the configured DPI and dimension limits, and uses optimized resizing that avoids unnecessary buffer copies.
crates/xberg/src/image · high confidence
Fix WASI imports to enable browser and bundler compatibility
Added a post-build script that patches the generated WebAssembly glue code to replace unresolvable 'env' and 'wasi\_snapshot\_preview1' imports with inline stubs. This change fixes runtime errors in browsers and bundlers (such as Deno) where module resolution for these system-level imports fails, ensuring the OCR functionality works correctly across all supported environments including Node.js, web browsers, and bundlers.
crates/xberg-wasm/scripts · high confidence
Fix image decode budget validation and add security audit tests
The image decoder now correctly validates memory budgets against byte counts rather than pixel counts, preventing false rejections for high-pixel-count images with small decoded sizes. Additionally, new tests ensure that the source audit correctly identifies various decoder families and aliases, strengthening the security posture of the image extraction process.
_crates/xberg/src/extraction/image\decode · high confidence
Fixes OCR cache returning incomplete results
The OCR cache now correctly preserves the structured hOCR document (including paragraph structure, bounding boxes, and confidence scores) alongside the extracted text. Previously, the cache silently dropped this structured data during serialization, causing cache hits to return strictly less information than a fresh OCR run (a cache miss). This fix ensures that cached results are lossless and identical to newly computed ones.
crates/xberg/src/ocr · high confidence
Fixes WASM build failures caused by duplicate libc symbols
The WebAssembly package now builds reliably in environments where the \RUSTFLAGS\ environment variable is set (such as CI pipelines that enforce \-D warnings\). Previously, setting \RUSTFLAGS\ caused Cargo to ignore the \--allow-multiple-definition\ linker flag defined in \.cargo/config.toml\, leading to link errors due to duplicate symbols from vendored libc stubs in Tesseract and tree-sitter. A new build script (\build.rs\) now explicitly emits this linker flag for the WASM target, ensuring the build succeeds regardless of external environment variables.
crates/xberg-wasm · high confidence
Fixes annotation offsets after text trimming
The extraction engine now correctly adjusts text annotation byte offsets when leading or trailing whitespace is trimmed from extracted content. Previously, annotations that extended into the removed whitespace were silently dropped, causing a loss of formatting information (such as bold or italic spans). The new \adjust\_annotations\_for\_trim\ utility shifts offsets to account for the trimmed characters and clamps them to the new text length, ensuring that annotations covering real words are preserved even if they originally touched the whitespace.
crates/xberg/src/extractors · high confidence
Fixes to legacy .ppt extraction: embedded objects, slide ordering, and notes attribution
The legacy PowerPoint (.ppt) extractor now correctly handles several structural nuances of the binary format. It extracts embedded OLE objects (such as inserted Excel tables) and associates them with the specific slide that displays them. Slide ordering and counting are resolved using the document's live persist chain rather than raw stream order, preventing stale revisions from appearing as extra slides. Additionally, slide titles are now read from the document outline collection, and speaker notes are accurately attributed to their slides using the slide ID reference, avoiding misattribution when intermediate slides lack notes.
crates/xberg/src/extraction/ppt · high confidence
Improved PDF table detection and layout validation
The PDF structure pipeline now includes a geometric table fallback that recovers borderless, numeric, and text-heavy key-value tables when the ML layout detector misses them, using column alignment and spacing heuristics. It also adds pixel-level validation to suppress false-positive table and picture regions that contain no text, prevents bare URLs from being promoted to headings, and improves table recognition by splitting side-by-side tables and tightening layout bounding boxes.
crates/xberg/src/pdf/structure/regions · high confidence
Improved table reconstruction accuracy by filtering OCR artifacts
The table reconstruction logic now better handles noisy OCR output from shaded or complex tables. It filters out 'shading marks' (such as low-confidence dashes or equals signs that appear between values due to shading artifacts) and removes fused underscore characters from word boundaries. Additionally, it ensures that words from a single Tesseract text line are aligned to a consistent vertical band, preventing isolated words from incorrectly starting new rows or causing column misalignment.
crates/xberg/src/ocr/table · high confidence
OCR pipeline refactored into modular submodules with improved text quality and numeric repair
The PDF OCR extraction logic has been reorganized into distinct submodules (\scoring\, \plausibility\, \rendering\, \document\, \pipeline\) to improve maintainability and separation of concerns. This change introduces a language/dictionary plausibility check to better detect and route incorrectly mapped text layers (such as ROT-shifted or mojibake text) to OCR, rather than relying solely on character-shape heuristics. It also adds an opt-in numeric token repair feature that fixes common OCR defects like missing thousands separators, misread decimal points, and split numbers. Additionally, the pipeline now handles per-page OCR confidence reporting, preserves structured elements like tables and bounding boxes through fallback routes, and ensures that OCR output is correctly scaled and assembled into the final document structure.
crates/xberg/src/extractors/pdf/ocr · high confidence
Test coverage
1 commit adding/updating tests in cli-proxy/pypi/tests; Added Bats test harness and shared path helpers; Added Kotlin Android end-to-end tests for batch extraction and code parsing; Added OCR test fixtures for hOCR output validation; Added comprehensive test suite for xberg extraction engine; Added concurrency, layout microbatch, and text quality benchmarks; Added diagnostic and tract-specific tests for PaddleOCR detection; Added integration and build-system tests for xberg-tesseract; Added integration tests for Candle-based GLiNER inference and WASM compatibility; Added integration tests for DeepSeek-OCR, GLM-OCR, PaddleOCR-VL, and TrOCR engines; Added regression tests for FFI byte-buffer length handling and vtable callbacks; Added shared test helpers for integration testing; Added smoke tests for HEIC/HEIF/AVIF decoding; Added test coverage for CSV, structured, and WordPerfect extractors; Added test coverage for core configuration, image encoding, and security validation modules; Added test environment shims for WASI imports; Added test fixture for bullet list word extraction; Added tests for extraction engine behavior and error handling; Added tests for format serialization and Tesseract config defaults; Added tests for the OCR enablement bridge; Added tests for the xberg-libwpd WordPerfect extraction crate; C\# end-to-end test suite scaffolding; Expanded CLI integration tests for extraction, configuration, and server commands; Integration test suite for xberg-native-pdf; Isolated test registry for pipeline initialization; Java E2E test suite for batch and contract extraction; Removed test fixtures and sample data files; Zig e2e test suite scaffolding.
Dependencies
Dependency updates across 105 manifests
This release updates dependencies across 105 manifests, including 2156 commits. Specific changes visible in the diff include updating mypy from \>=1.19 to \>=1.20.1, pytest-cov from \>=7.0.0 to \>=7.1.0, llama-index-core from \<0.15,\>=0.13 to \>=0.14.22,\<0.15, and html-to-markdown-rs to v3.1.0. Other updates include bumping glob from 10.5.0 to 13.0.6, pypdf, and various other packages across the workspace.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Score
- CAI 61 → 48 (-13.0)
- Rubric changed (rubric-2026.09.8 → rubric-2026.09.17) — scores are not directly comparable.
Lenses
- Code Health 64 → 66 (+1.7)
- Architecture 94 → 95 (+0.5)
- Maturity 76 → 73 (-3.0)
- Readiness 59 → 54 (-4.7)
- Security 57 → 39 (-18.4)
- Performance 48 (new)
Resolved (709)
- BenchmarkRunner::run (cognitive 113) (tools/benchmark-harness/src/runner.rs)
- BenchmarkRunner::run (cyclomatic 39) (tools/benchmark-harness/src/runner.rs)
- BenchmarkRunner::run_batch_iterations_static (cognitive 33) (tools/benchmark-harness/src/runner.rs)
- BenchmarkRunner::run_batch_iterations_static (cyclomatic 22) (tools/benchmark-harness/src/runner.rs)
- BenchmarkRunner::run_iterations_static (cognitive 61) (tools/benchmark-harness/src/runner.rs)
- BenchmarkRunner::run_iterations_static (cyclomatic 35) (tools/benchmark-harness/src/runner.rs)
- CaptioningProcessor::process (cognitive 34) (crates/xberg/src/plugins/processor/builtin/captioning.rs)
- CaptioningProcessor::process (cyclomatic 19) (crates/xberg/src/plugins/processor/builtin/captioning.rs)
- Change coupling: aggregate.rs ↔ output.rs (tools/benchmark-harness/src/aggregate.rs)
- Change coupling: docbook.rs ↔ mod.rs (crates/xberg/src/extractors/docbook.rs)
- Change coupling: docbook.rs ↔ mod.rs (crates/xberg/src/extractors/docbook.rs)
- Change coupling: paddleocr_vl_backend.rs ↔ trocr_backend.rs (crates/xberg/src/candle_ocr/paddleocr_vl_backend.rs)
- Change-coupling hub: subprocess.rs → native.rs, consolidate.rs, output.rs (tools/benchmark-harness/src/adapters/subprocess.rs)
- CitationExtractor::extract_content (cognitive 78) (crates/xberg/src/extractors/citation.rs)
- CitationExtractor::extract_content (cyclomatic 33) (crates/xberg/src/extractors/citation.rs)
- ClassTooLong: BenchmarkRunner (tools/benchmark-harness/src/runner.rs)
- ClassTooLong: ExtractionOverrides (crates/xberg-cli/src/commands/overrides.rs)
- ClassTooLong: HtmlWalker (crates/xberg/src/extraction/html/structure.rs)
- ClassTooLong: PdfFont (crates/xberg-pdfium-render/src/pdf/font.rs)
- ClassTooLong: ProfileReport (tools/benchmark-harness/src/profile_report.rs)
- …and 689 more
New (393)
- Change coupling: docbook.rs ↔ mod.rs (crates/xberg/src/extractors/docbook.rs)
- Change coupling: docbook.rs ↔ mod.rs (crates/xberg/src/extractors/docbook.rs)
- Change-coupling hub: subprocess.rs → native.rs, mod.rs, output.rs (tools/benchmark-harness/src/adapters/subprocess.rs)
- Documentation: no project overview (integrations/java/spring-ai/README.md)
- Duplicate intent between API response type and internal cache state type. CacheStatsResponse and CacheStats have identical properties (total_files, total_size_mb, available_space_mb, oldest_file_age_days, newest_file_age_days).
- Duplicated block (10 lines × 2) (crates/xberg-native-pdf/src/annotations.rs)
- Duplicated block (10 lines × 2) (crates/xberg-native-pdf/src/extractors/images.rs)
- Duplicated block (10 lines × 2) (crates/xberg-native-pdf/src/pipeline/reading_order/xycut.rs)
- Duplicated block (10 lines × 2) (crates/xberg-tesseract/src/leptonica.rs)
- Duplicated block (10 lines × 2) (crates/xberg/src/extraction/mathml/render.rs)
- Duplicated block (10 lines × 2) (crates/xberg/src/extraction/ooxml_embedded/mod.rs)
- Duplicated block (10 lines × 2) (crates/xberg/src/extraction/transform/content.rs)
- Duplicated block (10 lines × 2) (crates/xberg/src/extractors/epub/mod.rs)
- Duplicated block (10 lines × 2) (tools/benchmark-harness/src/aggregate/mod.rs)
- Duplicated block (10 lines × 2) (tools/benchmark-harness/src/comparison/execution.rs)
- Duplicated block (10–11 lines × 2) (crates/xberg-native-pdf/src/decoders/ccitt.rs)
- Duplicated block (10–11 lines × 2) (tools/benchmark-harness/src/sizes/package_measurements.rs)
- Duplicated block (11 lines × 2) (crates/xberg/src/extraction/excel/package_metadata.rs)
- Duplicated block (11 lines × 2) (crates/xberg/src/extractors/fictionbook.rs)
- Duplicated block (11 lines × 2) (crates/xberg/src/layout/models/pp_doclayout_v3.rs)
- …and 373 more
Changes since last survey
- 300 commits — 133 feature/other, 167 fixes
By area
- crates/xberg — 137 commits
- crates/xberg-native-pdf — 63 commits
- (root) — 47 commits
- (repo) — 20 commits
- docs-site/src — 19 commits
- .github/workflows — 2 commits
- plugin/.hermes — 2 commits
- .ai-rulez/skills — 1 commit
- .github/actions — 1 commit
- crates/xberg-ffi — 1 commit
- crates/xberg-tesseract — 1 commit
- e2e/zig — 1 commit
- fixtures/pdf — 1 commit
- integrations/python — 1 commit
- packages/elixir — 1 commit
- packages/kotlin-android — 1 commit
- scripts/ci — 1 commit
Notable commits
- fix: Merge origin/main into fix/zebra-table-cells
- fix: Revert "fix(pdf): split table regions from adjacent prose columns"
- fix: fix(alef): restore the poly.toml merge baselines two refactors destroyed
- fix: fix(api): drop the null union from omitted optional references
- fix: fix(build): gate config-dependent items for narrow feature legs
- fix: fix(build): keep the internal image-classify parameter bag out of the bindings
- fix: fix(build): package the JPEG 2000 fixtures the crate's own tests read
- fix: fix(ci): compare the Unreleased section and release dates in changelog sync
- fix: fix(ci): waive the int32-as-JAVA_LONG return of supports_language_for
- fix: fix(deps): re-pin skrifa to 0.46 so one read-fonts is in the graph
- fix: fix(deps): refresh the path-sourced uv lockfiles for 1.3.0
- fix: fix(doc): extract Word 6 and 7 binary documents
- fix: fix(docker): pass XBERG_BUILD_ID into image builds
- fix: fix(docker): require XBERG_BUILD_ID in every Rust-building image
- fix: fix(docs): restore OCR changelog entries
- fix: fix(excel): bypass ZIP validation for legacy OLE files
- fix: fix(ffi): add the config-aware vtable fields to the bytes-len test
- fix: fix(layout): restore the wasm32 gate on the custom-model dispatch (#1862)
- fix: fix(legacy): integrate legacy Office fixes (#1940)
- fix: fix(native-pdf): bound RunLength/LZW decode output mid-decode
- …and 280 more
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
xberg-io/xberg was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 29 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 1409e7be4bf23042f6d18d7405e2a7f560215109 — the exact code this score is about.
- Scored under rubric-2026.09.17 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-fbec9b1e08c2.