firecrawl/pdf-inspector
70.3
Strong · 28 September 2026
129.4k
lines of production code
Rust
primary language
2
measurements over time
What this system is
pdf-inspector is a high-performance, Rust-based library for PDF classification, text extraction, and Markdown conversion, designed for use in server-side and browser environments. It provides native bindings for Python, Node.js, and WebAssembly, enabling developers to process PDFs locally with features like position-aware text extraction, table detection, and optional local OCR. The system handles complex layout challenges, including multi-column documents, right-to-left text, and hidden text layers, while offering CLI tools for command-line usage and structured JSON output.
Features
Add pdf-inspector CLI for PDF text extraction and classification
A new command-line interface, pdf-inspector, is now available to extract text from PDF files into Markdown or classify document types. Users can run the tool on local files or via standard input, with options to output results as JSON, specify particular pages, or write output to a file. The CLI supports two main modes: default text extraction and document detection, providing a convenient way to interact with the library's core capabilities directly from the terminal.
napi/bin · high confidence
Detects hidden text layers masked by covering images
The detector now identifies pages where a text layer is present but invisible because it is obscured by a covering image. This is achieved by executing a byte-level scan of the page's content streams and Form XObjects to track text render modes and graphics state, specifically flagging cases where text operators use render mode 3 (invisible) or mode 7 (clip-only) and the paint that would reveal them is blocked by an image covering at least 50% of the page area.
src/detector · high confidence
Initial release of pdf-inspector 1.25.2 with Python/Node bindings and CLI tools
This entry marks the initial public release of the pdf-inspector library (version 1.25.2), providing a Rust-based engine for PDF classification and text extraction. Users can now install native bindings for Python and Node.js, as well as a WebAssembly build for browsers, alongside CLI binaries (\pdf2md\ and \detect-pdf\) for command-line usage. The library supports smart classification of PDF types (TextBased, Scanned, ImageBased, Mixed), position-aware text extraction with font metadata, and conversion to structured Markdown. It includes features like multi-column layout detection, table detection, and selective OCR routing, with comprehensive documentation and type stubs provided for the new bindings.
(repo-wide) · high confidence
Introduce NAPI bindings for Node.js with comprehensive PDF extraction API
This change adds the initial NAPI binding layer (\napi/src/lib.rs\), exposing the library's PDF processing capabilities to Node.js applications. The bindings provide string enums for classifying PDF types (\PdfType\), item types (\ItemType\), OCR modes (\OcrMode\), and bold detection sources (\BoldSource\). Users can now extract text with precise positioning (including rotation and axis-aligned boxes relative to the visible page box), retrieve document metadata (title, author, dates), and access detailed OCR diagnostics such as per-page reasons and ToUnicode CMap gap analysis. The API also supports region-based table extraction and vector grid detection, with async variants designed to keep the Node event loop free.
napi/src · high confidence
Introduce Node.js binding for PDF inspection and extraction
Adds a new N-API package (@firecrawl/pdf-inspector) that exposes native Rust-based PDF classification, text extraction with positions, and OCR capabilities to Node.js and Bun environments. The binding provides synchronous and asynchronous APIs (e.g., classifyPdf, extractTextWithPositions, processPdfWithOcr) that run on the libuv thread pool to keep the event loop free, supporting features like region-based extraction, font weight analysis, and rotation handling across Linux, macOS, and Windows platforms.
napi · high confidence
Introduce browser WebAssembly bindings for PDF inspection
This change adds the \@firecrawl/pdf-inspector-wasm\ package, enabling PDF classification and structured Markdown extraction directly in the browser. By embedding the Rust core as a WebAssembly binary, users can process \Uint8Array\ PDF data locally without uploading it to a server. The bindings expose functions such as \processPdf\, \detectPdf\, \classifyPdf\, and \extractText\, supporting options for page selection, password protection, and Markdown profiles (fidelity vs. compact). The package also includes embedded Adobe CMaps to ensure correct CJK font decoding without filesystem dependencies and provides document metadata extraction.
wasm · high confidence
Introduces optional local OCR pipeline with adaptive fusion and model management
The \src/vision\ module adds a new optional vision layer that enables local Optical Character Recognition (OCR) on PDF pages. This includes an OAR-based OCR engine (PP-OCRv6 Small), a checksum-verified model cache with HTTPS download support, and a routing system that can selectively apply OCR. The pipeline fuses OCR results with native PDF text extraction, allowing for adaptive merging where native text is retained when high-quality and OCR is used to supplement or replace it. Users can configure OCR modes (Off, Auto, Force), confidence thresholds, and offline model directories via the new \OcrOptions\ and \OcrPdfOptions\ APIs.
src/vision · high confidence
New Python example and RTL/symbolic font test fixtures
The \examples/basic\_usage.py\ script now demonstrates the Python bindings for the pdf-inspector library, covering full processing, detection, byte-based processing, text extraction, and region-based extraction. Additionally, new Rust examples (\rtl\_fixtures.rs\ and \symbolic\_font\_fixtures.rs\) generate synthetic test fixtures to validate right-to-left text handling and symbolic font decoding.
examples · high confidence
New benchmarking, evidence-probing, and version-unification scripts with tests
The scripts directory now includes three new tooling scripts and their corresponding unit tests. \bench\_opendataloader.py\ provides a benchmarking workflow for comparing OpenDataLoader builds, reporting metric and per-document deltas, and enforcing quality gates (e.g., minimum overall delta, maximum missing predictions, and reference lead requirements). \probe\_backend\_evidence.py\ is a diagnostic tool that compares pdf-inspector extraction evidence against MuPDF's structured text backend, flagging pages with significant text gains, alignment anchor differences, or image block discrepancies. \version.py\ enforces a single release version across all project artifacts (Rust crates, Python package, NAPI/WASM packages, Node optional dependencies, Bun locks, and the site WASM URL), providing functions to check for consistency and atomically update all locations. Comprehensive unit tests cover the comparison logic, gate evaluations, PDF fixture generation utilities, and version synchronization behavior.
scripts · high confidence
New landing page and branding assets for pdf-inspector
The site now features a new landing page (index.html) for the pdf-inspector project, replacing the previous content. This page includes a redesigned UI with a sticky navigation bar, a hero section highlighting the tool as an open-source PDF parser, and integration of new brand assets, specifically the firecrawl-mark.svg and firecrawl-wordmark.svg files.
site · high confidence
Architecture
New modular table detection and formatting system
The table extraction logic has been reorganized into a dedicated \src/tables\ module with specialized components for different detection strategies and output formatting. This change introduces separate modules for heuristic detection (\detect\_heuristic\), line-based detection from PDF path operators (\detect\_lines\), rectangle-based clustering (\detect\_rects\), and structure-tree extraction (\detect\_struct\). It also adds a new \cell\_text\ module to handle subscript/superscript-aware text joining and a \financial\ module for splitting consolidated financial values. The \format\ module now handles markdown conversion with specific logic for table-of-contents lists and footnote preservation. This modularization supports the various fixes and features seen in the commit history, such as improved borderless table detection, chart masking, and TOC handling, by isolating the logic for each capability.
src/tables · high confidence
Behavioural changes
Added Adobe CMap resources and license
The project now includes binary CMap files (CNS2-V, ETenms-B5-H, GB-H) and the associated Adobe Systems Incorporated license. These additions support improved CMap handling, enabling better text rendering for specific character sets through binary CMap support, inline ToUnicode, and fallback decoding.
external · high confidence
CLI binaries now support JSON output and configurable processing modes
The CLI tools in src/bin (pdf2md, detect\_pdf, and the new dump\_ops) have been updated to support structured JSON output via the --json flag, with proper escaping of special characters in strings. The pdf2md and detect\_pdf binaries now accept flags before the PDF path, support the --analyze flag for layout complexity detection, and integrate with the pdf\_inspector library to expose detailed extraction data including font properties, text colors, and OCR routing reasons. A new dump\_ops binary has been added for debugging PDF content streams.
src/bin · high confidence
Major refactor of Markdown conversion with improved heading, table, and layout detection
The \src/markdown\ module has been restructured into distinct components (\analysis\, \classify\, \convert\, \furniture\, \heading\, \postprocess\, \preprocess\) to support a more robust and deterministic Markdown extraction pipeline. Key behavioral improvements include: rarity-based heading detection that identifies section titles by font size frequency rather than just absolute size; improved table detection that now handles borderless layouts, chart-bar clusters, and side-by-side tables via X-gap pre-splitting; and better handling of document furniture, including stripping running headers/footers on short documents using positional evidence. The conversion loop now correctly interleaves tables and images with prose, respects chart regions to prevent text overlap, and merges wrapped heading lines and hyphenated words more accurately. Additionally, the system now suppresses overused struct-tree heading tags to reduce false positives and improves list item classification to distinguish between bullet points and numbered headings.
src/markdown · high confidence
New internal modules for PDF text extraction and repair
The library adds several new internal source files to improve text extraction accuracy and robustness. bidi.rs and bidi\_mirroring.rs implement the Unicode Bidirectional Algorithm to correctly reverse visual-order right-to-left text (Arabic, Hebrew) and un-mirror paired brackets. Glyph mapping is enhanced with adobe\_korea1.rs for Korean CID-to-Unicode lookup, glyph\_names.rs for Adobe Glyph List support, and mac\_glyph\_order.rs as a fallback for subsetted TrueType fonts. Form XObject rendering issues are addressed by form\_bbox\_repair.rs and overlong\_numerals.rs, which widen zero-area or excessively large BBoxes to prevent blank pages. Additionally, process\_mode.rs introduces a configurable pipeline mode (DetectOnly, Analyze, Full), and python.rs exposes the library's capabilities via PyO3 bindings.
src · high confidence
Fixes
Improved text extraction accuracy for standard fonts, layout, and security
The extractor now supplies built-in glyph metrics for the 14 standard PDF fonts (e.g., Times, Helvetica, Courier) when width data is missing, ensuring correct spacing and layout detection. It handles right-to-left text by reversing visual order using glyph geometry and honors PDF text horizontal scaling. Layout detection is hardened to preserve contextual digit runs, unfuse independent column runs sharing a baseline, and correctly clip text painted wholly outside its rectangular clip. Security is improved by bounding content-stream operator counts and decompressed bytes to prevent memory exhaustion, and by preventing stack-overflow denial-of-service from AcroForm self-cycles.
src/extractor · high confidence
Test coverage
Added comprehensive test coverage for PDF extraction, OCR, and Python bindings; Added script to generate font metadata test fixtures.
Dependencies
pdf-inspector v1.25.2 release with unified package structure and new binaries
This release updates the core Rust library, N-API bindings, and WASM bindings to version 1.25.2. The N-API package now includes a CLI binary (\pdf-inspector\) and distributes platform-specific binaries as optional dependencies for Linux (x64/arm64), macOS (arm64), and Windows (x64). The core library has been renamed from \pdf-to-markdown\ to \pdf-inspector\, updated to require Rust 1.88, and now includes a new \dump\_ops\ binary alongside existing tools. The WASM bindings have been updated to use \wasm-bindgen\ 0.2 and \serde-wasm-bindgen\ 0.6, with explicit configuration to disable \wasm-opt\ to ensure compatibility with current tooling.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Score
- CAI 73 → 70 (-2.8)
- Rubric changed (rubric-2026.09.8 → rubric-2026.09.16) — scores are not directly comparable.
Lenses
- Code Health 74 → 75 (+0.4)
- Architecture 100 → 90 (-9.6)
- Maturity 69 → 68 (-0.7)
- Readiness 68 → 65 (-3.5)
- Security 91 → 94 (+2.5)
Resolved (57)
- Documentation: no installation or build instructions (README.md)
- Duplicated block (10 lines × 3) (src/tables/financial.rs)
- Duplicated block (11 lines × 2) (src/lib.rs)
- Duplicated block (11 lines × 2) (src/tounicode.rs)
- Duplicated block (12–17 lines × 2) (src/extractor/content_stream.rs)
- Duplicated block (14 lines × 2) (src/tounicode.rs)
- Duplicated block (15 lines × 2) (src/detector.rs)
- Duplicated block (15 lines × 2) (src/extractor/xobjects.rs)
- Duplicated block (15 lines × 3) (src/extractor/content_stream.rs)
- Duplicated block (17 lines × 2) (src/lib.rs)
- Duplicated block (19–21 lines × 2) (src/extractor/content_stream.rs)
- Duplicated block (28 lines × 2) (src/extractor/content_stream.rs)
- Duplicated block (35–36 lines × 2) (src/extractor/content_stream.rs)
- Duplicated block (49–50 lines × 2) (src/extractor/content_stream.rs)
- Duplicated block (6 lines × 2) (src/bin/detect_pdf.rs)
- Duplicated block (6 lines × 2) (src/bin/detect_pdf.rs)
- Duplicated block (6 lines × 2) (src/extractor/mod.rs)
- Duplicated block (7 lines × 2) (src/tounicode.rs)
- Duplicated block (7–8 lines × 2) (src/detector.rs)
- Duplicated block (8 lines × 2) (src/tables/detect_rects.rs)
- …and 37 more
New (118)
- ClassTooLong: ContentScanState (src/detector/content_scan.rs)
- ClassTooLong: ToUnicodeCMap (src/tounicode.rs)
- Dependency hygiene PARTLY measured — Cargo dependencies read, no committed lock to grade for currency
- Duplicate Data Structures for CMap Statistics
- Duplicate Result Types for Identical Data Structure
- Duplicated block (10 lines × 2) (napi/src/lib.rs)
- Duplicated block (10 lines × 2) (src/overlong_numerals.rs)
- Duplicated block (10 lines × 3) (src/extractor/content_stream.rs)
- Duplicated block (10 lines × 3) (src/tables/financial.rs)
- Duplicated block (10–14 lines × 2) (src/extractor/content_stream.rs)
- Duplicated block (11 lines × 2) (src/detector/content_resources.rs)
- Duplicated block (11 lines × 2) (src/extractor/content_stream.rs)
- Duplicated block (11 lines × 2) (src/extractor/xobjects.rs)
- Duplicated block (12 lines × 2) (src/overlong_numerals.rs)
- Duplicated block (12 lines × 2) (src/tounicode.rs)
- Duplicated block (14 lines × 2) (src/bin/detect_pdf.rs)
- Duplicated block (14–15 lines × 2) (src/extractor/content_stream.rs)
- Duplicated block (14–15 lines × 2) (src/extractor/xobjects.rs)
- Duplicated block (15 lines × 2) (src/tounicode.rs)
- Duplicated block (15 lines × 3) (src/extractor/content_stream.rs)
- …and 98 more
Changes since last survey
- 43 commits — 15 feature/other, 28 fixes
By area
- src/extractor — 19 commits
- (root) — 17 commits
- tests/fixtures — 2 commits
- .github/workflows — 1 commit
- src/bin — 1 commit
- src/detector — 1 commit
- src/lib.rs — 1 commit
- tests/integration_tests.rs — 1 commit
Notable commits
- fix: fix(cli): accept flags before the PDF path and embed CMaps (#570)
- fix: fix(detector): a text layer nobody sees under a covering image is a scan (#566)
- fix: fix(detector): judge a CID-font page's text decoded before calling it vector text (#551)
- fix: fix(extract): read right-to-left lines back through the Unicode Bidirectional Algorithm (#552)
- fix: fix(extractor): Type1 fonts read through their program's built-in encoding (#596)
- fix: fix(extractor): a return from a zero-advance sign placed behind the pen is no word gap (#564)
- fix: fix(extractor): a super- or subscript run is sized by its letters and digits (#603)
- fix: fix(extractor): compare a short string's readings without its short common words (#597)
- fix: fix(extractor): compose a detached spacing accent with the letter it is painted over (#563)
- fix: fix(extractor): keep tracked display text whole when its letter spacing is written as TJ offsets (#548)
- fix: fix(extractor): leave out text painted wholly outside its rectangular clip (#539)
- fix: fix(extractor): lines of small type keep apart — the line window follows the type where the type is small (#580)
- fix: fix(extractor): read word gaps written as character spacing (#530)
- fix: fix(extractor): simple fonts read one byte per code under a two-byte ToUnicode codespace (#595)
- fix: fix(extractor): start TJ sub-runs at their first painted glyph (#527)
- fix: fix(extractor): underline detection no longer goes quadratic on pages drawn from thin rects; add extractTextWithPositionsAsync (#592)
- fix: fix(fonts): decode simple fonts through their base and built-in encodings and glyph names (#553)
- fix: fix(fonts): read a re-encoded font by its Differences under a stale ToUnicode CMap; keep one glyph's characters together through the visual-order read-back (#562)
- fix: fix(fonts): read ligature glyph names by their components instead of dropping them (#558)
- fix: fix(load): saturate form /BBox numerals no parser holds so the form is read (#560)
- …and 23 more
Architecture
- Containers 0 added · 0 removed · contexts 1 added · 0 removed · edges 0 added · 0 removed
Added bounded contexts (1)
- pdf-inspector
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
firecrawl/pdf-inspector was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 28 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit ef52f77850b29189048797a48b24862251f54fe7 — the exact code this score is about.
- Scored under rubric-2026.09.16 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-d46da229e3fd.