yob/pdf-reader
64.6
Adequate · 19 September 2026
8.5k
lines of production code
Ruby
primary language
1
measurement over time
What this system is
This system is a Ruby library for parsing PDF files, providing capabilities to extract text, metadata, fonts, and images from documents. It supports reading encrypted PDFs using AES-256 encryption and includes utilities for inspecting low-level PDF structures and stream filters. The library ensures robust text extraction through embedded font metrics and strict type checking, while offering command-line tools and examples for analysis and debugging.
How it got here
2006–2010 — Library initialization and core feature development
10 changes.
This period established the foundational structure of the pdf-reader library, introducing core parsing capabilities, command-line utilities, and comprehensive test suites. It focused on building a robust API for text extraction and object inspection, while adding support for modern encryption standards and strict type checking.
2011–2012 — PDF parsing robustness and refactoring
7 changes.
This period focused on refactoring core PDF parsing components, such as stream filters and glyph width calculators, into dedicated classes with improved error handling and strict typing. The work included embedding standard font metrics to ensure accurate text extraction and adding comprehensive test suites and development utilities to validate these changes and enhance overall library stability.
2013–2021 — CI migration and type safety adoption
5 changes.
The project migrated its continuous integration infrastructure from Travis to Buildkite, introducing automated release pipelines and consistent Docker-based development environments. Concurrently, it implemented Sorbet static type checking across the codebase, enforcing strict typing requirements and generating RBI stubs for dependencies. These efforts were supported by expanded test coverage for core PDF parsing components and new scripts for memory benchmarking and quality assurance.
Features
Added memory benchmarking script and strict typing enforcement check
A new script, scripts/benchmark\_allocations.rb, has been added to measure and compare memory allocations for various PDF parsing scenarios (such as cairo-unicode, type1-arial, and truetype-arial) using the benchmark-memory gem. Additionally, a new build step script, scripts/require-strict-typing, enforces that any source files added to the lib/ directory since commit feb9203 must include the 'typed: strict' Sorbet directive, failing the build if this requirement is not met.
scripts · high confidence
Initial Sorbet type checking configuration and gem RBI stubs
Added Sorbet static type checking support by creating the \sorbet/config\ file to configure the type checker (scanning the root directory while ignoring vendor bundles) and generating initial RBI stub files for internal annotations and external gems (including Rainbow, Ascii85, AFM, AST, Cane, CodeRay, Commander, Diff-LCS, FFI, Hashery, HighLine, MethodSource, and Morecane) to enable type safety across the codebase.
sorbet · high confidence
Initial project scaffolding and configuration
Establishes the foundational structure for the pdf-reader library by adding essential configuration and metadata files. This includes a .gitattributes file to prevent Windows checkouts from altering line endings in PDF test data, a .gitignore to exclude build artifacts and lockfiles, and an .rspec file to standardize test runner options. The release also introduces a CHANGELOG documenting the history up to v2.16.0, a MIT-LICENSE file, a README with usage instructions, a Rakefile for building and testing, and a TODO list outlining future development goals.
(repo-wide) · high confidence
Introduce PDF::Reader as a new PDF parsing library
The lib/pdf directory now includes the PDF::Reader library, providing a new entry point for parsing PDF files. This addition introduces a page-based iteration model for extracting content, metadata, and fonts, along with support for encrypted files via password options. The library exposes methods for accessing document info, XML metadata, and page counts, and allows direct access to raw PDF objects through an ObjectHash interface.
lib/pdf · high confidence
New and updated example scripts for PDF extraction and inspection
The examples directory has been reorganized and expanded with new scripts demonstrating how to extract specific PDF content, including images (JPG, TIFF, CMYK/RGB/Gray), TrueType fonts, and Bates numbers, as well as how to inspect page callbacks, raw PDF objects, metadata, page counts, and PDF versions. Existing examples for text extraction and RSpec testing have also been updated to align with the current API.
examples · high confidence
New command-line utilities for PDF inspection and text extraction
This release introduces three new executable scripts in the bin directory to aid in PDF analysis. The \pdf\_text\ script extracts and prints text content from all pages of a PDF file, supporting both file arguments and standard input. The \pdf\_object\ script allows users to retrieve and inspect specific PDF objects by their ID and generation number, displaying their structure or stream data. Additionally, the \pdf\_callbacks\ script provides a way to walk through page content streams and print receiver events, useful for debugging or understanding low-level PDF structure.
bin · high confidence
New development and testing utilities for PDF parsing
Added several new scripts in the tools directory to aid in development and quality assurance: bench.rb for benchmarking text extraction performance and profiling object allocations; fuzz.rb for generating mutated PDF files to test robustness and catch exceptions; page\_bench for per-page text extraction timing; profile.rb to orchestrate parallel profiling runs; and read-pdf.rb as a diagnostic script to load, parse, and display metadata and content of PDF files.
tools · high confidence
New text filtering and AES-256 encryption support
PDF::Reader now includes an AdvancedTextRunFilter to allow users to filter extracted text runs based on attributes like font size, text content, and width using logical operators. It also adds support for AES-256 (AESV3) encryption, enabling the reading of PDFs secured with revision 5 or 6 encryption standards.
lib/pdf/reader · high confidence
Behavioural changes
Add lib/pdf-reader.rb entry point with strict typing and encoding markers
The library now includes a new lib/pdf-reader.rb file that serves as the main entry point. This file requires the 'pdf/reader' gem and explicitly sets UTF-8 encoding, strict type checking, and frozen string literals to ensure consistent behavior and performance across the library.
lib · high confidence
Embedded standard PDF font metrics for accurate text rendering
The PDF reader now includes embedded AFM (Adobe Font Metrics) files for the 14 standard PDF fonts, including Courier (Regular, Bold, Oblique, BoldOblique) and Helvetica (Regular, Bold, Oblique, BoldOblique). By bundling these metrics directly in the \lib/pdf/reader/afm\ directory, the library can accurately calculate character widths and layout for these fonts without relying on external system fonts, ensuring consistent and correct text rendering in generated or analyzed PDF documents.
lib/pdf/reader/afm · high confidence
Migrate CI pipeline from Travis to Buildkite
The continuous integration system has moved from Travis CI to Buildkite, introducing a new pipeline structure defined in .buildkite/pipeline.yml and .buildkite/pipeline.release.yml. The new build process enforces stricter quality gates by running Sorbet type checking and a 'typed: strict' requirement check before executing the RSpec test matrix. The test suite now covers a broader range of Ruby implementations, including MRI versions 2.1 through 4.0, JRuby 9.1 through 10.1, and TruffleRuby 33.0.1 and 34.0.1, with TruffleRuby jobs configured to soft-fail. Additionally, a dedicated release pipeline is introduced to handle gem publishing using the rubygems-oidc plugin and Docker.
.buildkite · high confidence
New automation scripts for CI, testing, and gem release
This change introduces a suite of new shell scripts in the \auto/\ directory to streamline development and release workflows. The \with-ruby\ script provides a consistent Docker-based environment for running commands, defaulting to Ruby 3.4. New scripts \run-specs\, \run-quality\, and \run-sorbet\ automate the execution of RSpec tests, quality checks, and Sorbet type checking (with support for inline RBS comments). Additionally, \bundle-exec\ ensures dependencies are installed before running commands, while \release-gem\ and \upload-release-steps\ automate the generation of RBI/RBS signatures, building the gem, and uploading it to RubyGems via Buildkite.
auto · high confidence
Refactored PDF stream filters into dedicated classes with improved error handling
The PDF stream filter implementations (Ascii85, AsciiHex, Depredict, Flate, LZW, Null, and RunLength) have been split into individual, single-responsibility classes within the \lib/pdf/reader/filter\ directory. This change introduces stricter Sorbet typing (\\# typed: strict\) and enables \frozen\_string\_literal\ for better performance and safety. Notably, the Flate filter now includes a fallback mechanism to handle zlib inflation failures by retrying with the final byte removed, and all filters now raise \MalformedPDFError\ with specific context when decoding fails, rather than propagating generic exceptions.
lib/pdf/reader/filter · high confidence
Refactored glyph width calculation into specialized calculator classes
The PDF reader's text extraction logic has been restructured to use dedicated width calculator classes for different font types (Built-in, Composite, TrueType, Type0, and Type1/Type3). This change improves the reliability of text extraction by ensuring numeric values are always returned for glyph widths, handling missing or empty font metrics gracefully, and correctly processing built-in standard fonts using embedded AFM metrics. Users will experience more robust text extraction from PDFs containing various font types, particularly those with missing width data or control characters.
_lib/pdf/reader/width\calculator · high confidence
Test coverage
Added invalid PDF test fixtures for robustness testing; Added shared example for WidthCalculator duck type; Added test coverage for PDF stream filters; Added test fixtures for PDF parsing edge cases; Added test support helpers for PDF parsing and encoding validation; Expanded test coverage for PDF::Reader components; Expanded test suite for PDF parsing integrity and encoding.
Dependencies
Update pdf-reader to version 2.16.0 with modernized build tooling
This release updates the pdf-reader gem to version 2.16.0 and modernizes its development environment. The gem now requires Ruby 2.1 or higher and includes Sorbet (0.6.12872) and Tapioca (0.17.10) for static type checking and RBI generation, alongside Spoom and rbs-inline for RBS annotation support. Development dependencies have been updated to include RSpec (\~\> 3.5), Cane (\~\> 3.0), Morecane (\~\> 0.2), Pry, and RDoc, while runtime dependencies now explicitly allow Ascii85 versions \>= 1.0 and \< 3.0 (excluding 2.0.0), Hashery (\~\> 2.0), and AFM (\>= 0.2.1, \< 2). The gemspec also adds project metadata for bug tracker, changelog, and documentation URIs.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Baseline
- First survey — no prior run to compare against. CAI 65.
Lenses
- Code Health 94
- Architecture 100
- Maturity 55
- Readiness 58
- Security 81
Changes since last survey
- 300 commits — 273 feature/other, 27 fixes
By area
- (repo) — 115 commits
- lib/pdf — 109 commits
- (root) — 39 commits
- .buildkite/pipeline.yml — 15 commits
- rbi/pdf-reader.rbi — 6 commits
- .buildkite/pipeline.release.yml — 4 commits
- spec/reader — 3 commits
- spec/data — 2 commits
- auto/release-gem — 1 commit
- auto/run-sorbet — 1 commit
- auto/with-ruby — 1 commit
- scripts/benchmark_allocations.rb — 1 commit
- sorbet/rbi — 1 commit
- spec/encoding_spec.rb — 1 commit
- spec/integration_spec.rb — 1 commit
Notable commits
- fix: Fix RBI signatures
- fix: Fix RC4 for Ruby 2.1-2.3 compatibility
- fix: Fix ToUnicode handling for symbolic fonts
- fix: Fix assigned but unused variable warnings
- fix: Fix broken type sigs for resource methods
- fix: Fix decoding of some UTF-16 strings that use surrogate pairs
- fix: Fix indent in readme
- fix: Fix sorbet strict typing and split rc4 specs per review
- fix: Fix white space on type declaration
- fix: Fix xref objid reset when object 1 is a legitimate free entry
- fix: Merge pull request #471 from yob/fix-type-sigs-for-resources
- fix: Merge pull request #494 from ShockwaveNN/fix/typo-in-documentation
- fix: Merge pull request #498 from yob/fix-reading-tempfile
- fix: Merge pull request #502 from fac/fix-warnings
- fix: Merge pull request #512 from iMacTia/mg/fix-sorbet-sig
- fix: Merge pull request #529 from yob/fix-utf16-surrogate-pairs
- fix: Merge pull request #540 from yob/fix-ruby-2-ci
- fix: Merge pull request #550 from olivier-thatch/olivier-fix-rbi
- fix: Merge pull request #567 from yob/fix-stack-overflow-in-page-ancestors
- fix: Merge pull request #573 from yob/tounicode-fix
- …and 280 more
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
yob/pdf-reader was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 19 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 0178487145988c3fa8371a444b1abbd67d3a96cf — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-13a154b7f5d1.