Skip to content
CAI
Software that uses CAICheck a score

mozilla/readability

44.9

Weak · 2 October 2026

3.9k

lines of production code

JavaScript

primary language

2

measurements over time

CAI band scale
CAI trend line
CAI lens gauges

What this system is

This system is a standalone JavaScript library that extracts the main article content and metadata from complex HTML web pages. It provides a lightweight DOM parser for non-browser environments and cleans the extracted content by removing scripts, styles, and hidden elements while preserving text, images, and embedded media. The library also detects whether a page is likely to contain readable article content and handles diverse metadata sources and international text directions.

How it got here

2015 — Standalone release and test expansion

39 changes.

The project released Readability.js as a standalone npm package for Node.js environments, introducing core parsing logic and TypeScript definitions. This period was dominated by extensive test coverage additions, validating the parser's ability to extract content and metadata from a wide variety of publisher sites and handle specific HTML edge cases.

2016–2018 — Test coverage expansion

26 changes.

This period focused on expanding the automated test suite by adding fixtures for a wide variety of publisher sites, including BBC, New York Times, Yahoo, and Medium. It also introduced tests for specific parsing edge cases such as right-to-left text direction, hidden nodes, and base URL resolution to ensure robust content extraction.

2019–2025 — Comprehensive test coverage expansion

30 changes.

This period focused on systematically expanding the test suite with fixtures for diverse content sources, including major news outlets, technical blogs, and documentation sites. The work verified the parser's ability to handle complex metadata schemas, embedded media, lazy-loaded images, and specific HTML edge cases like entity unescaping and hidden content filtering.

Features

Initial release of standalone Readability.js library

This change introduces the standalone version of the Readability library, originally used for Firefox Reader View, as a publishable npm package. It includes the core parsing logic in Readability.js and a lightweight DOM parser (JSDOMParser.js) for environments without a native DOM implementation. The release also provides TypeScript definitions (index.d.ts) for the Readability and isProbablyReaderable APIs, a pre-configured ESLint setup, and a CHANGELOG.md documenting the history from version 0.3.0 through 0.6.0. This makes the library available for use in Node.js and other JavaScript environments outside of the Firefox browser.

(repo-wide) · high confidence

Test coverage

Add WebMD test page with expected metadata and HTML output; Add Yahoo test page for article extraction; Add eHow test case for terrarium article parsing; Add test case for The Independent article parsing; Add test coverage for CNET article parsing; Add test coverage for IETF remoteStorage draft page; Add test coverage for Libération article with embedded videos and JSON-LD metadata; Add test coverage for Medium article extraction; Add test coverage for Yahoo! News Japan article parsing; Add test coverage for article with embedded videos and published time metadata; Add test coverage for lazy-loaded images on Kinja sites; Add test coverage for pages with missing paragraphs; Add test fixtures for frontend JavaScript testing and Quanta Magazine articles; Add test page for CNET SVG class handling; Add test page for IAB article on digital ad UX; Add test page for Salon article on Uber and the sharing economy; Add test page for article author tag extraction; Added BBC News test page for article extraction; Added Engadget test case for Xbox One X review; Added RTL test pages to verify text direction handling; Added test case for 'Bartleby the Scrivener' article cleaning; Added test case for data URL image handling; Added test case for title and H1 discrepancy; Added test coverage for 'toc-missing' article extraction; Added test coverage for \<base\> tag URL resolution; Added test coverage for Ars Technica article parsing; Added test coverage for BR replacement logic; Added test coverage for Blogger site parsing; Added test coverage for CNN article parsing; Added test coverage for Daring Fireball; Added test coverage for Fandom (Wikia) article parsing; Added test coverage for Folha de S.Paulo article parsing; Added test coverage for Google SRE Book chapter 6; Added test coverage for HTML entity unescaping; Added test coverage for Le Monde article parsing; Added test coverage for Libération.fr article parsing; Added test coverage for Parsely metadata extraction; Added test coverage for QQ.com articles; Added test coverage for SVG parsing; Added test coverage for The Guardian article parser; Added test coverage for The New York Times en Español; Added test coverage for The Verge article parsing; Added test coverage for V8 blog post on standalone WebAssembly; Added test coverage for Washington Post article parsing; Added test coverage for Yahoo News article parsing; Added test coverage for a New York Times article with complex metadata and schema markup; Added test coverage for base URL resolution; Added test coverage for basic tag cleaning; Added test coverage for embedded video handling; Added test coverage for gmw.cn article parsing; Added test coverage for heise.de article parsing; Added test coverage for hidden node handling; Added test coverage for image lists and figures; Added test coverage for javascript: link replacement; Added test coverage for keeping article images; Added test coverage for lazy-loaded images with alt text extensions; Added test coverage for links-in-tables parsing; Added test coverage for metadata extraction edge cases; Added test coverage for metadata extraction with missing content attributes; Added test coverage for nested font tag replacement; Added test coverage for ordered list preservation; Added test coverage for paragraph reordering; Added test coverage for pixnet.net; Added test coverage for removing aria-hidden elements; Added test coverage for schema.org context object handling; Added test coverage for script tag removal and extra paragraph cleanup; Added test coverage for script tags containing HTML comments; Added test coverage for social button removal; Added test coverage for style tag removal and space normalization; Added test coverage for trailing \<br\> removal; Added test coverage for visibility-hidden content filtering; Added test fixtures for LWN.net weekly edition; Added test fixtures for Medium article extraction; Added test fixtures for TMZ and Washington Post articles; Added test fixtures for Tumblr page parsing; Added test fixtures for aktualne.cz article parsing; Added test fixtures for two New York Times articles; Added test infrastructure and coverage for JSDOMParser and isProbablyReaderable; Added test page for Archive of Our Own fanfiction; Added test page for Factorio blog tabular data; Added test page for Fetch API article; Added test page for Firefox Developer Edition first-run experience; Added test page for Firefox Nightly blog issue 85; Added test page for Firefox customization; Added test page for GitLab blog article; Added test page for New York Times article on NYC infrastructure; Added test page for WebMD article parsing; Added test page for Wikipedia article on films featuring time loops; Added test page for eHow article extraction; Added test pages for Mercurial documentation and en-dash title handling; Added test pages for lazy image handling and topic seed content extraction; Initial test coverage for WordPress blog parsing; Updated test expectations for The Seattle Times article parsing.

Dependencies

Release 0.6.0 with updated dependencies and Node 14 requirement

This release updates the project to version 0.6.0 and introduces a new package-lock.json to manage dependencies. The minimum supported Node.js version is now 14.0.0, and several development dependencies have been updated, including eslint to 8.57.0, mocha to 11.7.6, jsdom to 20.0.2, and release-it to 20.2.1.

(dependencies) · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

This is the PUBLIC form of this artifact. Findings are listed in full, but the details of SECURITY findings — which rule fired, in which file, on which line, and how to fix it — are deliberately withheld, and any secret-scanner results are excluded entirely. Where detail is absent here it was REMOVED FOR PUBLICATION; it is not missing from the analysis. The complete artifact is available from the repository owner.

Score

  • CAI 45 → 45 (-0.3)
  • Rubric changed (rubric-2026.09.12 → rubric-2026.09.18) — scores are not directly comparable.

Lenses

  • Code Health 30 → 30 (-0.8)
  • Architecture 99 → 100 (+0.1)
  • Maturity 52 → 52 (+0.0)
  • Readiness 47 → 48 (+1.1)
  • Security 87 → 85 (-2.3)

Resolved (5)

  • Dependency hygiene PARTLY measured — npm pinning read, dependency currency not (no pnpm-resolved versions to grade)
  • Documentation: no licence statement (README.md)
  • High CVE: [GHSA redacted] (package-lock.json)
  • High CVE: [GHSA redacted] (package-lock.json)
  • High CVE: [GHSA redacted] (package-lock.json)

New (8)

  • Dependency hygiene PARTLY measured — npm pinning read, dependency currency not (the committed lockfile resolved no direct production dependency)
  • Documentation: no architecture or design documentation (README.md)
  • FileTooLong: JSDOMParser.js (JSDOMParser.js)
  • High CVE: [GHSA redacted] (package-lock.json)
  • High CVE: [GHSA redacted] (package-lock.json)
  • High CVE: [GHSA redacted] (package-lock.json)
  • High CVE: [GHSA redacted] (package-lock.json)
  • Medium: security finding (details withheld)

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

mozilla/readability was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 2 October 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit ab4027a8b37669745016869a37a504727992b2ba — the exact code this score is about.
  • Scored under rubric-2026.09.18 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-e569280dd5e2.