Skip to content
CAI
Software that uses CAICheck a score

firecrawl/anydoc

72.2

Strong · 28 September 2026

25.8k

lines of production code

Rust

primary language

2

measurements over time

CAI band scale
CAI trend line
CAI lens gauges

What this system is

This system is a high-fidelity document conversion library that transforms a wide variety of file formats—including legacy Office, OpenDocument, EPUB, and PDF—into GitHub-Flavored Markdown. It provides a structured, format-agnostic document model that preserves complex elements like nested lists, tables, LaTeX math, and embedded assets, while handling edge cases such as scanned PDFs and encrypted files through explicit error reporting. The core Rust engine is exposed via official bindings for Node.js, Python, and WebAssembly, enabling integration into server-side applications and browser-based interfaces.

Features

Add browser-based anydoc demo with high-DPI-accurate UI

A new interactive demo page is now available at wasm/www, allowing users to convert documents to Markdown entirely in the browser via WebAssembly. The page includes a branded UI with light and dark mode support, custom fonts, and a layout frame drawn with SVG to ensure crisp, hairline-accurate corners on high-DPI displays.

wasm/www · high confidence

Added version consistency check script

A new shell script, scripts/check-versions.sh, has been added to verify that the release version is synchronized across five locations: Cargo.toml, python/Cargo.toml, wasm/Cargo.toml, node/package.json, and the version guard in node/index.js. This tool helps prevent release errors by ensuring all package definitions and generated artifacts agree on the same version number.

scripts · high confidence

Content-based format detection and PDF Markdown conversion

The system now automatically detects document formats from their binary signatures and container structures (such as OLE compound files and ZIP packages) rather than relying solely on file extensions, with plain-text formats like CSV falling back to extension-based detection. Additionally, PDF files are now supported via direct Markdown extraction using the pdf-inspector library, which reports specific pages requiring OCR when text cannot be extracted, rather than attempting to parse them into the standard document model.

src/formats · high confidence

Initial EPUB format support with spine-aware navigation

Added a new EPUB parser that reads the book's spine to determine reading order, ensuring that intra-book links and navigation work correctly by converting relative links to scoped anchors for spine items while keeping other resources (like images or non-linear content) as relative paths. The implementation also handles chapter titles, extracts CSS stylesheets (with caching to avoid redundant parsing), and manages asset extraction for images.

src/formats/epub · high confidence

Initial OpenDocument Format (ODF) support for text, spreadsheets, and presentations

This change introduces the \src/formats/odf\ module, enabling the conversion of OpenDocument Text (.odt), Spreadsheet (.ods), and Presentation (.odp) files. The implementation parses \content.xml\ and \styles.xml\ to reconstruct document structure, supporting features such as heading levels, nested lists with composite numbering, tables with header rows and cell spanning, and embedded images. It also handles specific content types like presentation slides (including speaker notes as blockquotes), spreadsheet form controls (rendered as checkboxes), and mathematical equations (converted to LaTeX). The parser includes resource-limit protections against table expansion attacks and correctly identifies encrypted ODF packages.

src/formats/odf · high confidence

Introduce @firecrawl/anydoc Node.js bindings and CLI

The npm package @firecrawl/anydoc is now available, providing Node.js bindings and a CLI for converting Word, PowerPoint, Excel, OpenDocument, RTF, EPUB, CSV, and PDF files into GitHub-Flavored Markdown. The library exposes \toMarkdown\, \toMarkdownBytes\, and \toDocument\ functions, along with format detection utilities. For PDFs with scanned or image-only pages, the default behavior is to reject with a \needsOcr\ error listing the affected pages; users can opt into hosted OCR via Firecrawl Parse by setting \ocr: 'hosted'\. The package also ships an \anydoc\ CLI command for converting documents from the terminal or stdin.

node · high confidence

Introduce @firecrawl/anydoc-wasm WebAssembly bindings and demo

This change adds the \@firecrawl/anydoc-wasm\ package, providing WebAssembly bindings that mirror the Rust library's API for in-memory document conversion (e.g., \toMarkdownBytes\, \toDocument\) without filesystem access. It introduces specific error handling via \Error\ objects with \code\ properties (such as \needsOcr\, \encrypted\, \unsupported\) to allow callers to handle issues like scanned PDF pages or encryption explicitly. The release includes a browser-based demo site and Node.js smoke tests to verify round-trip conversion and error reporting.

wasm · high confidence

Introduce Python bindings for anydoc as the firecrawl-anydoc package

Users can now install the \firecrawl-anydoc\ package via PyPI to convert Word, PowerPoint, Excel, OpenDocument, RTF, EPUB, CSV, and PDF files into GitHub-Flavored Markdown. The bindings expose \to\_markdown\ and \to\_markdown\_bytes\ for direct conversion, as well as \to\_document\ to access the underlying document model with embedded assets. The package includes type stubs for static analysis and handles format detection from file content. For PDFs with scanned pages, the library raises a \NeedsOcrError\ by default, but users can opt into hosted OCR via Firecrawl Parse by setting \ocr="hosted"\, which sends the document to the Firecrawl API for processing.

python · high confidence

Introduce WebAssembly bindings for document conversion and model access

New WebAssembly bindings allow converting documents to Markdown or parsing them into a structured document model directly in the browser. The \toMarkdownBytes\ function handles conversion from various formats (including DOCX, PDF, and EPUB), while \toDocument\ exposes a detailed object model containing blocks, inline elements, notes, and assets for programmatic manipulation. The bindings also include utilities to detect file formats from bytes or extensions, and provide specific error handling for cases like scanned PDFs requiring OCR.

wasm/src · high confidence

Introduce structured document model with support for LaTeX math and task-list checkboxes

The \src/model\ module now provides a comprehensive, format-agnostic document representation that replaces ad-hoc parsing structures. This new model introduces specific support for LaTeX math formulas (rendered as \Inline::Math\ and \Block::Math\) and checkbox controls (rendered as \Inline::Checkbox\ with task-list token support), enabling accurate preservation of these elements during conversion. It also standardizes the handling of embedded assets, hyperlinks, lists with resolved numbering, and complex table grids with span management, ensuring that all content is fully resolved and self-contained before reaching the renderer.

src/model · high confidence

New DOCX and PPTX import formats

Added support for importing Microsoft Word (.docx) and PowerPoint (.pptx) files. The DOCX parser resolves styles, numbering, and relationships to convert document blocks, tables, and notes into the internal model, while the PPTX parser handles the full text-property cascade (slide to layout to master) and includes speaker notes as quotes.

src/formats/docx · high confidence

New RTF parser with robust lexer and table support

Added a new RTF format parser in src/formats/rtf that includes a position-explicit lexer, prelude table parsing, and table assembly. The lexer correctly treats a backslash before a line break as a paragraph mark and handles binary payloads safely. The parser supports nested tables, list numbering from list tables, and style-based heading emphasis.

src/formats/rtf · high confidence

New binary .doc and .ppt parsers with robust list and style resolution

Added new parsers for legacy Word (.doc) and PowerPoint (.ppt) binary formats, introducing dedicated modules for list table resolution (PlfLst/PlfLfo), style sheet (STSH) inheritance, and property modifier (sprm) application. The Word parser now correctly handles complex list numbering, including restart limits and start-at overrides, while the PowerPoint parser supports master style inheritance and includes speaker notes as quote blocks. These changes improve fidelity for legacy Office documents by implementing the published MS-DOC and MS-PPT resolution algorithms.

src/formats/doc · high confidence

New competitor benchmark harness for anydoc

A new benchmarking suite has been added to the \bench\ directory to evaluate anydoc against well-known document-to-markdown converters (such as markitdown, pandoc, docling, unstructured, mammoth, and LibreOffice). The harness measures speed and quality across Office, text, and presentation documents (excluding PDFs) using a corpus located in \samples/\. It provides deterministic metrics (structure counts and word-trigram containment) and an LLM-based quality judge (using Claude Sonnet 5) that compares anydoc's output against competitors using ground-truth page images. The suite includes scripts for running conversions (\convert.py\), generating ground truth (\render\_truth.py\), calculating metrics (\metrics.py\), performing LLM judging (\judge.py\), and aggregating results into a final report (\report.py\).

bench · high confidence

New shared conversion utilities and resource safety bounds

The \src/shared\ module introduces a suite of new components to support document conversion, including an \AssetSink\ that enforces a hard cap on embedded asset bytes and deduplicates repeated references, and a \StyleChains\ system that safely traverses style inheritance while detecting cycles. It adds \blockstyle.rs\ to map producer-specific paragraph style names (like 'Quote' or 'Code') into semantic block containers, and \header.rs\ to automatically detect header rows in CSV and spreadsheet data based on column type consistency. The module also includes \binary.rs\ for safe, bounded reading of legacy OLE2 streams, \fields.rs\ for parsing Word hyperlink instructions, \grid.rs\ for assembling complex table grids with merged cells, and \drawingml.rs\ for extracting chart and SmartArt data. Additionally, it provides \math/mathml.rs\ for converting MathML to LaTeX, \html.rs\ for a minimal CSS subset and HTML-to-block conversion, and \list.rs\ for assembling nested lists with correct numbering identities.

src/shared · high confidence

Removals

Removal of legacy shared support modules

The shared infrastructure modules in src/support (fields, html, text, xml, zip) have been removed. This eliminates the previous implementations for Word field instruction parsing, HTML-to-block conversion, text normalization, XML parsing, and ZIP archive reading, indicating a shift in how these capabilities are handled within the application.

src/support · high confidence

Behavioural changes

Content-based format detection and typed error reporting

The library now detects document formats by inspecting file content signatures (via \Format::from\_bytes\) rather than relying solely on file extensions, with extensions serving as a fallback for signature-less formats like CSV. This change introduces a new \ConvertError\ enum that provides specific, machine-readable error codes (e.g., \needsOcr\, \unsupported\, \encrypted\) instead of generic errors, allowing callers to handle issues like missing OCR pages or unsupported formats programmatically. The public API functions \to\_markdown\ and \to\_markdown\_bytes\ now return \Result\<String, ConvertError\>\ to expose these detailed errors, and the internal intermediate representation (\ir\) has been replaced by a new \model\ module.

src · high confidence

In-house Excel reader replaces Calamine with support for checkboxes and improved number formatting

The spreadsheet parser in \src/formats/sheet\ has been rewritten to read all Excel containers (\.xlsx\, \.xlsm\, \.xlsb\, and legacy \.xls\) in-house, removing the external Calamine dependency. This change introduces support for form control checkboxes (\.xlsx\/\.xlsm\ VML drawings and \.xls\ OBJ records) which are now rendered inline with cell text. It also improves numeric fidelity by rendering floats at 15 significant digits to avoid binary noise and correctly formats time serials as \hh:mm:ss\ without the Excel epoch date offset.

src/formats/sheet · high confidence

Secure, limited ZIP archive parsing for OOXML/ODF/EPUB packages

The \src/package\ module now provides a hardened, in-house ZIP archive reader for OOXML, ODF, and EPUB documents, replacing external dependencies like calamine for container handling. It enforces strict resource limits (128 MiB per entry, 512 MiB total, 100k entries) to prevent decompression bombs, caches repeated reads to avoid budget exhaustion, and implements safe OPC/EPUB path resolution that rejects encoded path traversal and separators. The module also includes namespace-aware XML parsing with depth/node caps and a probe for legacy OLE/encrypted files, ensuring that malformed or malicious archives fail safely rather than consuming excessive resources.

src/package · high confidence

Unified document-to-Markdown conversion examples with asset extraction

The examples directory now provides consistent CLI tools for converting documents to Markdown across Node.js, Python, and Rust. These examples support explicit format specification via the \-f\ flag, output redirection with \-o\, and the extraction of embedded images and objects to a specified directory using \--assets\. The Rust example also removes the previous \--bench\ benchmarking mode in favor of this unified feature set.

examples · high confidence

Fixes

Fixes Markdown rendering for anchors, escaping, and tables

The Markdown serializer now correctly handles internal linking by only emitting HTML anchor tags for targets that are actually referenced by links in the document, preventing noise in the output. It applies context-sensitive escaping to prevent Markdown syntax characters (like emphasis delimiters, backticks, and pipes) from being misinterpreted, specifically fixing issues where delimiters paired incorrectly across run seams or where dollar signs formed unintended math spans. Table rendering has also been improved to properly escape pipes within code spans and to trim cell padding instead of encoding it as HTML entities.

src/render · high confidence

Test coverage

Add comprehensive test fixtures and helpers for document conversion; Add fuzzing infrastructure for document conversion targets.

Dependencies

Release v0.2.4 with Node, Python, and WebAssembly bindings

This release publishes the core document conversion library as v0.2.4 and introduces official bindings for Node.js (@firecrawl/anydoc), Python (firecrawl-anydoc), and WebAssembly (@firecrawl/anydoc-wasm). The Node and Python packages are built using NAPI and PyO3 respectively, enabling direct integration into JavaScript/TypeScript and Python applications. The WebAssembly bindings allow document conversion to run in the browser. Additionally, the core library has been updated to require Rust edition 2024 (Rust 1.88+), and the dependency on the calamine Excel parser has been removed in favor of an in-house implementation.

(dependencies) · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

Score

  • CAI 73 → 72 (-0.8)
  • Rubric changed (rubric-2026.09.8 → rubric-2026.09.16) — scores are not directly comparable.

Lenses

  • Code Health 78 → 78 (+0.0)
  • Architecture 100 → 93 (-7.0)
  • Maturity 70 → 70 (-0.0)
  • Readiness 71 → 67 (-3.9)
  • Security 74 → 83 (+8.5)
  • Accessibility 82 → 82 (+0.0)
  • Performance 100 (new)

Resolved (19)

  • Documentation: no architecture or design documentation (README.md)
  • Documentation: no installation or build instructions (README.md)
  • Hotspot: node/cli.js (node/cli.js)
  • Hotspot: src/formats/doc/lists.rs (src/formats/doc/lists.rs)
  • Hotspot: src/formats/doc/mod.rs (src/formats/doc/mod.rs)
  • Hotspot: src/formats/doc/sprm.rs (src/formats/doc/sprm.rs)
  • Hotspot: src/formats/odf/mod.rs (src/formats/odf/mod.rs)
  • Hotspot: src/formats/odf/table.rs (src/formats/odf/table.rs)
  • Hotspot: src/formats/ppt/mod.rs (src/formats/ppt/mod.rs)
  • Hotspot: src/formats/pptx/mod.rs (src/formats/pptx/mod.rs)
  • Hotspot: src/formats/rtf/tables.rs (src/formats/rtf/tables.rs)
  • Hotspot: src/formats/sheet/xls.rs (src/formats/sheet/xls.rs)
  • Hotspot: src/formats/sheet/xlsb.rs (src/formats/sheet/xlsb.rs)
  • Hotspot: src/formats/sheet/xlsx.rs (src/formats/sheet/xlsx.rs)
  • Hotspot: src/package/xml.rs (src/package/xml.rs)
  • Hotspot: src/render/markdown/mod.rs (src/render/markdown/mod.rs)
  • Hotspot: src/shared/grid.rs (src/shared/grid.rs)
  • Hotspot: src/shared/header.rs (src/shared/header.rs)
  • Hotspot: src/shared/list.rs (src/shared/list.rs)

New (16)

  • Documentation: contradicts the code (bench/README.md)
  • Duplication of format detection logic. The anydoc module methods accept an optional Format hint, but the Format type itself provides multiple static constructors (from_bytes, from_extension, from_path) to detect format. It is unclear if the format parameter in to_document is used for validation, optimization, or if the library ignores it and auto-detects using the Format constructors. This creates a confusing API surface where format detection is split between the consumer-facing functions and the type constructors.
  • Inconsistent input source naming and parameter structure. One method takes a file path, the other takes raw bytes. While the domains differ (file vs memory), the naming convention 'to_markdown' vs 'to_markdown_bytes' is redundant and confusing. It implies 'to_markdown' might also accept bytes or that there is a missing 'to_markdown_path' equivalent. Furthermore, the top-level module exposes these methods directly, but the Format type has its own from_* constructors, creating a split in responsibility for format detection.
  • Off the main sequence: anydoc
  • Outdated: encoding_rs
  • Outdated: flate2
  • Outdated: js-sys
  • Outdated: log
  • Outdated: napi
  • Outdated: napi-build
  • Outdated: napi-derive
  • Outdated: pdf-inspector
  • Outdated: pyo3
  • Outdated: wasm-bindgen
  • Projects may be oversized for their cohesion
  • Type inconsistency in Asset model. Asset.bytes is defined as u8 (a single byte), which is almost certainly a typo for Vec<u8> or &[u8]. This is a critical API error that breaks the type's intended purpose of storing asset data.

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

firecrawl/anydoc was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 28 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit 261fc257d17c3eab0f673be31c408fd9fdc2171a — the exact code this score is about.
  • Scored under rubric-2026.09.16 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-d46da229e3fd.