Skip to content
CAI
Software that uses CAICheck a score

microsoft/markitdown

67.7

Adequate · 5 August 2026

6.5k

lines of production code

Python

primary language

4

measurements over time

CAI band scale
CAI trend line
CAI lens gauges

What this system is

MarkItDown is a Python library and ecosystem designed to convert various document formats into Markdown. It provides a core conversion engine that handles text, images, and structured data from sources like PDFs, Word documents, and spreadsheets, while also supporting external services like Azure and LLMs for advanced extraction. The system is extensible through a plugin architecture, allowing third-party converters for formats like RTF or OCR-based image text extraction. Additionally, it includes an MCP server interface for programmatic access.

Features

Add OCR plugin for extracting text from images in PDF, DOCX, PPTX, and XLSX files

The \markitdown-ocr\ package introduces a new plugin that enables LLM-based text extraction from images embedded in PDF, DOCX, PPTX, and XLSX files. The plugin registers enhanced converters for each supported format that extract embedded images, send them to a configured LLM (such as GPT-4o) for text extraction, and insert the resulting OCR text inline or alongside the document content. The plugin requires an \llm\_client\ and \llm\_model\ to be passed to \MarkItDown\ to activate the OCR functionality; without them, the plugin loads but OCR is silently skipped. The plugin is registered with a priority of -1.0, ensuring it runs before the built-in converters.

packages/markitdown-ocr · high confidence

Add new converters for audio, CSV, EPUB, and external services

The library now supports converting audio files (WAV, MP3, M4A, MP4) by extracting metadata and transcribing speech, and converts CSV files into Markdown tables. It also adds support for EPUB e-books, Bing search result pages, and integrates with Azure services: Document Intelligence for document analysis and Azure Content Understanding for multi-modal extraction. Additionally, the HTML converter now handles deeply nested structures that previously caused a RecursionError by falling back to plain-text extraction, and the markdownify processor now supports checkbox rendering and prevents JavaScript links.

packages/markitdown/src/markitdown/converters · high confidence

Add sample RTF plugin with versioning

A new sample plugin for converting RTF files to Markdown has been added to the MarkItDown ecosystem. The plugin registers an RtfConverter that accepts .rtf files and text/rtf or application/rtf MIME types, using the striprtf library to extract text. The package includes versioning (\_\about\\.py, \\init\\_.py) and type hints (py.typed) to support plugin discovery and integration.

packages/markitdown-sample-plugin/src · high confidence

MarkItDown MCP server now supports HTTP transport and environment-based plugin configuration

The MarkItDown MCP server now supports running as an HTTP service via Streamable HTTP (preferred) and the deprecated SSE transport, accessible through the new \--http\ and \--sse\ CLI flags. Additionally, the server now reads the \MARKITDOWN\_ENABLE\_PLUGINS\ environment variable to control plugin activation, allowing users to enable or disable plugins at runtime without code changes.

packages/markitdown-mcp/src · high confidence

Behavioural changes

MarkItDown library release 0.1.7

The MarkItDown library has been updated to version 0.1.7. This release includes a variety of fixes and improvements, such as better handling of linked images in DOCX files, support for preserving base64-encoded images, and the addition of CSV to Markdown table conversion. The CLI now supports extension, MIME type, and charset hints, and the library has switched from puremagic to magika for file type detection. Additionally, the library now supports Azure Content Understanding and Document Intelligence for cloud-based conversion, and includes a plugin system for third-party converters.

packages/markitdown/src/markitdown · high confidence

Fixes

Improved math equation rendering in .docx files

The .docx converter now renders math equations as LaTeX, improving the accuracy of mathematical content in Word documents. This update includes a fix for invalid LaTeX macros for mu, nu, tau, and down-arrow in equation conversion, ensuring that special characters and symbols are correctly translated. Additionally, the OMML template bugs have been resolved, leading to more reliable conversion of mathematical expressions. Users will see enhanced support for complex mathematical notations in their converted documents.

_packages/markitdown/src/markitdown/converter\utils · high confidence

Test coverage

Added comprehensive test coverage for markitdown CLI, module, and converters; Added test fixtures for diverse document formats; Added test infrastructure for the MarkItDown MCP server; Added tests for the MarkItDown sample plugin's RTF conversion.

Dependencies

Initial release of MarkItDown and its plugins

This change introduces the core MarkItDown library along with three new packages: markitdown-mcp (an MCP server), markitdown-ocr (an OCR plugin), and markitdown-sample-plugin (a sample plugin). The main markitdown package now depends on magika for file type detection, defusedxml for secure XML parsing, and optional dependencies for various file formats (e.g., mammoth for DOCX, pdfminer.six and pdfplumber for PDFs). The mcp package is pinned to version 1.8.0, and mammoth is specified at version 1.11.0.

(dependencies) · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

Score

  • CAI 68 → 68 (-0.4)
  • Rubric changed (rubric-2026.08.18 → rubric-2026.08.19) — scores are not directly comparable.

Lenses

  • Code Health 80 → 80 (+0.3)
  • Architecture 100 → 100 (+0.0)
  • Maturity 72 → 70 (-1.6)
  • Readiness 57 → 57 (+0.0)
  • Security 88 → 87 (-0.5)

Resolved (2)

  • Duplicated block (11 lines × 2) (packages/markitdown-ocr/src/markitdown_ocr/_pptx_converter_with_ocr.py)
  • Off-boarding risk: anonymized user #1

New (6)

  • Duplicated block (7 lines × 2) (packages/markitdown-ocr/src/markitdown_ocr/_pptx_converter_with_ocr.py)
  • Hotspot: packages/markitdown/src/markitdown/main.py (packages/markitdown/src/markitdown/main.py)
  • Hotspot: packages/markitdown/src/markitdown/_markitdown.py (packages/markitdown/src/markitdown/_markitdown.py)
  • Medium IaC: CKV_DOCKER_2 (Dockerfile)
  • Medium IaC: CKV_DOCKER_2 (packages/markitdown-mcp/Dockerfile)
  • Off-boarding risk: anonymized user #1

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

microsoft/markitdown was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 5 August 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit fd239d5d2be43d9b68329730206b9312c7d5a388 — the exact code this score is about.
  • Scored under rubric-2026.08.19 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer latest.