microsoft/markitdown
67.7
Adequate · 5 August 2026
6.5k
lines of production code
Python
primary language
4
measurements over time
What this system is
MarkItDown is a Python library and ecosystem designed to convert various document formats into Markdown. It provides a core conversion engine that handles text, images, and structured data from sources like PDFs, Word documents, and spreadsheets, while also supporting external services like Azure and LLMs for advanced extraction. The system is extensible through a plugin architecture, allowing third-party converters for formats like RTF or OCR-based image text extraction. Additionally, it includes an MCP server interface for programmatic access.
Features
Add OCR plugin for extracting text from images in PDF, DOCX, PPTX, and XLSX files
The \markitdown-ocr\ package introduces a new plugin that enables LLM-based text extraction from images embedded in PDF, DOCX, PPTX, and XLSX files. The plugin registers enhanced converters for each supported format that extract embedded images, send them to a configured LLM (such as GPT-4o) for text extraction, and insert the resulting OCR text inline or alongside the document content. The plugin requires an \llm\_client\ and \llm\_model\ to be passed to \MarkItDown\ to activate the OCR functionality; without them, the plugin loads but OCR is silently skipped. The plugin is registered with a priority of -1.0, ensuring it runs before the built-in converters.
packages/markitdown-ocr · high confidence
Add new converters for audio, CSV, EPUB, and external services
The library now supports converting audio files (WAV, MP3, M4A, MP4) by extracting metadata and transcribing speech, and converts CSV files into Markdown tables. It also adds support for EPUB e-books, Bing search result pages, and integrates with Azure services: Document Intelligence for document analysis and Azure Content Understanding for multi-modal extraction. Additionally, the HTML converter now handles deeply nested structures that previously caused a RecursionError by falling back to plain-text extraction, and the markdownify processor now supports checkbox rendering and prevents JavaScript links.
packages/markitdown/src/markitdown/converters · high confidence
Add sample RTF plugin with versioning
A new sample plugin for converting RTF files to Markdown has been added to the MarkItDown ecosystem. The plugin registers an RtfConverter that accepts .rtf files and text/rtf or application/rtf MIME types, using the striprtf library to extract text. The package includes versioning (\_\about\\.py, \\init\\_.py) and type hints (py.typed) to support plugin discovery and integration.
packages/markitdown-sample-plugin/src · high confidence
MarkItDown MCP server now supports HTTP transport and environment-based plugin configuration
The MarkItDown MCP server now supports running as an HTTP service via Streamable HTTP (preferred) and the deprecated SSE transport, accessible through the new \--http\ and \--sse\ CLI flags. Additionally, the server now reads the \MARKITDOWN\_ENABLE\_PLUGINS\ environment variable to control plugin activation, allowing users to enable or disable plugins at runtime without code changes.
packages/markitdown-mcp/src · high confidence
Behavioural changes
MarkItDown library release 0.1.7
The MarkItDown library has been updated to version 0.1.7. This release includes a variety of fixes and improvements, such as better handling of linked images in DOCX files, support for preserving base64-encoded images, and the addition of CSV to Markdown table conversion. The CLI now supports extension, MIME type, and charset hints, and the library has switched from puremagic to magika for file type detection. Additionally, the library now supports Azure Content Understanding and Document Intelligence for cloud-based conversion, and includes a plugin system for third-party converters.
packages/markitdown/src/markitdown · high confidence
Fixes
Improved math equation rendering in .docx files
The .docx converter now renders math equations as LaTeX, improving the accuracy of mathematical content in Word documents. This update includes a fix for invalid LaTeX macros for mu, nu, tau, and down-arrow in equation conversion, ensuring that special characters and symbols are correctly translated. Additionally, the OMML template bugs have been resolved, leading to more reliable conversion of mathematical expressions. Users will see enhanced support for complex mathematical notations in their converted documents.
_packages/markitdown/src/markitdown/converter\utils · high confidence
Test coverage
Added comprehensive test coverage for markitdown CLI, module, and converters; Added test fixtures for diverse document formats; Added test infrastructure for the MarkItDown MCP server; Added tests for the MarkItDown sample plugin's RTF conversion.
Dependencies
Initial release of MarkItDown and its plugins
This change introduces the core MarkItDown library along with three new packages: markitdown-mcp (an MCP server), markitdown-ocr (an OCR plugin), and markitdown-sample-plugin (a sample plugin). The main markitdown package now depends on magika for file type detection, defusedxml for secure XML parsing, and optional dependencies for various file formats (e.g., mammoth for DOCX, pdfminer.six and pdfplumber for PDFs). The mcp package is pinned to version 1.8.0, and mammoth is specified at version 1.11.0.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Score
- CAI 68 → 68 (-0.4)
- Rubric changed (rubric-2026.08.18 → rubric-2026.08.19) — scores are not directly comparable.
Lenses
- Code Health 80 → 80 (+0.3)
- Architecture 100 → 100 (+0.0)
- Maturity 72 → 70 (-1.6)
- Readiness 57 → 57 (+0.0)
- Security 88 → 87 (-0.5)
Resolved (2)
- Duplicated block (11 lines × 2) (packages/markitdown-ocr/src/markitdown_ocr/_pptx_converter_with_ocr.py)
- Off-boarding risk: anonymized user #1
New (6)
- Duplicated block (7 lines × 2) (packages/markitdown-ocr/src/markitdown_ocr/_pptx_converter_with_ocr.py)
- Hotspot: packages/markitdown/src/markitdown/main.py (packages/markitdown/src/markitdown/main.py)
- Hotspot: packages/markitdown/src/markitdown/_markitdown.py (packages/markitdown/src/markitdown/_markitdown.py)
- Medium IaC: CKV_DOCKER_2 (Dockerfile)
- Medium IaC: CKV_DOCKER_2 (packages/markitdown-mcp/Dockerfile)
- Off-boarding risk: anonymized user #1
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
microsoft/markitdown was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 5 August 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit fd239d5d2be43d9b68329730206b9312c7d5a388 — the exact code this score is about.
- Scored under rubric-2026.08.19 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer latest.