datalab-to/marker
61.7
Adequate · 18 September 2026
12.5k
lines of production code
Python
primary language
1
measurement over time
What this system is
Marker is a modular document conversion system that transforms PDFs and various other file formats into structured Markdown, HTML, JSON, or chunked text. It employs a server-based inference architecture to offload heavy layout and OCR tasks to external services, while utilizing a pluggable LLM layer for content correction and extraction. The system supports multiple output renderers and provides CLI, API, and web interfaces for batch or single-file processing.
How it got here
2023–2024 — v2.0 architecture and modular pipeline
22 changes.
The project underwent a major v2.0 overhaul, shifting to a server-based inference architecture and migrating its build system to Hatch. This period introduced a comprehensive modular framework with specialized builders, processors, and renderers, while expanding input support to include DOCX, EPUB, and other formats. Extensive test coverage was established for the new pipeline, alongside the integration of LLM-based processors for enhanced content correction and extraction.
2025–2026 — Interface expansion and LLM integration
7 changes.
This period focused on expanding user access to the PDF conversion engine by introducing dedicated CLI, REST API, and Streamlit interfaces, alongside a modular architecture for pluggable LLM providers. The work also included optimizing batch processing performance, adding comprehensive test coverage for configuration and service initialization, and establishing a reproducible benchmarking suite for quality and throughput evaluation.
Features
Add example for deploying Marker on Modal
Added a new example in the \examples\ directory that demonstrates how to deploy the Marker PDF conversion service using Modal. This includes a Python script (\marker\_modal\_deployment.py\) setting up a FastAPI-based web endpoint with GPU support and model caching, along with updated documentation in \examples/README.md\ explaining prerequisites, deployment commands, and usage via CLI or HTTP requests.
examples · high confidence
Added LaTeX to Markdown conversion utility
A new shell script (data/latex\_to\_md.sh) has been introduced to automate the conversion of LaTeX files into Markdown format. This tool processes .tex files from the 'latex' directory, converts them using Pandoc, cleans up non-breaking spaces and citation commands, and outputs the results to the 'references' directory. A corresponding .gitignore file has also been added to exclude generated LaTeX, PDF, and reference artifacts from version control.
data · high confidence
Dynamic CLI configuration via automatic config crawling
The CLI now automatically discovers and exposes configuration options for all Builders, Processors, Converters, Providers, and Renderers. By crawling class annotations and inheritance hierarchies, the system generates shared command-line flags (e.g., \--debug\) and class-specific flags (e.g., \--MarkdownRenderer\_output\_dir\) without manual registration. This allows users to override settings for any component directly from the command line or JSON config files, with the parser correctly handling defaults, types, and tri-state boolean logic.
marker/config · high confidence
Introduce modular LLM service providers for Marker
The \marker/services\ directory now provides a unified, pluggable architecture for Large Language Model integration. A new \BaseService\ abstract class defines the interface for inference, including retry logic, timeout handling, and image processing. Concrete implementations have been added for OpenAI, Azure OpenAI, Anthropic Claude, Google Gemini, Google Vertex AI, OpenRouter, and local Ollama models. This change allows users to select their preferred LLM provider via configuration, supporting structured output parsing and multimodal inputs across all supported backends.
marker/services · high confidence
Introduce structured document schema and block type registry
The \marker/schema\ package now provides a formalized data model for document processing. It defines a \BlockTypes\ enumeration covering all supported content elements (such as Text, Table, Equation, and ComplexRegion) and establishes a registry that maps these types to their corresponding Pydantic model classes. This structure enables consistent serialization, layout label mapping, and block resolution across the pipeline.
marker/schema · high confidence
Introduce structured schema groups for document elements
The document model now includes dedicated group classes (FigureGroup, TableGroup, ListGroup, PictureGroup, and PageGroup) that organize related content blocks. These groups enable more accurate HTML rendering by handling specific structures like lists with continuations and tables with captions, while the PageGroup manages page-level resources such as images and references.
marker/schema/groups · high confidence
Introduce structured text schema with rich formatting support
This change introduces the core data model for text content in the \marker/schema/text\ module, defining \Char\, \Span\, and \Line\ classes. It enables the system to preserve and render rich text formatting, including bold, italic, code, underline, highlighting, and specifically superscripts and subscripts. The implementation includes logic to handle hyphenated line breaks, merge lines while preserving structure, and convert text spans into HTML with support for inline and block math environments.
marker/schema/text · high confidence
LLM-based content correction and extraction processors
Added a suite of new LLM-powered processors in the marker/processors/llm module to improve document fidelity. These include LLMTableProcessor for correcting table structure and content, LLMEquationProcessor for generating accurate LaTeX math, LLMHandwritingProcessor for OCRing handwritten text, LLMFormProcessor for structuring form data, LLMImageDescriptionProcessor for captioning figures, LLMSectionHeaderProcessor for fixing heading hierarchy, LLMComplexRegionProcessor for general text correction, and LLMMathBlockProcessor for inline math correction. A meta-processor (LLMSimpleBlockMetaProcessor) orchestrates these tasks in parallel, and LLMPageCorrectionProcessor handles high-level page structure reordering and block type correction.
marker/processors/llm · high confidence
New CLI, API, and Streamlit entrypoints for PDF conversion
The marker package now includes dedicated scripts for running the conversion tool in different environments. A new Streamlit app (run via \marker/scripts/run\_streamlit\_app.py\) provides an interactive web interface for uploading PDFs/images and selecting output formats (Markdown, JSON, HTML, chunks) and modes (fast, balanced, auto). A FastAPI server script (\marker/scripts/server.py\) exposes a REST API with endpoints for converting PDFs via file path or direct upload, supporting the same configuration options. Additionally, new CLI entrypoints (\marker/scripts/convert.py\ and \convert\_single.py\) allow batch or single-file conversion from the command line, with support for multi-node sharding and worker configuration. These scripts centralize the user-facing interfaces for the underlying conversion logic.
marker/scripts · high confidence
New document structure processors for layout refinement
The marker/processors module now includes a suite of new processors that refine document structure and output quality. These include BlankPageProcessor to filter out blank pages, BlockRelabelProcessor to heuristically relabel blocks based on confidence, BlockquoteProcessor to detect and tag blockquotes, CodeProcessor to format code blocks, DebugProcessor to dump layout and PDF debug data, DocumentTOCProcessor to generate a table of contents, EquationProcessor to recognize equations and inline math via OCR, FootnoteProcessor to push footnotes to the bottom and assign superscripts, IgnoreTextProcessor to filter out common repetitive text like headers and footers, LineMergeProcessor to merge inline math lines, LineNumbersProcessor to ignore line numbers, ListProcessor to merge lists across pages and columns, MarginaliaProcessor to relabel running headers/footers, PageHeaderProcessor to move page headers to the top, ReferenceProcessor to add references to the document, and SectionHeaderProcessor to recognize section headers.
marker/processors · high confidence
New renderer architecture with configurable output formats
The document rendering system has been refactored into a modular renderer framework located in \marker/renderers\. This introduces a \BaseRenderer\ that supports configurable options for image extraction (including high/low resolution modes), page header/footer retention, and block ID injection. Specific renderers now handle distinct output formats: \HTMLRenderer\ produces paginated or flat HTML with optimized content resolution; \MarkdownRenderer\ supports configurable math delimiters, HTML table preservation, and chemical structure fencing; \JSONRenderer\ and \OCRJSONRenderer\ provide structured block-level data with polygon coordinates and optional character-level OCR details; and a new \ChunkRenderer\ flattens the document tree into individual block chunks for granular processing.
marker/renderers · high confidence
New reproducible benchmark harness for quality and throughput
A new benchmarking suite has been added to the \benchmarks/\ directory to provide reproducible quality and throughput measurements for the top-level README. The harness uses the olmOCR-bench dataset (1,403 PDFs) to measure quality via pass rates and reports deployment-relevant sustained throughput (pages/sec) rather than single-stream latency. It includes a main inference script for the Marker system (supporting balanced, fast, and accurate modes with optional LLM refinement and Surya server integration), a post-processing step to normalize output for accurate scoring, and a summarization tool to calculate Overall and Digital-only macro-averages. Additionally, it provides competitor scripts for Liteparse, Docling, and MinerU, allowing for fair, like-for-like throughput comparisons across different systems and concurrency models.
benchmarks · high confidence
Project initialization with Apache 2.0 licensing and pre-commit tooling
The repository has been initialized with a new project structure, establishing the Apache License 2.0 for the main codebase and the OpenRAIL-M license for the AI models. A Contributor License Agreement (CLA) has been added to define intellectual property rights for contributions. Development workflow is standardized with a pre-commit configuration that enforces code quality using Ruff (v0.9.10) for linting and formatting. The project also includes entry-point scripts for the CLI, single-file conversion, Streamlit app, and server, along with a pytest configuration for test execution.
(repo-wide) · high confidence
Support for DOCX, EPUB, HTML, PowerPoint, and Spreadsheet files
The document processing pipeline now accepts DOCX, EPUB, HTML, PowerPoint (PPTX), and spreadsheet (XLSX) files in addition to PDFs and images. These new formats are handled by dedicated providers that convert the source files into PDFs using WeasyPrint, allowing the existing PDF extraction and OCR engine to process them uniformly. A file-type registry automatically detects the input format and routes it to the appropriate provider.
marker/providers · high confidence
Architecture
Architectural shift to server-based inference and simplified configuration
Marker now offloads heavy model inference (layout, OCR, and error detection) to shared external servers, allowing multiple worker processes to share a single GPU without each loading a model copy. This change removes the local Nougat equation model, the LayoutLMv3 segmentation model, and the Tesseract OCR fallback, replacing them with a thin client architecture. The configuration is simplified: the \TORCH\_DEVICE\ setting now only applies to lightweight local models (like the OCR error detector), while the main VLM runs on the inference server. Additionally, the output encoding is now explicitly configurable via \OUTPUT\_ENCODING\ (defaulting to UTF-8) and images are converted to RGB before saving as JPEG.
marker · high confidence
Behavioural changes
Introduce modular converter architecture with specialized PDF, OCR, and table extraction modes
The \marker/converters\ package has been restructured into a modular system centered on a \BaseConverter\ that handles dependency resolution and processor initialization, including a new meta-processor for batching LLM requests. This introduces three distinct conversion paths: the standard \PdfConverter\ (which now defaults to 'balanced' mode on GPU and 'fast' mode on CPU/MPS), an \OCRConverter\ for forced OCR workflows, and a new \TableConverter\ that isolates table/form extraction by disabling general OCR and filtering document structure. Users can now choose the appropriate converter for their specific data type, with the table converter offering a dedicated pipeline for complex table and form structures.
marker/converters · high confidence
New block schema and HTML rendering logic for document elements
The \marker/schema/blocks\ module now defines a comprehensive set of block types (including Text, Equation, Table, ListItem, Footnote, and ComplexRegion) with dedicated HTML assembly logic. This change introduces structured HTML output for document components, supporting features such as nested blockquotes, list indentation, table cell spanning, and configurable inclusion of page headers/footers via \block\_config\ options like \keep\_pageheader\_in\_output\ and \add\_block\_ids\.
marker/schema/blocks · high confidence
Optimized batch concurrency and blank-content detection
The marker utility layer now includes a new batch processing module that calculates worker counts based on physical CPU cores and server parallelism settings, allowing for more efficient resource utilization during conversion tasks. Additionally, the image processing logic has been refactored to use a faster, bit-for-bit equivalent method for detecting blank images by leveraging adaptive thresholding directly, eliminating expensive connected-components operations. This change also introduces a specific fix to treat polygons with identical coordinates as blank, ensuring consistent filtering of empty content.
marker/utils · high confidence
Refactored document processing pipeline with new builder architecture
The marker/builders module has been completely rewritten to introduce a modular builder architecture (BaseBuilder, DocumentBuilder, LayoutBuilder, LineBuilder, OcrBuilder, StructureBuilder). This refactoring changes how documents are processed: layout detection now supports a 'fast' mode using a lightweight detector for CPU/MPS devices, while 'balanced' mode uses a VLM model on GPU. Reading order is now determined by the PDF's character stream on text-rich pages to improve accuracy, and high-resolution images are rendered lazily only for pages that require them (e.g., those with tables or equations). OCR is now performed as a full-page request for pages with bad embedded text, with a fallback to block-level OCR for specific garbled regions, and structure grouping (lists, captions) is handled by a dedicated StructureBuilder.
marker/builders · high confidence
Test coverage
Added comprehensive test coverage for PDF converters and modes; Added comprehensive test suite for document builders; Added test coverage for document processors; Added test coverage for renderer components; Added test for list grouping structure validation; Added tests for LLM service initialization; Added tests for document, image, and PDF providers; Added unit tests for CLI configuration parsing; Initial test infrastructure for PDF conversion pipeline.
Dependencies
Migration from Poetry to Hatch and major dependency overhaul
The project has switched its build system from Poetry to Hatch (pyproject.toml), removing the legacy poetry.lock and requirements.txt files. This update bumps the package version to 2.0.0 and significantly updates dependencies: it replaces PyMuPDF with pdftext, swaps nougat-ocr for surya-ocr, and updates core libraries like Pillow, Pydantic, and transformers to newer major versions. It also adds support for new document formats via optional dependencies (mammoth, openpyxl, python-pptx, ebooklib, weasyprint) and introduces new CLI entry points for server and single-file conversion.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Baseline
- First survey — no prior run to compare against. CAI 62.
Lenses
- Code Health 87
- Architecture 98
- Maturity 60
- Readiness 55
- Security 63
Changes since last survey
- 300 commits — 232 feature/other, 68 fixes
By area
- (repo) — 72 commits
- signatures/version1 — 50 commits
- marker/processors — 39 commits
- (root) — 37 commits
- marker/builders — 22 commits
- marker/services — 14 commits
- marker/renderers — 11 commits
- marker/schema — 7 commits
- .github/workflows — 5 commits
- examples/README_MODAL.md — 5 commits
- marker/config — 4 commits
- marker/scripts — 4 commits
- tests/builders — 4 commits
- marker/converters — 3 commits
- marker/models.py — 3 commits
- tests/processors — 3 commits
- marker/extractors — 2 commits
- tests/converters — 2 commits
- .github/ISSUE_TEMPLATE — 1 commit
- benchmarks/overall — 1 commit
Notable commits
- fix: Bugfix
- fix: Bump to new surya with latex hotfix
- fix: CI fixes
- fix: CI fixes
- fix: Copy + file name fix
- fix: Error fixes
- fix: Filter fix
- fix: Fix CLA allowlist
- fix: Fix ToC cells
- fix: Fix asserts
- fix: Fix attributes
- fix: Fix blocks
- fix: Fix charts
- fix: Fix crash on layout-detected groups with no children
- fix: Fix issues with OCRConverter
- fix: Fix license
- fix: Fix math rendering - Less newlines in output
- fix: Fix merge conflicts
- fix: Fix misc bugs
- fix: Fix numpy issue
- …and 280 more
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
datalab-to/marker was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 18 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 8a1d2344de25d7ec4c5209133aed7af565874ff0 — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-5d04157a340d.