Skip to content
CAI
Software that uses CAICheck a score

microsoft/unilm

58.7

Adequate · 19 September 2026

813.6k

lines of production code

Python

primary language

1

measurement over time

CAI band scale
CAI lens gauges

What this system is

This system is a comprehensive research repository for multimodal machine learning, providing PyTorch implementations of state-of-the-art models for natural language processing, computer vision, and speech recognition. It enables users to pre-train, fine-tune, and evaluate architectures such as transformers, diffusion models, and document understanding systems across diverse tasks including translation, OCR, and dense retrieval. The codebase also includes specialized toolkits for optimizing inference speed, handling long-context sequences, and performing domain adaptation.

How it got here

2020–2021 — multimodal and multilingual model expansion

23 changes.

This period focused on releasing a diverse suite of pre-trained models and fine-tuning toolkits for text, document, image, and speech modalities. Key additions included MiniLM, LayoutLM, BEiT, WavLM, and TrOCR, alongside specialized toolkits for sequence-to-sequence tasks, domain adaptation, and cross-lingual transfer. The work established comprehensive infrastructure for multimodal understanding and generation across various benchmarks.

2022 — multimodal and speech model releases

22 changes.

This period focused on releasing a broad suite of new models and frameworks for document AI, vision, and speech processing, including EdgeFormer, DiT, LayoutLMv3, BEiT-2/3, SpeechLM, and SpeechT5. The work involved implementing core architectures, training pipelines, and evaluation scripts for tasks such as object detection, text detection, dense retrieval, and audio tokenization. These releases expanded the project's capabilities in multimodal understanding and efficient sequence generation.

2023–2026 — multimodal and reasoning model releases

19 changes.

This period focused on releasing and expanding several multimodal and reasoning models, including TextDiffuser, Kosmos-2, and LatentLM, alongside their associated training and evaluation tooling. It also introduced specialized architectures for long-context inference and efficient attention, such as YOCO, ReSA, and Diff Transformer, while establishing new benchmarks for mathematical reasoning and text rendering.

Features

Add BEATs audio pre-training and tokenization models

This release introduces the official PyTorch implementation of BEATs (Audio Pre-Training with Acoustic Tokenizers), including the core model architectures in BEATs.py and Tokenizers.py, supporting transformer-based feature extraction and acoustic tokenization via vector quantization. The addition includes the necessary backbone and module components (backbone.py, modules.py, quantizer.py) to load pre-trained and fine-tuned models for tasks such as audio representation learning and classification, along with updated documentation and download links for the model checkpoints.

beats · high confidence

Add BEiT semantic segmentation support for ADE20K

Introduces a new semantic segmentation module for the ADE20K dataset using the BEiT transformer backbone. This includes the BEiT backbone implementation, UperNet model configuration, and training/evaluation scripts for both base (12-layer) and large (24-layer) variants. The feature supports fine-tuning from ImageNet pre-trained weights and includes multi-scale testing configurations for improved segmentation accuracy.

_beit/semantic\segmentation · high confidence

Add BEiT v2 semantic segmentation support for ADE20K

This change introduces a complete implementation for fine-tuning BEiT v2 (base and large variants) on the ADE20K semantic segmentation dataset using the UperNet architecture. It adds the BEiT backbone code, custom training utilities including a layer-decay optimizer constructor and an Apex-based AMP runner, and configuration files for 160k and 320k iteration schedules with 512x512 and 640x640 crop sizes. Users can now train and evaluate BEiT v2 models on ADE20K using the provided mmsegmentation-compatible configs and pretrained checkpoints.

_beit2/semantic\segmentation · high confidence

Add BEiT-3 vision and vision-language model implementation

Introduces the official PyTorch implementation and pretrained models for BEiT-3, a vision and vision-language pretraining model. This addition includes the core model architecture (leveraging the torchscale library), dataset handling for image-text pairs, and specific fine-tuning engines for downstream tasks such as image classification, visual question answering (VQAv2), visual reasoning (NLVR2), image captioning, and image-text retrieval. The release also provides pretrained checkpoints for base and large model sizes, along with documentation and setup instructions.

beit3 · high confidence

Add DiT image classification fine-tuning support for RVL-CDIP

This change introduces a new classification module within the DiT project, enabling users to fine-tune and evaluate the Document Image Transformer on the RVL-CDIP dataset. The addition includes the full training and evaluation pipeline (\run\_class\_finetuning.py\, \engine\_for\_finetuning.py\), model definitions for fine-tuning (\modeling\_finetune.py\), and necessary data handling utilities (\datasets.py\, \dataset\_folder.py\). It also provides DeepSpeed configuration for distributed training and detailed usage instructions in the README.

dit/classification · high confidence

Add DiT-based text detection training and evaluation scripts

Introduces a new text detection module that leverages the DiT (Document Image Transformer) backbone within a Mask R-CNN framework. The change adds a training script (\train\_net.py\) capable of registering the FUNSD dataset, initializing the DiT configuration, and launching distributed training or evaluation using Detectron2. It also includes a README documenting fine-tuned model weights, data preparation steps, and usage examples for training and evaluating text detection models on FUNSD.

_dit/text\detection · high confidence

Add LayoutLMv3 object detection training examples and infrastructure

This release adds a complete set of examples and supporting code for training LayoutLMv3 on object detection tasks. It includes a YAML configuration file for a cascade LayoutLMv3 model on the PubLayNet dataset, along with a custom Detectron2 backbone implementation that integrates the LayoutLMv3 transformer. The package also provides dataset conversion scripts to transform ICDAR data into COCO format, an adaptive binarization utility for image preprocessing, and a custom training loop and checkpointer to handle model weight loading and position embedding interpolation.

layoutlmv3/examples · high confidence

Add MARIOEval evaluation and generation scripts

The textdiffuser/eval directory now includes scripts to generate images and evaluate them using the MARIOEval benchmark. MARIOEval\_generate.py supports image generation via Stable Diffusion, ControlNet, and DeepFloyd pipelines, while MARIOEval\_evaluate.py computes CLIPScore and FID metrics. Additionally, ocr\_eval.py provides OCR-based evaluation (precision, recall, accuracy) for text rendering, and supporting modules for CLIPScore and FID calculations are included.

textdiffuser/eval · high confidence

Add PFPO (ICLR 2025) analysis and pseudo-test-case pipelines

Introduces a new suite of scripts under \PFPO/scripts/apps/\ to support the PFPO methodology. This includes \analyze/\ tools for computing output frequencies and generating visualization histograms, \prm/\ scripts for constructing process reward model samples and sampling code steps, and \pseudo\_test\_cases/\ utilities for generating, cleaning, and executing synthetic test cases to build DPO preference pairs.

(repo-wide) · high confidence

Add SpeechLM speech pre-training and fine-tuning code

Added the SpeechLM module, including the main model implementation (SpeechLM.py), supporting modules (modules.py), and comprehensive documentation (README.md). This addition provides pre-trained and fine-tuned models for speech tasks, along with scripts and instructions for feature extraction, automatic speech recognition (ASR) on LibriSpeech, and speech translation (ST) on CoVoST-2. The codebase includes a submodule for fairseq and requires specific dependencies like sacrebleu for setup.

speechlm · high confidence

Add SpeechT5 unified speech and text processing implementation

Introduces the SpeechT5 module, providing a unified encoder-decoder architecture for spoken language processing tasks. This addition includes the core model components, data handling via a MultitaskDataset, and specific loss functions (criterions) for pre-training, speech-to-text (ASR), and text-to-speech (TTS) tasks. It also provides a script for generating speech outputs and integrates with the fairseq framework via a submodule, enabling users to load pre-trained models and perform inference or fine-tuning for speech and text modalities.

speecht5 · high confidence

Add TrOCR text recognition module with augmentation and inference support

Introduces the TrOCR (Transformer-based Optical Character Recognition) module, providing an end-to-end text recognition approach using pre-trained image and text Transformers. This addition includes the core model implementations (ViTTRModel, TrOCRModel), a custom GPT2BPEEnhancedSpace tokenizer, and data handling utilities for datasets like IAM and SROIE. It also provides a comprehensive suite of image augmentation techniques (blur, noise, geometry, weather effects) to improve model robustness, along with scripts for fine-tuning, evaluation, and inference.

trocr · high confidence

Add VLMo vision-language pre-training implementation

Introduces the VLMo (Unified Vision-Language Pre-Training with Mixture-of-Modality-Experts) module, providing the PyTorch implementation, configuration, and data handling for pre-training and fine-tuning. This includes dataset preparation scripts and data modules for GCC, SBU, Visual Genome, COCO, Flickr30K, VQAv2, NLVR2, and WikiBK, along with a multi-task data loader and training entry point using PyTorch Lightning.

vlmo · high confidence

Add ViT-based text detection backbone and evaluation utilities

Introduces a new text detection module under \dit/text\_detection/ditod\ that integrates Multi-Path Vision Transformer (ViT) backbones (including BEiT, DiT, DeiT, and MAE variants) with a Feature Pyramid Network (FPN) via the Detectron2 registry. This change also adds supporting infrastructure for model training and evaluation, including a custom dataset mapper, checkpointer, trainer, and a FUNSD-specific evaluator, alongside utility modules for configuration management, image conversion, and standard ICDAR detection metrics (IoU, DetEval, ICDAR2013, and MTWI2018).

_dit/text\detection/ditod · high confidence

Add XDoc fine-tuning code and documentation

This change introduces the fine-tuning infrastructure for the XDoc model, including a comprehensive README with instructions for SQuAD, FUNSD, and WebSRC tasks, alongside the \layoutlmft\ library. The library provides custom dataset builders for FUNSD and XFUN, data collators that handle image and bounding-box inputs, and model implementations for LayoutLMv2 and LayoutXLM, enabling users to fine-tune XDoc on document understanding benchmarks.

xdoc · high confidence

Add angle and quadrilateral text rendering extensions

New training and inference scripts have been added to the TextDiffuser-2 extensions directory to support rendering text with specific geometric layouts: angles and quadrilaterals. This includes Python scripts for training (\train\_textdiffuser2\_t2i\_full\_angle.py\, \train\_textdiffuser2\_t2i\_full\_quadrilateral.py\) and inference (\inference\_textdiffuser2\_t2i\_full\_angle.py\, \inference\_textdiffuser2\_t2i\_full\_quadrilateral.py\), along with corresponding shell launch scripts and template input files (\angle\_template\_file.txt\, \quadrilateral\_template\_file.txt\) that define the coordinate-based layout for these shapes.

textdiffuser-2/extensions · high confidence

Add deprecated LayoutLM document understanding module

This change introduces the \layoutlm/deprecated\ directory, containing the original LayoutLM implementation for multimodal document understanding (text, layout, and image). It includes the core model classes (\LayoutlmConfig\, \LayoutlmForSequenceClassification\, \LayoutlmForTokenClassification\), data processors for the FUNSD and RVL-CDIP datasets, and example scripts for fine-tuning on sequence labeling and classification tasks. The module also adds standard Python development configuration files (\.flake8\, \.isort.cfg\, \.pre-commit-config.yaml\) to enforce code style and linting.

layoutlm/deprecated · high confidence

Add object detection inference and training scripts for DiT

This change introduces the \dit/object\_detection\ module, providing scripts to run inference and train Mask R-CNN and Cascade Mask R-CNN models using the DiT backbone on PubLayNet and ICDAR 2019 cTDaR datasets. Users can now perform document layout analysis via the new \inference.py\ script, which supports visualization of detected objects, or fine-tune the models using \train\_net.py\ with configurable datasets and GPU settings. The module also includes utilities for data preparation, such as \convert\_to\_coco\_format.py\ for converting ICDAR annotations to COCO format and \adaptive\_binarize.py\ for image preprocessing.

_dit/object\detection · high confidence

Add relation extraction decoder module

The library now includes a new relation extraction (RE) decoder component within the modules/decoders package. This addition introduces a BiaffineAttention operator and a REDecoder class that enable binary relation classification by combining entity embeddings with hidden states, allowing users to extract structured relationships between entities in their documents.

layoutlmft/layoutlmft/modules/decoders · high confidence

Add support for MiniLM and UniLM models in s2s-ft

The s2s-ft module now supports fine-tuning MiniLM and UniLM models in addition to BERT and RoBERTa. This includes new configuration classes (MinilmConfig, UnilmConfig), tokenizers (MinilmTokenizer, UnilmTokenizer), and state-dict converters to handle model-specific weight formats (e.g., UniLM's QKV linear layers). The modeling code has been updated to load these models, resize position embeddings if necessary, and apply source/target type IDs for sequence-to-sequence tasks.

_s2s-ft/s2s\ft · high confidence

Added DynamicConv, LightConv, and Product Quantization modules

This change introduces new neural network layer implementations and quantization utilities to the library. It adds \DynamicconvLayer\ and \LightconvLayer\, which provide optimized CUDA-accelerated convolution operations with support for incremental decoding during inference. Additionally, it includes a Product Quantization (PQ) subsystem under \modules/quantization/pq\, featuring an Expectation-Maximization (EM) algorithm for clustering weights and specific quantized module classes (\PQConv2d\, \PQEmbedding\, \PQLinear\) that allow for model compression by storing centroids and assignments rather than full-precision weights.

_edgelm/fairseq/modules/dynamicconv\_layer, edgelm/fairseq/modules/lightconv\layer, edgelm/fairseq/modules/quantization · high confidence

Added Mario LAION example dataset and processing scripts

The \textdiffuser/data\ directory now includes a new example dataset under \mario-laion-example\, containing structured entries with captions, OCR text with bounding box coordinates, and metadata JSON files for specific image keys. Additionally, a Python script \mario-laion-unzip.py\ has been added to facilitate the multiprocessed extraction of LAION OCR data archives, and a Jupyter notebook \visualize\_charseg.ipynb\ is provided for visualizing character segmentation results using the new data.

textdiffuser/data · high confidence

Added SpeechLM dataset configurations and assets for CommonVoice and LibriSpeech

This change introduces dataset configuration files and supporting assets for the SpeechLM project, specifically for English-to-German translation using CommonVoice v4 and speech recognition using LibriSpeech. For CommonVoice, it adds YAML configs (\config\_base\_ende.yaml\, \config\_large\_ende.yaml\) defining SentencePiece tokenization, audio input settings (16kHz, mono), and a 100-sample development manifest (\dev-sample100\_st\_en\_de\_local.tsv\) linking English audio files to German translations. It also includes the corresponding SentencePiece vocabulary files (\spm\_char\_st\_en\_de.txt\, \spm\_char\_st\_en\_de.vocab\). For LibriSpeech, it adds configuration and dictionary files for both phone-level and hidden-unit models, including character-level (\dict.ltr.txt\) and phone-level (\dict.phn.txt\) vocabularies, and a 100-sample training manifest (\train\_sample100.ltr\) with phoneme-transcribed text.

speechlm/dataset · high confidence

Added TextDiffuser demo assets and example cases

This update adds demonstration assets for the TextDiffuser model, including Python implementation files (modeling utilities, DDPMScheduler, and UNet2DConditionModel) and a collection of text prompt examples for text inpainting, text-to-image, and text-to-image-with-template scenarios. These files provide the necessary code structure and input cases to run and test the TextDiffuser demo.

textdiffuser/assets · high confidence

Added TextDiffuser model components for keyword layout prediction

This change introduces the core model architecture for TextDiffuser, a system that predicts the spatial layout of keywords within user prompts. It adds a \LayoutTransformer\ and \TextConditioner\ (leveraging CLIP) to determine where text should be placed, along with a UNet-based \TextSegmenter\ for segmentation tasks. These components enable the model to generate structured layouts for text-based image generation.

textdiffuser/model, textdiffuser/textdiffuser-ckpt · high confidence

Added fairseq documentation and project metadata

The infoxlm/fairseq directory now includes the Sphinx documentation source files (RST), configuration, and static assets, providing users with guides on getting started, command-line tools, and extending the toolkit. This change also adds standard project metadata files, including the MIT License, Code of Conduct, and Contributing guidelines, to support community engagement and clarify usage rights.

(repo-wide) · high confidence

Adds Aggressive Decoding (GAD) plugins for non-autoregressive sequence generation

This change introduces a new set of Fairseq plugins under \decoding/GAD/block\_plugins\ that enable Aggressive Decoding (GAD) for sequence-to-sequence tasks. The update adds a \BlockNAT\ model (\block\ architecture) which implements a non-autoregressive transformer with specific forward-pass logic for GAD-style iterative refinement. It also registers a \glat\_loss\ criterion for training with label smoothing and dual imitation, and a \translation\_lev\_modified\ task that supports various noise injection strategies (random delete, random mask, block mask) to facilitate the training of these non-autoregressive models.

decoding · high confidence

Initial release of AdaLM domain adaptation toolkit

This change introduces the AdaLM repository, providing code to fine-tune pre-trained language models for specific domains (biomedical and computer science) and to generate incremental domain-specific vocabularies. The \finetune\ directory contains scripts for sequence classification (\run\_classifier.py\) and named entity recognition (\run\_ner.py\, \run\_pico.py\), supporting models like BERT, RoBERTa, and DistilBERT via the Hugging Face \transformers\ library. The \incr\_bpe\ directory includes tools to build domain-specific subword vocabularies by extending base BERT vocabularies, utilizing a simplified version of the \tensor2tensor\ library's BPE implementation. Documentation and pre-trained model links are also included.

adalm · high confidence

Initial release of BEiT image transformer implementation

This change introduces the BEiT (BERT Pre-Training of Image Transformers) codebase, providing the official PyTorch implementation for self-supervised pre-training and fine-tuning of image transformers. The release includes the core model architecture, training engines for both pre-training and fine-tuning, and data augmentation pipelines that support masked image modeling. It also integrates a DALL-E discrete VAE tokenizer for generating visual tokens and provides documentation and scripts for fine-tuning on ImageNet-1k and semantic segmentation on ADE20K.

beit · high confidence

Initial release of EdgeFormer documentation and project scaffolding

This change introduces the \edgelm\ directory, establishing the project structure for EdgeFormer (a parameter-efficient transformer for on-device seq2seq generation). It adds the core project metadata files, including the MIT License, Code of Conduct, and Contributing guidelines. It also provides a comprehensive README detailing pretrained model checkpoints, benchmark performance on tasks like CoNLL-14, XSUM, and SQuAD-NQG, and setup/fine-tuning instructions using fairseq. Additionally, it includes a full Sphinx documentation suite (\docs/\) covering command-line tools, library references (models, tasks, optimizers), and tutorials for extending the framework.

edgelm · high confidence

Initial release of Kosmos-2 grounding toolkit and documentation

This change introduces the initial codebase and documentation for Kosmos-2, a multimodal large language model grounded to the world. It includes a comprehensive README detailing model checkpoints, setup instructions (Docker and Conda), and evaluation metrics for tasks like phrase grounding and visual question answering. The release provides data preparation scripts for the GRIT dataset, a Gradio-based interactive demo for local hosting, and utility modules for decoding bounding box coordinates from model captions and visualizing them on images.

kosmos-2 · high confidence

Initial release of Kosmos-2.5 multimodal literate model

This change introduces the Kosmos-2.5 repository, a multimodal model designed for machine reading of text-intensive images. The release includes the inference code (\inference.py\), model architecture definitions (\kosmos2\_5/models/\), and task handling (\kosmos2\_5/tasks/\) required to run the model. Users can now perform two primary tasks: text recognition (OCR) that generates spatially-aware text blocks, and image-to-markdown conversion that captures document structure and styles. The implementation relies on Fairseq for the language model backbone and integrates with Hugging Face Transformers for image processing, supporting GPU acceleration via Flash Attention2.

kosmos-2.5 · high confidence

Initial release of LayoutLMv3 document AI model

This change introduces the LayoutLMv3 model, a multimodal transformer for Document AI that uses unified text and image masking for pre-training. The release includes the \layoutlmv3-base\, \layoutlmv3-large\, and \layoutlmv3-base-chinese\ pre-trained models, along with setup scripts and documentation for fine-tuning on tasks such as form understanding (FUNSD, XFUND) and document layout analysis (PubLayNet).

layoutlmv3 · high confidence

Initial release of LayoutReader for reading order detection

This change introduces the LayoutReader component, a seq2seq model that captures text and layout information to predict reading order, significantly improving OCR engine results. The release includes the core training script (run\_seq2seq.py) and decoding script (decode\_seq2seq.py) which leverage the LayoutLM architecture and the s2s-ft library, along with a setup.py defining the s2s-ft package dependencies (including transformers \<= 2.10.0). A comprehensive README provides installation instructions, usage examples for training and decoding on the ReadingBank dataset, and performance benchmarks.

layoutreader · high confidence

Initial release of LayoutXLM and XFUN dataset support

This change introduces the LayoutXLM model and its associated tokenizer, extending the existing LayoutLMv2 architecture to support multilingual document understanding. It also adds the XFUN dataset loader for cross-lingual form understanding, along with the necessary data collators, evaluation metrics for relation extraction, and utility functions for bounding box normalization and image loading.

layoutlmft/layoutlmft · high confidence

Initial release of MarkupLM pre-trained models and fine-tuning code

This change introduces the MarkupLM component, a multimodal pre-training method for text and markup language designed for visually-rich document understanding tasks such as webpage QA and information extraction. The release includes the \setup.py\ package definition for \markuplmft\ (version 0.1) and a comprehensive \README.md\ detailing installation, fine-tuning procedures for datasets like WebSRC and SWDE, and links to pre-trained models (MarkupLM-Base and MarkupLM-Large) available on Hugging Face.

markuplm · high confidence

Initial release of PFPO source code and configuration

This change introduces the source code and configuration files for Preference Optimization for Reasoning with Pseudo Feedback (PFPO), a method for improving reasoning in large language models. The release includes a README detailing experimental results on mathematical reasoning (MATH, GSM8K) and coding tasks (LiveCodeBench, HumanEval, MBPP), along with Hydra configuration files for running inference and training pipelines using models like Mathstral-7B and Deepseek-Coder-7B-v1.5 via vLLM.

PFPO · high confidence

Initial release of TextDiffuser 2 assets and training data

This change introduces the foundational components for TextDiffuser 2, including a custom attention processor implementation in \assets/attention\_processor.py\ that supports advanced features like LoRA and xFormers optimization. It also adds a \reference\_requirements.txt\ file pinning the environment dependencies (such as \diffusers==0.24.0\, \torch\, and \gradio==3.50.2\) and provides a 5,000-sample dataset (\data/layout\_planner\_data\_5k.json\) for training the model to plan visual text layouts, along with a utility script to visualize these layout plans.

textdiffuser-2/assets, textdiffuser-2/data · high confidence

Initial release of TextDiffuser-2 codebase and demos

This change introduces the complete TextDiffuser-2 repository, including training scripts for the layout planner (using FastChat/Vicuna) and diffusion models (full-parameter and LoRA fine-tuning), as well as inference scripts and Gradio-based interactive demos for text-to-image generation and text inpainting. The codebase supports text rendering with layout planning via language models, handles coordinate-based text positioning, and includes specific support for angle/quadrilateral text layouts and inpainting tasks.

textdiffuser-2 · high confidence

Initial release of s2s-ft sequence-to-sequence fine-tuning toolkit

The s2s-ft package is introduced as a PyTorch-based toolkit for fine-tuning pre-trained Transformer models (including BERT, RoBERTa, XLM-RoBERTa, Electra, UniLM, and MiniLM) for sequence-to-sequence tasks. This release provides the core training script (run\_seq2seq.py) and decoding utilities (decode\_seq2seq.py), along with specialized tokenizers and model configurations for UniLM and MiniLM. It also includes a comprehensive suite of evaluation scripts for standard summarization benchmarks (XSum, CNN/Daily Mail, Gigaword) and a utility for generating sequences from beam search traces, enabling users to fine-tune and evaluate models on these datasets out of the box.

s2s-ft · high confidence

Initial release of s2s\_ft module for LayoutReader

This change introduces the s2s\_ft (sequence-to-sequence fine-tuning) module to the LayoutReader package, providing the core infrastructure for loading and fine-tuning transformer models. It adds configuration classes for BERT, MiniLM, and UniLM models, along with tokenizers and model implementations that support pre-trained weights from Azure Blob storage and Hugging Face. The module includes utilities for converting state dictionaries between different model formats (e.g., RoBERTa to BERT, LayoutLM to BERT) and data loading pipelines for sequence-to-sequence tasks, enabling the layout reader to leverage these specific model architectures for document understanding.

_layoutreader/s2s\ft · high confidence

Initial release of the layoutlmft document understanding toolkit

This change introduces the \layoutlmft\ package, a multimodal fine-tuning toolkit for document understanding that supports LayoutLM, LayoutLMv2, and LayoutXLM models. It provides example scripts for fine-tuning on the FUNSD dataset for key-value extraction (\run\_funsd.py\) and the XFUN dataset for relation extraction (\run\_xfun\_re.py\) and sequence entity recognition (\run\_xfun\_ser.py\), along with the necessary project configuration files like \setup.py\ and \Makefile\.

layoutlmft · high confidence

Initial release of xTune cross-lingual fine-tuning scripts and documentation

Adds the xTune repository, providing a two-stage training process for cross-lingual fine-tuning of XLM-Roberta models on XTREME datasets (XNLI, PANX, PAWS-X, UDPOS, MLQA, TyDiQA, XQuAD). The change includes a README with setup instructions, shell scripts for downloading datasets and models, and specific training configurations for both 'cross-lingual-transfer' and 'translate-train-all' settings.

xtune · high confidence

Initial support for Vision Transformer (ViT) backbones in object detection

This change introduces a new \ditod\ module for the object detection pipeline, enabling the use of Vision Transformer backbones (such as BEiT, DIT, DeiT, and MAE) integrated with Feature Pyramid Networks (FPN). It adds configuration options for these models, a custom dataset mapper for DETR-style training, and specialized evaluation tools for table detection (ICDAR). Additionally, it includes a custom checkpointer to handle position embedding interpolation when loading pre-trained weights into different input resolutions.

_dit/object\detection/ditod · high confidence

Introduce BEiT v2 implementation with vector-quantized visual tokenizers

Adds the official PyTorch codebase for BEiT v2, a masked image modeling approach that uses vector-quantized visual tokenizers (VQ-KD) instead of discrete pixel tokens. This release includes the core training engines for pretraining, fine-tuning, and VQ-KD tokenizer training, along with dataset utilities and data augmentation pipelines specific to the v2 architecture. Documentation and configuration examples are provided for pretraining on ImageNet-1k and fine-tuning for image classification and semantic segmentation.

beit2 · high confidence

Introduce DeltaLM encoder-decoder model with Fairseq integration

Adds the DeltaLM model, an encoder-decoder architecture for multilingual language generation and translation, built as a submodular dependency on Fairseq. The implementation includes the core model definitions in deltalm/models/deltalm.py, which register with Fairseq's model registry and support loading pretrained checkpoints for fine-tuning. Entry-point scripts (train.py, generate.py, preprocess.py, interactive.py) are provided to expose Fairseq's CLI commands for the new architecture, alongside example shell scripts for data preparation, tokenization, training, and evaluation on the IWSLT14 German-English dataset.

deltalm · high confidence

Introduce Diff Transformer attention modules with FlashAttention support

This change adds a new Diff Transformer implementation to the repository, providing both standard PyTorch and FlashAttention-optimized versions of multi-head differential attention. The codebase includes \MultiheadDiffAttn\ for baseline comparison, \MultiheadFlashDiff1\ for libraries supporting different Q/K/V dimensions (like xformers or customized flash-attention), and \MultiheadFlashDiff2\ for standard flash-attention. It also introduces a Triton-based rotary embedding kernel (\kernel/rotary.py\) and a fallback RMSNorm implementation (\rms\_norm.py\), along with an example script to compare parameter counts between the differential and conventional attention mechanisms.

Diff-Transformer · high confidence

Introduce Differential Transformer V2 (DIFF V2) with FlashAttention integration

Added a new \MultiheadFlashDiffV2\ attention module that implements the Differential Transformer V2 architecture. This change improves inference efficiency by allowing the use of standard FlashAttention without custom kernels, enhances training stability by removing per-head RMSNorm, and simplifies parameterization by using token-specific, head-wise projected lambda values instead of a globally shared scalar. The implementation is provided in \multihead\_flashdiffv2.py\ alongside updated documentation.

Diff-Transformer/Diff-Transformer-V2 · high confidence

Introduce LatentLM multimodal latent language modeling toolkit

Adds the LatentLM component, providing an official PyTorch implementation for multimodal latent language modeling with next-token diffusion. This includes core model architectures (DiT, Transformer), supporting modules (RMSNorm, EMA, attention kernels), and evaluation utilities for computing FID and Inception Score metrics, along with scripts for inference speed benchmarking and fidelity evaluation.

LatentLM · high confidence

Introduce Rectified Sparse Attention (ReSA) for efficient long-context inference

Adds the ReSA module, which improves sparse decoding by periodically refreshing the KV cache to prevent error accumulation and maintain generation quality. This change introduces a new \KVManager\ for block-sparse attention, custom Triton kernels for sparse decoding and attention with KV cache, and integration into the model's attention layer to enable up to 2.42x speedup at 256K context length while preserving near-lossless performance.

ReSA · high confidence

Introduce SpeechLM pre-training, fine-tuning, and decoding configurations

This change adds the configuration files for the SpeechLM model, including Hydra configs for pre-training (base and large variants), fine-tuning (100h and 960h subsets), and inference decoding (fairseq LM, KenLM, and Viterbi). These configs define the task, model architecture, optimization, and dataset settings for the SpeechLM pipeline.

speechlm/speechlm · high confidence

Introduce WavLM self-supervised speech model

Adds the WavLM model implementation and documentation, providing pre-trained checkpoints (Base, Base+, Large) and code for feature extraction and fine-tuning on speech tasks such as speaker verification, separation, diarization, and recognition.

wavlm · high confidence

Introduce YOCO decoder-decoder architecture with long-context evaluation capabilities

This change adds the YOCO (You Only Cache Once) implementation, a decoder-decoder architecture for large language models designed to improve long-context retrieval. The update includes the core model components, such as the GateRetention mechanism, cross-attention, and optimized kernels for SwiGLU and rotary embeddings. It also introduces specific evaluation criteria for the 'Needle in a Haystack' and multi-needle tasks, along with scripts for training and harness-based evaluation, enabling users to train and benchmark the model's performance on long-sequence tasks.

YOCO · high confidence

New CUDA-accelerated Dynamic and Light Convolution layers for faster sequence generation

This change introduces two new neural network modules, \DynamicconvLayer\ and \LightconvLayer\, located in \fairseq/modules/dynamicconv\_layer\ and \fairseq/modules/lightconv\_layer\. These modules provide optimized, CUDA-accelerated implementations of convolution operations designed to speed up sequence-to-sequence generation. The implementation includes custom C++/CUDA extensions (built via \setup.py\) and Python wrappers that handle both training (using custom CUDA kernels for forward and backward passes) and inference (using incremental state buffers for efficient decoding). The \DynamicconvLayer\ supports dynamic filter sizes and padding, while \LightconvLayer\ offers a lighter-weight alternative with fixed filters, both aiming to reduce latency during generation.

(repo-wide) · high confidence

New MWPBench evaluation suite with Fresh-GaokaoMath-2023 dataset

The mathscale repository now includes MWPBench, a unified benchmark for evaluating math instruction-tuned models. This addition provides a Dockerfile for setting up the evaluation environment, along with dataset files including \full\_test.json\ (aggregating sources like GSM8K) and \fresh\_gaokao\_math\_2023.json\ (containing 30 recent Gaokao math problems). Users can now run standardized evaluations against these datasets using the provided scripts.

mathscale · high confidence

New attention head selection capability for speech-to-text and text transformers

This change introduces a new \attention\_head\_selection\ example module that enables dynamic selection of attention heads based on task context (such as language or domain) for both standard Transformer and Speech-to-Text (S2T) models. It adds a \HeadSelectionLoss\ for KL regularization, a \MultiheadAttentionSelection\ module to apply selected heads, and new model classes (\HeadSelectionTransformerModel\, \HeadSelectionS2TTransformerModel\) that integrate these components. A new \SpeechToTextHeadSelectionTask\ is provided to handle dataset loading with domain/language metadata and to compute the selection loss during training.

(repo-wide) · high confidence

New fine-tuning examples for SWDE and WebSRC datasets

Added new example scripts in \markuplm/examples/fine\_tuning\ to support fine-tuning MarkupLM on the SWDE (information extraction) and WebSRC (question answering) datasets. The SWDE example (\run\_swde/\) includes utilities for packing HTML data, extracting XPath features, and performing evaluation with site-level voting and page-level constraints. The WebSRC example (\run\_websrc/\) provides dataset generation from CSV sources, conversion to SQuAD-style JSON, and training/evaluation scripts for question answering tasks.

markuplm/examples · high confidence

Release E5 embedding models with evaluation code and multilingual support

This release introduces the E5 text embedding models, including English variants (small, base, large, v2, and unsupervised), multilingual variants (small, base, large, and large-instruct), and the LLM-based e5-mistral-7b-instruct. It provides a complete evaluation suite for the BEIR and MTEB benchmarks, featuring scripts and Python code that support both standard query/passage prefixes and instruction-based prefixes for instruct models. The implementation includes position-weighted mean pooling and specific handling for multilingual tasks, enabling users to evaluate these models on retrieval, classification, clustering, and other NLP tasks.

e5 · high confidence

Release SimLM pre-training and fine-tuning code for dense retrieval

This release provides the complete codebase for SimLM, a retrieval-oriented pre-training architecture that uses a representation bottleneck and a replaced language modeling objective. It includes a four-stage supervised fine-tuning pipeline for training state-of-the-art dense retrievers and cross-encoder re-rankers on the MS-MARCO passage ranking task. The package contains scripts for downloading pre-processed data, training bi-encoder retrievers (with BM25 hard negatives and knowledge distillation), training cross-encoder re-rankers, and evaluating models against standard benchmarks like TREC DL 2019/2020. It also provides utilities for data preparation, metric computation, and hard negative mining, compatible with Hugging Face transformers and DeepSpeed for distributed training.

simlm · high confidence

Release of MiniLM v2 and Multilingual MiniLM v1 models with fine-tuning examples

This update introduces the MiniLM v2 pre-trained models, which utilize self-attention relation distillation to compress teacher models (such as RoBERTa-Large and XLMR-Large) into smaller, faster variants (e.g., L6xH384, L12xH384) while maintaining competitive performance on NLU and NLG tasks. It also includes the previously released Multilingual MiniLM v1 models. To support adoption, the repository now provides example scripts for fine-tuning these models, including \run\_xnli.py\ for cross-lingual natural language inference and instructions for sequence-to-sequence tasks like abstractive summarization.

minilm · high confidence

Release of UniLM v1 codebase and documentation

This change introduces the UniLM v1 release, providing the source code and pre-trained models for the NeurIPS 2019 paper 'Unified Language Model Pre-training for Natural Language Understanding and Generation'. The entry includes the \README.md\ with setup instructions (including Docker configuration) and links to pre-trained models, along with the core implementation in \src/biunilm/\ for fine-tuning and decoding sequence-to-sequence tasks (such as abstractive summarization on Gigaword and CNN/Daily Mail). It also adds evaluation utilities in \src/cnndm/\ for computing ROUGE scores.

unilm-v1 · high confidence

TextDiffuser demo and evaluation scripts released

The textdiffuser location now includes a Gradio-based web demo (gradio\_app.py) for interactive text rendering, along with standalone evaluation (evaluate.py) and inference (inference.py) scripts. These tools allow users to run the TextDiffuser model in three modes—text-to-image, text-to-image-with-template, and text-inpainting—either via command line or through the web interface, supporting the generation of images with visually appealing, coherent text.

textdiffuser · high confidence

Behavioural changes

Introduce InfoXLM multi-task training with XLM-Align data loading fix

This change introduces the InfoXLM framework, enabling multi-task training that combines Masked Language Modeling (MLM), Token-Level Multilingual (TLM), and eXtreme Language Contrastive Optimization (XlCo). It adds new Fairseq tasks (\infoxlm\, \xlm\_align\, \tlm\, \mlm\) and models (\infoxlm\, \xlm\_align\, \reload\_roberta\) along with corresponding data loaders and criteria. Specifically, it fixes a bug in the XLM-Align data loading process (referenced as \#404) by implementing the \XlmAlignTask\ and its associated dataset utilities (\xlm\_align.py\ in data and criterions), ensuring correct handling of word alignment and offset data during training.

infoxlm/src-infoxlm · high confidence

Dependencies

Add dependency manifests for new model implementations

This change introduces requirements.txt and pyproject.toml files for several new model implementations, including PFPO, YOCO, BEiT v2/v3, Kosmos-2.5, InfoXLM, and MWPBench. These files define the specific Python package versions and build dependencies required to run these models, such as PyTorch, Transformers, and DeepSpeed, ensuring that users can install the correct environment for each new component.

(dependencies) · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

Baseline

  • First survey — no prior run to compare against. CAI 59.

Lenses

  • Code Health 78
  • Architecture 86
  • Maturity 70
  • Readiness 40
  • Security 74
  • Domain Modelling 100

Changes since last survey

  • 300 commits — 283 feature/other, 17 fixes

By area

  • (repo) — 53 commits
  • (root) — 36 commits
  • textdiffuser/README.md — 24 commits
  • textdiffuser-2/README.md — 13 commits
  • kosmos-2/README.md — 12 commits
  • retnet/README.md — 10 commits
  • kosmos-2.5/README.md — 9 commits
  • Diff-Transformer/README.md — 7 commits
  • dit/README.md — 7 commits
  • bitnet/README.md — 5 commits
  • kosmos-g/README.md — 5 commits
  • Diff-Transformer/imgs — 4 commits
  • e5/README.md — 4 commits
  • layoutreader/README.md — 4 commits
  • mathscale/README.md — 4 commits
  • textdiffuser/data — 4 commits
  • Diff-Transformer/multihead_flashdiff_1.py — 3 commits
  • beit/README.md — 3 commits
  • e5/utils.py — 3 commits
  • kosmos-2/evaluation — 3 commits

Notable commits

  • fix: Fix a model download link
  • fix: Fix simlm training error with newer versions of transformers
  • fix: Readme fix
  • fix: fix Azure links for BEATs
  • fix: fix Azure links for SpeechLM
  • fix: fix Azure links for WavLM
  • fix: fix beit-1
  • fix: fix beit-2
  • fix: fix beit-3
  • fix: fix bug
  • fix: fix import flex_head_fa
  • fix: fix link
  • fix: fix paper released year
  • fix: fix setup
  • fix: fix sparse indices
  • fix: fix swa
  • fix: fix urls
  • change: remove unused params and add comments
  • change: Add C-MTEB eval instructions
  • change: Add MWPBench
  • …and 280 more

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

microsoft/unilm was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 19 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit 50224e387211f15ac6a3b2685730b9a0c850f145 — the exact code this score is about.
  • Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-13a154b7f5d1.