microsoft/unilm
58.7
Adequate · 19 September 2026
813.6k
lines of production code
Python
primary language
1
measurement over time
What this system is
This system is a comprehensive research repository for multimodal machine learning, providing PyTorch implementations of state-of-the-art models for natural language processing, computer vision, and speech recognition. It enables users to pre-train, fine-tune, and evaluate architectures such as transformers, diffusion models, and document understanding systems across diverse tasks including translation, OCR, and dense retrieval. The codebase also includes specialized toolkits for optimizing inference speed, handling long-context sequences, and performing domain adaptation.
How it got here
2020–2021 — multimodal and multilingual model expansion
23 changes.
This period focused on releasing a diverse suite of pre-trained models and fine-tuning toolkits for text, document, image, and speech modalities. Key additions included MiniLM, LayoutLM, BEiT, WavLM, and TrOCR, alongside specialized toolkits for sequence-to-sequence tasks, domain adaptation, and cross-lingual transfer. The work established comprehensive infrastructure for multimodal understanding and generation across various benchmarks.
2022 — multimodal and speech model releases
22 changes.
This period focused on releasing a broad suite of new models and frameworks for document AI, vision, and speech processing, including EdgeFormer, DiT, LayoutLMv3, BEiT-2/3, SpeechLM, and SpeechT5. The work involved implementing core architectures, training pipelines, and evaluation scripts for tasks such as object detection, text detection, dense retrieval, and audio tokenization. These releases expanded the project's capabilities in multimodal understanding and efficient sequence generation.
2023–2026 — multimodal and reasoning model releases
19 changes.
This period focused on releasing and expanding several multimodal and reasoning models, including TextDiffuser, Kosmos-2, and LatentLM, alongside their associated training and evaluation tooling. It also introduced specialized architectures for long-context inference and efficient attention, such as YOCO, ReSA, and Diff Transformer, while establishing new benchmarks for mathematical reasoning and text rendering.
Features
Add BEATs audio pre-training and tokenization models
This release introduces the official PyTorch implementation of BEATs (Audio Pre-Training with Acoustic Tokenizers), including the core model architectures in BEATs.py and Tokenizers.py, supporting transformer-based feature extraction and acoustic tokenization via vector quantization. The addition includes the necessary backbone and module components (backbone.py, modules.py, quantizer.py) to load pre-trained and fine-tuned models for tasks such as audio representation learning and classification, along with updated documentation and download links for the model checkpoints.
beats · high confidence
Add BEiT semantic segmentation support for ADE20K
Introduces a new semantic segmentation module for the ADE20K dataset using the BEiT transformer backbone. This includes the BEiT backbone implementation, UperNet model configuration, and training/evaluation scripts for both base (12-layer) and large (24-layer) variants. The feature supports fine-tuning from ImageNet pre-trained weights and includes multi-scale testing configurations for improved segmentation accuracy.
_beit/semantic\segmentation · high confidence
Add BEiT v2 semantic segmentation support for ADE20K
This change introduces a complete implementation for fine-tuning BEiT v2 (base and large variants) on the ADE20K semantic segmentation dataset using the UperNet architecture. It adds the BEiT backbone code, custom training utilities including a layer-decay optimizer constructor and an Apex-based AMP runner, and configuration files for 160k and 320k iteration schedules with 512x512 and 640x640 crop sizes. Users can now train and evaluate BEiT v2 models on ADE20K using the provided mmsegmentation-compatible configs and pretrained checkpoints.
_beit2/semantic\segmentation · high confidence
Add BEiT-3 vision and vision-language model implementation
Introduces the official PyTorch implementation and pretrained models for BEiT-3, a vision and vision-language pretraining model. This addition includes the core model architecture (leveraging the torchscale library), dataset handling for image-text pairs, and specific fine-tuning engines for downstream tasks such as image classification, visual question answering (VQAv2), visual reasoning (NLVR2), image captioning, and image-text retrieval. The release also provides pretrained checkpoints for base and large model sizes, along with documentation and setup instructions.
beit3 · high confidence
Add DiT image classification fine-tuning support for RVL-CDIP
This change introduces a new classification module within the DiT project, enabling users to fine-tune and evaluate the Document Image Transformer on the RVL-CDIP dataset. The addition includes the full training and evaluation pipeline (\run\_class\_finetuning.py\, \engine\_for\_finetuning.py\), model definitions for fine-tuning (\modeling\_finetune.py\), and necessary data handling utilities (\datasets.py\, \dataset\_folder.py\). It also provides DeepSpeed configuration for distributed training and detailed usage instructions in the README.
dit/classification · high confidence
Add DiT-based text detection training and evaluation scripts
Introduces a new text detection module that leverages the DiT (Document Image Transformer) backbone within a Mask R-CNN framework. The change adds a training script (\train\_net.py\) capable of registering the FUNSD dataset, initializing the DiT configuration, and launching distributed training or evaluation using Detectron2. It also includes a README documenting fine-tuned model weights, data preparation steps, and usage examples for training and evaluating text detection models on FUNSD.
_dit/text\detection · high confidence
Add LayoutLMv3 object detection training examples and infrastructure
This release adds a complete set of examples and supporting code for training LayoutLMv3 on object detection tasks. It includes a YAML configuration file for a cascade LayoutLMv3 model on the PubLayNet dataset, along with a custom Detectron2 backbone implementation that integrates the LayoutLMv3 transformer. The package also provides dataset conversion scripts to transform ICDAR data into COCO format, an adaptive binarization utility for image preprocessing, and a custom training loop and checkpointer to handle model weight loading and position embedding interpolation.
layoutlmv3/examples · high confidence
Add MARIOEval evaluation and generation scripts
The textdiffuser/eval directory now includes scripts to generate images and evaluate them using the MARIOEval benchmark. MARIOEval\_generate.py supports image generation via Stable Diffusion, ControlNet, and DeepFloyd pipelines, while MARIOEval\_evaluate.py computes CLIPScore and FID metrics. Additionally, ocr\_eval.py provides OCR-based evaluation (precision, recall, accuracy) for text rendering, and supporting modules for CLIPScore and FID calculations are included.
textdiffuser/eval · high confidence
Add PFPO (ICLR 2025) analysis and pseudo-test-case pipelines
Introduces a new suite of scripts under \PFPO/scripts/apps/\ to support the PFPO methodology. This includes \analyze/\ tools for computing output frequencies and generating visualization histograms, \prm/\ scripts for constructing process reward model samples and sampling code steps, and \pseudo\_test\_cases/\ utilities for generating, cleaning, and executing synthetic test cases to build DPO preference pairs.
(repo-wide) · high confidence
Add SpeechLM speech pre-training and fine-tuning code
Added the SpeechLM module, including the main model implementation (SpeechLM.py), supporting modules (modules.py), and comprehensive documentation (README.md). This addition provides pre-trained and fine-tuned models for speech tasks, along with scripts and instructions for feature extraction, automatic speech recognition (ASR) on LibriSpeech, and speech translation (ST) on CoVoST-2. The codebase includes a submodule for fairseq and requires specific dependencies like sacrebleu for setup.
speechlm · high confidence
Add SpeechT5 unified speech and text processing implementation
Introduces the SpeechT5 module, providing a unified encoder-decoder architecture for spoken language processing tasks. This addition includes the core model components, data handling via a MultitaskDataset, and specific loss functions (criterions) for pre-training, speech-to-text (ASR), and text-to-speech (TTS) tasks. It also provides a script for generating speech outputs and integrates with the fairseq framework via a submodule, enabling users to load pre-trained models and perform inference or fine-tuning for speech and text modalities.
speecht5 · high confidence
Add TrOCR text recognition module with augmentation and inference support
Introduces the TrOCR (Transformer-based Optical Character Recognition) module, providing an end-to-end text recognition approach using pre-trained image and text Transformers. This addition includes the core model implementations (ViTTRModel, TrOCRModel), a custom GPT2BPEEnhancedSpace tokenizer, and data handling utilities for datasets like IAM and SROIE. It also provides a comprehensive suite of image augmentation techniques (blur, noise, geometry, weather effects) to improve model robustness, along with scripts for fine-tuning, evaluation, and inference.
trocr · high confidence
Add VLMo vision-language pre-training implementation
Introduces the VLMo (Unified Vision-Language Pre-Training with Mixture-of-Modality-Experts) module, providing the PyTorch implementation, configuration, and data handling for pre-training and fine-tuning. This includes dataset preparation scripts and data modules for GCC, SBU, Visual Genome, COCO, Flickr30K, VQAv2, NLVR2, and WikiBK, along with a multi-task data loader and training entry point using PyTorch Lightning.
vlmo · high confidence
Add ViT-based text detection backbone and evaluation utilities
Introduces a new text detection module under \dit/text\_detection/ditod\ that integrates Multi-Path Vision Transformer (ViT) backbones (including BEiT, DiT, DeiT, and MAE variants) with a Feature Pyramid Network (FPN) via the Detectron2 registry. This change also adds supporting infrastructure for model training and evaluation, including a custom dataset mapper, checkpointer, trainer, and a FUNSD-specific evaluator, alongside utility modules for configuration management, image conversion, and standard ICDAR detection metrics (IoU, DetEval, ICDAR2013, and MTWI2018).
_dit/text\detection/ditod · high confidence
Add XDoc fine-tuning code and documentation
This change introduces the fine-tuning infrastructure for the XDoc model, including a comprehensive README with instructions for SQuAD, FUNSD, and WebSRC tasks, alongside the \layoutlmft\ library. The library provides custom dataset builders for FUNSD and XFUN, data collators that handle image and bounding-box inputs, and model implementations for LayoutLMv2 and LayoutXLM, enabling users to fine-tune XDoc on document understanding benchmarks.
xdoc · high confidence
Add angle and quadrilateral text rendering extensions
New training and inference scripts have been added to the TextDiffuser-2 extensions directory to support rendering text with specific geometric layouts: angles and quadrilaterals. This includes Python scripts for training (\train\_textdiffuser2\_t2i\_full\_angle.py\, \train\_textdiffuser2\_t2i\_full\_quadrilateral.py\) and inference (\inference\_textdiffuser2\_t2i\_full\_angle.py\, \inference\_textdiffuser2\_t2i\_full\_quadrilateral.py\), along with corresponding shell launch scripts and template input files (\angle\_template\_file.txt\, \quadrilateral\_template\_file.txt\) that define the coordinate-based layout for these shapes.
textdiffuser-2/extensions · high confidence
Add deprecated LayoutLM document understanding module
This change introduces the \layoutlm/deprecated\ directory, containing the original LayoutLM implementation for multimodal document understanding (text, layout, and image). It includes the core model classes (\LayoutlmConfig\, \LayoutlmForSequenceClassification\, \LayoutlmForTokenClassification\), data processors for the FUNSD and RVL-CDIP datasets, and example scripts for fine-tuning on sequence labeling and classification tasks. The module also adds standard Python development configuration files (\.flake8\, \.isort.cfg\, \.pre-commit-config.yaml\) to enforce code style and linting.
layoutlm/deprecated · high confidence
Add object detection inference and training scripts for DiT
This change introduces the \dit/object\_detection\ module, providing scripts to run inference and train Mask R-CNN and Cascade Mask R-CNN models using the DiT backbone on PubLayNet and ICDAR 2019 cTDaR datasets. Users can now perform document layout analysis via the new \inference.py\ script, which supports visualization of detected objects, or fine-tune the models using \train\_net.py\ with configurable datasets and GPU settings. The module also includes utilities for data preparation, such as \convert\_to\_coco\_format.py\ for converting ICDAR annotations to COCO format and \adaptive\_binarize.py\ for image preprocessing.
_dit/object\detection · high confidence
Add relation extraction decoder module
The library now includes a new relation extraction (RE) decoder component within the modules/decoders package. This addition introduces a BiaffineAttention operator and a REDecoder class that enable binary relation classification by combining entity embeddings with hidden states, allowing users to extract structured relationships between entities in their documents.
layoutlmft/layoutlmft/modules/decoders · high confidence
Add support for MiniLM and UniLM models in s2s-ft
The s2s-ft module now supports fine-tuning MiniLM and UniLM models in addition to BERT and RoBERTa. This includes new configuration classes (MinilmConfig, UnilmConfig), tokenizers (MinilmTokenizer, UnilmTokenizer), and state-dict converters to handle model-specific weight formats (e.g., UniLM's QKV linear layers). The modeling code has been updated to load these models, resize position embeddings if necessary, and apply source/target type IDs for sequence-to-sequence tasks.
_s2s-ft/s2s\ft · high confidence
Added DynamicConv, LightConv, and Product Quantization modules
This change introduces new neural network layer implementations and quantization utilities to the library. It adds \DynamicconvLayer\ and \LightconvLayer\, which provide optimized CUDA-accelerated convolution operations with support for incremental decoding during inference. Additionally, it includes a Product Quantization (PQ) subsystem under \modules/quantization/pq\, featuring an Expectation-Maximization (EM) algorithm for clustering weights and specific quantized module classes (\PQConv2d\, \PQEmbedding\, \PQLinear\) that allow for model compression by storing centroids and assignments rather than full-precision weights.
_edgelm/fairseq/modules/dynamicconv\_layer, edgelm/fairseq/modules/lightconv\layer, edgelm/fairseq/modules/quantization · high confidence
Added Mario LAION example dataset and processing scripts
The \textdiffuser/data\ directory now includes a new example dataset under \mario-laion-example\, containing structured entries with captions, OCR text with bounding box coordinates, and metadata JSON files for specific image keys. Additionally, a Python script \mario-laion-unzip.py\ has been added to facilitate the multiprocessed extraction of LAION OCR data archives, and a Jupyter notebook \visualize\_charseg.ipynb\ is provided for visualizing character segmentation results using the new data.
textdiffuser/data · high confidence
Added SpeechLM dataset configurations and assets for CommonVoice and LibriSpeech
This change introduces dataset configuration files and supporting assets for the SpeechLM project, specifically for English-to-German translation using CommonVoice v4 and speech recognition using LibriSpeech. For CommonVoice, it adds YAML configs (\config\_base\_ende.yaml\, \config\_large\_ende.yaml\) defining SentencePiece tokenization, audio input settings (16kHz, mono), and a 100-sample development manifest (\dev-sample100\_st\_en\_de\_local.tsv\) linking English audio files to German translations. It also includes the corresponding SentencePiece vocabulary files (\spm\_char\_st\_en\_de.txt\, \spm\_char\_st\_en\_de.vocab\). For LibriSpeech, it adds configuration and dictionary files for both phone-level and hidden-unit models, including character-level (\dict.ltr.txt\) and phone-level (\dict.phn.txt\) vocabularies, and a 100-sample training manifest (\train\_sample100.ltr\) with phoneme-transcribed text.
speechlm/dataset · high confidence
Added TextDiffuser demo assets and example cases
This update adds demonstration assets for the TextDiffuser model, including Python implementation files (modeling utilities, DDPMScheduler, and UNet2DConditionModel) and a collection of text prompt examples for text inpainting, text-to-image, and text-to-image-with-template scenarios. These files provide the necessary code structure and input cases to run and test the TextDiffuser demo.
textdiffuser/assets · high confidence
Added TextDiffuser model components for keyword layout prediction
This change introduces the core model architecture for TextDiffuser, a system that predicts the spatial layout of keywords within user prompts. It adds a \LayoutTransformer\ and \TextConditioner\ (leveraging CLIP) to determine where text should be placed, along with a UNet-based \TextSegmenter\ for segmentation tasks. These components enable the model to generate structured layouts for text-based image generation.
textdiffuser/model, textdiffuser/textdiffuser-ckpt · high confidence
Added fairseq documentation and project metadata
The infoxlm/fairseq directory now includes the Sphinx documentation source files (RST), configuration, and static assets, providing users with guides on getting started, command-line tools, and extending the toolkit. This change also adds standard project metadata files, including the MIT License, Code of Conduct, and Contributing guidelines, to support community engagement and clarify usage rights.
(repo-wide) · high confidence
Adds Aggressive Decoding (GAD) plugins for non-autoregressive sequence generation
This change introduces a new set of Fairseq plugins under \decoding/GAD/block\_plugins\ that enable Aggressive Decoding (GAD) for sequence-to-sequence tasks. The update adds a \BlockNAT\ model (\block\ architecture) which implements a non-autoregressive transformer with specific forward-pass logic for GAD-style iterative refinement. It also registers a \glat\_loss\ criterion for training with label smoothing and dual imitation, and a \translation\_lev\_modified\ task that supports various noise injection strategies (random delete, random mask, block mask) to facilitate the training of these non-autoregressive models.
decoding · high confidence
Initial release of AdaLM domain adaptation toolkit
This change introduces the AdaLM repository, providing code to fine-tune pre-trained language models for specific domains (biomedical and computer science) and to generate incremental domain-specific vocabularies. The \finetune\ directory contains scripts for sequence classification (\run\_classifier.py\) and named entity recognition (\run\_ner.py\, \run\_pico.py\), supporting models like BERT, RoBERTa, and DistilBERT via the Hugging Face \transformers\ library. The \incr\_bpe\ directory includes tools to build domain-specific subword vocabularies by extending base BERT vocabularies, utilizing a simplified version of the \tensor2tensor\ library's BPE implementation. Documentation and pre-trained model links are also included.
adalm · high confidence
Initial release of BEiT image transformer implementation
This change introduces the BEiT (BERT Pre-Training of Image Transformers) codebase, providing the official PyTorch implementation for self-supervised pre-training and fine-tuning of image transformers. The release includes the core model architecture, training engines for both pre-training and fine-tuning, and data augmentation pipelines that support masked image modeling. It also integrates a DALL-E discrete VAE tokenizer for generating visual tokens and provides documentation and scripts for fine-tuning on ImageNet-1k and semantic segmentation on ADE20K.
beit · high confidence
Initial release of EdgeFormer documentation and project scaffolding
This change introduces the \edgelm\ directory, establishing the project structure for EdgeFormer (a parameter-efficient transformer for on-device seq2seq generation). It adds the core project metadata files, including the MIT License, Code of Conduct, and Contributing guidelines. It also provides a comprehensive README detailing pretrained model checkpoints, benchmark performance on tasks like CoNLL-14, XSUM, and SQuAD-NQG, and setup/fine-tuning instructions using fairseq. Additionally, it includes a full Sphinx documentation suite (\docs/\) covering command-line tools, library references (models, tasks, optimizers), and tutorials for extending the framework.
edgelm · high confidence
Initial release of Kosmos-2 grounding toolkit and documentation
This change introduces the initial codebase and documentation for Kosmos-2, a multimodal large language model grounded to the world. It includes a comprehensive README detailing model checkpoints, setup instructions (Docker and Conda), and evaluation metrics for tasks like phrase grounding and visual question answering. The release provides data preparation scripts for the GRIT dataset, a Gradio-based interactive demo for local hosting, and utility modules for decoding bounding box coordinates from model captions and visualizing them on images.
kosmos-2 · high confidence
Initial release of Kosmos-2.5 multimodal literate model
This change introduces the Kosmos-2.5 repository, a multimodal model designed for machine reading of text-intensive images. The release includes the inference code (\inference.py\), model architecture definitions (\kosmos2\_5/models/\), and task handling (\kosmos2\_5/tasks/\) required to run the model. Users can now perform two primary tasks: text recognition (OCR) that generates spatially-aware text blocks, and image-to-markdown conversion that captures document structure and styles. The implementation relies on Fairseq for the language model backbone and integrates with Hugging Face Transformers for image processing, supporting GPU acceleration via Flash Attention2.
kosmos-2.5 · high confidence
Initial release of LayoutLMv3 document AI model
This change introduces the LayoutLMv3 model, a multimodal transformer for Document AI that uses unified text and image masking for pre-training. The release includes the \layoutlmv3-base\, \layoutlmv3-large\, and \layoutlmv3-base-chinese\ pre-trained models, along with setup scripts and documentation for fine-tuning on tasks such as form understanding (FUNSD, XFUND) and document layout analysis (PubLayNet).
layoutlmv3 · high confidence
Initial release of LayoutReader for reading order detection
This change introduces the LayoutReader component, a seq2seq model that captures text and layout information to predict reading order, significantly improving OCR engine results. The release includes the core training script (run\_seq2seq.py) and decoding script (decode\_seq2seq.py) which leverage the LayoutLM architecture and the s2s-ft library, along with a setup.py defining the s2s-ft package dependencies (including transformers \<= 2.10.0). A comprehensive README provides installation instructions, usage examples for training and decoding on the ReadingBank dataset, and performance benchmarks.
layoutreader · high confidence
Initial release of LayoutXLM and XFUN dataset support
This change introduces the LayoutXLM model and its associated tokenizer, extending the existing LayoutLMv2 architecture to support multilingual document understanding. It also adds the XFUN dataset loader for cross-lingual form understanding, along with the necessary data collators, evaluation metrics for relation extraction, and utility functions for bounding box normalization and image loading.
layoutlmft/layoutlmft · high confidence
Initial release of MarkupLM pre-trained models and fine-tuning code
This change introduces the MarkupLM component, a multimodal pre-training method for text and markup language designed for visually-rich document understanding tasks such as webpage QA and information extraction. The release includes the \setup.py\ package definition for \markuplmft\ (version 0.1) and a comprehensive \README.md\ detailing installation, fine-tuning procedures for datasets like WebSRC and SWDE, and links to pre-trained models (MarkupLM-Base and MarkupLM-Large) available on Hugging Face.
markuplm · high confidence
Initial release of PFPO source code and configuration
This change introduces the source code and configuration files for Preference Optimization for Reasoning with Pseudo Feedback (PFPO), a method for improving reasoning in large language models. The release includes a README detailing experimental results on mathematical reasoning (MATH, GSM8K) and coding tasks (LiveCodeBench, HumanEval, MBPP), along with Hydra configuration files for running inference and training pipelines using models like Mathstral-7B and Deepseek-Coder-7B-v1.5 via vLLM.
PFPO · high confidence
Initial release of TextDiffuser 2 assets and training data
This change introduces the foundational components for TextDiffuser 2, including a custom attention processor implementation in \assets/attention\_processor.py\ that supports advanced features like LoRA and xFormers optimization. It also adds a \reference\_requirements.txt\ file pinning the environment dependencies (such as \diffusers==0.24.0\, \torch\, and \gradio==3.50.2\) and provides a 5,000-sample dataset (\data/layout\_planner\_data\_5k.json\) for training the model to plan visual text layouts, along with a utility script to visualize these layout plans.
textdiffuser-2/assets, textdiffuser-2/data · high confidence
Initial release of TextDiffuser-2 codebase and demos
This change introduces the complete TextDiffuser-2 repository, including training scripts for the layout planner (using FastChat/Vicuna) and diffusion models (full-parameter and LoRA fine-tuning), as well as inference scripts and Gradio-based interactive demos for text-to-image generation and text inpainting. The codebase supports text rendering with layout planning via language models, handles coordinate-based text positioning, and includes specific support for angle/quadrilateral text layouts and inpainting tasks.
textdiffuser-2 · high confidence
Initial release of s2s-ft sequence-to-sequence fine-tuning toolkit
The s2s-ft package is introduced as a PyTorch-based toolkit for fine-tuning pre-trained Transformer models (including BERT, RoBERTa, XLM-RoBERTa, Electra, UniLM, and MiniLM) for sequence-to-sequence tasks. This release provides the core training script (run\_seq2seq.py) and decoding utilities (decode\_seq2seq.py), along with specialized tokenizers and model configurations for UniLM and MiniLM. It also includes a comprehensive suite of evaluation scripts for standard summarization benchmarks (XSum, CNN/Daily Mail, Gigaword) and a utility for generating sequences from beam search traces, enabling users to fine-tune and evaluate models on these datasets out of the box.
s2s-ft · high confidence
Initial release of s2s\_ft module for LayoutReader
This change introduces the s2s\_ft (sequence-to-sequence fine-tuning) module to the LayoutReader package, providing the core infrastructure for loading and fine-tuning transformer models. It adds configuration classes for BERT, MiniLM, and UniLM models, along with tokenizers and model implementations that support pre-trained weights from Azure Blob storage and Hugging Face. The module includes utilities for converting state dictionaries between different model formats (e.g., RoBERTa to BERT, LayoutLM to BERT) and data loading pipelines for sequence-to-sequence tasks, enabling the layout reader to leverage these specific model architectures for document understanding.
_layoutreader/s2s\ft · high confidence
Initial release of the layoutlmft document understanding toolkit
This change introduces the \layoutlmft\ package, a multimodal fine-tuning toolkit for document understanding that supports LayoutLM, LayoutLMv2, and LayoutXLM models. It provides example scripts for fine-tuning on the FUNSD dataset for key-value extraction (\run\_funsd.py\) and the XFUN dataset for relation extraction (\run\_xfun\_re.py\) and sequence entity recognition (\run\_xfun\_ser.py\), along with the necessary project configuration files like \setup.py\ and \Makefile\.
layoutlmft · high confidence
Initial release of xTune cross-lingual fine-tuning scripts and documentation
Adds the xTune repository, providing a two-stage training process for cross-lingual fine-tuning of XLM-Roberta models on XTREME datasets (XNLI, PANX, PAWS-X, UDPOS, MLQA, TyDiQA, XQuAD). The change includes a README with setup instructions, shell scripts for downloading datasets and models, and specific training configurations for both 'cross-lingual-transfer' and 'translate-train-all' settings.
xtune · high confidence
Initial support for Vision Transformer (ViT) backbones in object detection
This change introduces a new \ditod\ module for the object detection pipeline, enabling the use of Vision Transformer backbones (such as BEiT, DIT, DeiT, and MAE) integrated with Feature Pyramid Networks (FPN). It adds configuration options for these models, a custom dataset mapper for DETR-style training, and specialized evaluation tools for table detection (ICDAR). Additionally, it includes a custom checkpointer to handle position embedding interpolation when loading pre-trained weights into different input resolutions.
_dit/object\detection/ditod · high confidence
Introduce BEiT v2 implementation with vector-quantized visual tokenizers
Adds the official PyTorch codebase for BEiT v2, a masked image modeling approach that uses vector-quantized visual tokenizers (VQ-KD) instead of discrete pixel tokens. This release includes the core training engines for pretraining, fine-tuning, and VQ-KD tokenizer training, along with dataset utilities and data augmentation pipelines specific to the v2 architecture. Documentation and configuration examples are provided for pretraining on ImageNet-1k and fine-tuning for image classification and semantic segmentation.
beit2 · high confidence
Introduce DeltaLM encoder-decoder model with Fairseq integration
Adds the DeltaLM model, an encoder-decoder architecture for multilingual language generation and translation, built as a submodular dependency on Fairseq. The implementation includes the core model definitions in deltalm/models/deltalm.py, which register with Fairseq's model registry and support loading pretrained checkpoints for fine-tuning. Entry-point scripts (train.py, generate.py, preprocess.py, interactive.py) are provided to expose Fairseq's CLI commands for the new architecture, alongside example shell scripts for data preparation, tokenization, training, and evaluation on the IWSLT14 German-English dataset.
deltalm · high confidence
Introduce Diff Transformer attention modules with FlashAttention support
This change adds a new Diff Transformer implementation to the repository, providing both standard PyTorch and FlashAttention-optimized versions of multi-head differential attention. The codebase includes \MultiheadDiffAttn\ for baseline comparison, \MultiheadFlashDiff1\ for libraries supporting different Q/K/V dimensions (like xformers or customized flash-attention), and \MultiheadFlashDiff2\ for standard flash-attention. It also introduces a Triton-based rotary embedding kernel (\kernel/rotary.py\) and a fallback RMSNorm implementation (\rms\_norm.py\), along with an example script to compare parameter counts between the differential and conventional attention mechanisms.
Diff-Transformer · high confidence
Introduce Differential Transformer V2 (DIFF V2) with FlashAttention integration
Added a new \MultiheadFlashDiffV2\ attention module that implements the Differential Transformer V2 architecture. This change improves inference efficiency by allowing the use of standard FlashAttention without custom kernels, enhances training stability by removing per-head RMSNorm, and simplifies parameterization by using token-specific, head-wise projected lambda values instead of a globally shared scalar. The implementation is provided in \multihead\_flashdiffv2.py\ alongside updated documentation.
Diff-Transformer/Diff-Transformer-V2 · high confidence
Introduce LatentLM multimodal latent language modeling toolkit
Adds the LatentLM component, providing an official PyTorch implementation for multimodal latent language modeling with next-token diffusion. This includes core model architectures (DiT, Transformer), supporting modules (RMSNorm, EMA, attention kernels), and evaluation utilities for computing FID and Inception Score metrics, along with scripts for inference speed benchmarking and fidelity evaluation.
LatentLM · high confidence
Introduce Rectified Sparse Attention (ReSA) for efficient long-context inference
Adds the ReSA module, which improves sparse decoding by periodically refreshing the KV cache to prevent error accumulation and maintain generation quality. This change introduces a new \KVManager\ for block-sparse attention, custom Triton kernels for sparse decoding and attention with KV cache, and integration into the model's attention layer to enable up to 2.42x speedup at 256K context length while preserving near-lossless performance.
ReSA · high confidence
Introduce SpeechLM pre-training, fine-tuning, and decoding configurations
This change adds the configuration files for the SpeechLM model, including Hydra configs for pre-training (base and large variants), fine-tuning (100h and 960h subsets), and inference decoding (fairseq LM, KenLM, and Viterbi). These configs define the task, model architecture, optimization, and dataset settings for the SpeechLM pipeline.
speechlm/speechlm · high confidence
Introduce WavLM self-supervised speech model
Adds the WavLM model implementation and documentation, providing pre-trained checkpoints (Base, Base+, Large) and code for feature extraction and fine-tuning on speech tasks such as speaker verification, separation, diarization, and recognition.
wavlm · high confidence
Introduce YOCO decoder-decoder architecture with long-context evaluation capabilities
This change adds the YOCO (You Only Cache Once) implementation, a decoder-decoder architecture for large language models designed to improve long-context retrieval. The update includes the core model components, such as the GateRetention mechanism, cross-attention, and optimized kernels for SwiGLU and rotary embeddings. It also introduces specific evaluation criteria for the 'Needle in a Haystack' and multi-needle tasks, along with scripts for training and harness-based evaluation, enabling users to train and benchmark the model's performance on long-sequence tasks.
YOCO · high confidence
New CUDA-accelerated Dynamic and Light Convolution layers for faster sequence generation
This change introduces two new neural network modules, \DynamicconvLayer\ and \LightconvLayer\, located in \fairseq/modules/dynamicconv\_layer\ and \fairseq/modules/lightconv\_layer\. These modules provide optimized, CUDA-accelerated implementations of convolution operations designed to speed up sequence-to-sequence generation. The implementation includes custom C++/CUDA extensions (built via \setup.py\) and Python wrappers that handle both training (using custom CUDA kernels for forward and backward passes) and inference (using incremental state buffers for efficient decoding). The \DynamicconvLayer\ supports dynamic filter sizes and padding, while \LightconvLayer\ offers a lighter-weight alternative with fixed filters, both aiming to reduce latency during generation.
(repo-wide) · high confidence
New MWPBench evaluation suite with Fresh-GaokaoMath-2023 dataset
The mathscale repository now includes MWPBench, a unified benchmark for evaluating math instruction-tuned models. This addition provides a Dockerfile for setting up the evaluation environment, along with dataset files including \full\_test.json\ (aggregating sources like GSM8K) and \fresh\_gaokao\_math\_2023.json\ (containing 30 recent Gaokao math problems). Users can now run standardized evaluations against these datasets using the provided scripts.
mathscale · high confidence
New attention head selection capability for speech-to-text and text transformers
This change introduces a new \attention\_head\_selection\ example module that enables dynamic selection of attention heads based on task context (such as language or domain) for both standard Transformer and Speech-to-Text (S2T) models. It adds a \HeadSelectionLoss\ for KL regularization, a \MultiheadAttentionSelection\ module to apply selected heads, and new model classes (\HeadSelectionTransformerModel\, \HeadSelectionS2TTransformerModel\) that integrate these components. A new \SpeechToTextHeadSelectionTask\ is provided to handle dataset loading with domain/language metadata and to compute the selection loss during training.
(repo-wide) · high confidence
New fine-tuning examples for SWDE and WebSRC datasets
Added new example scripts in \markuplm/examples/fine\_tuning\ to support fine-tuning MarkupLM on the SWDE (information extraction) and WebSRC (question answering) datasets. The SWDE example (\run\_swde/\) includes utilities for packing HTML data, extracting XPath features, and performing evaluation with site-level voting and page-level constraints. The WebSRC example (\run\_websrc/\) provides dataset generation from CSV sources, conversion to SQuAD-style JSON, and training/evaluation scripts for question answering tasks.
markuplm/examples · high confidence
Release E5 embedding models with evaluation code and multilingual support
This release introduces the E5 text embedding models, including English variants (small, base, large, v2, and unsupervised), multilingual variants (small, base, large, and large-instruct), and the LLM-based e5-mistral-7b-instruct. It provides a complete evaluation suite for the BEIR and MTEB benchmarks, featuring scripts and Python code that support both standard query/passage prefixes and instruction-based prefixes for instruct models. The implementation includes position-weighted mean pooling and specific handling for multilingual tasks, enabling users to evaluate these models on retrieval, classification, clustering, and other NLP tasks.
e5 · high confidence
Release SimLM pre-training and fine-tuning code for dense retrieval
This release provides the complete codebase for SimLM, a retrieval-oriented pre-training architecture that uses a representation bottleneck and a replaced language modeling objective. It includes a four-stage supervised fine-tuning pipeline for training state-of-the-art dense retrievers and cross-encoder re-rankers on the MS-MARCO passage ranking task. The package contains scripts for downloading pre-processed data, training bi-encoder retrievers (with BM25 hard negatives and knowledge distillation), training cross-encoder re-rankers, and evaluating models against standard benchmarks like TREC DL 2019/2020. It also provides utilities for data preparation, metric computation, and hard negative mining, compatible with Hugging Face transformers and DeepSpeed for distributed training.
simlm · high confidence
Release of MiniLM v2 and Multilingual MiniLM v1 models with fine-tuning examples
This update introduces the MiniLM v2 pre-trained models, which utilize self-attention relation distillation to compress teacher models (such as RoBERTa-Large and XLMR-Large) into smaller, faster variants (e.g., L6xH384, L12xH384) while maintaining competitive performance on NLU and NLG tasks. It also includes the previously released Multilingual MiniLM v1 models. To support adoption, the repository now provides example scripts for fine-tuning these models, including \run\_xnli.py\ for cross-lingual natural language inference and instructions for sequence-to-sequence tasks like abstractive summarization.
minilm · high confidence
Release of UniLM v1 codebase and documentation
This change introduces the UniLM v1 release, providing the source code and pre-trained models for the NeurIPS 2019 paper 'Unified Language Model Pre-training for Natural Language Understanding and Generation'. The entry includes the \README.md\ with setup instructions (including Docker configuration) and links to pre-trained models, along with the core implementation in \src/biunilm/\ for fine-tuning and decoding sequence-to-sequence tasks (such as abstractive summarization on Gigaword and CNN/Daily Mail). It also adds evaluation utilities in \src/cnndm/\ for computing ROUGE scores.
unilm-v1 · high confidence
TextDiffuser demo and evaluation scripts released
The textdiffuser location now includes a Gradio-based web demo (gradio\_app.py) for interactive text rendering, along with standalone evaluation (evaluate.py) and inference (inference.py) scripts. These tools allow users to run the TextDiffuser model in three modes—text-to-image, text-to-image-with-template, and text-inpainting—either via command line or through the web interface, supporting the generation of images with visually appealing, coherent text.
textdiffuser · high confidence
Behavioural changes
Introduce InfoXLM multi-task training with XLM-Align data loading fix
This change introduces the InfoXLM framework, enabling multi-task training that combines Masked Language Modeling (MLM), Token-Level Multilingual (TLM), and eXtreme Language Contrastive Optimization (XlCo). It adds new Fairseq tasks (\infoxlm\, \xlm\_align\, \tlm\, \mlm\) and models (\infoxlm\, \xlm\_align\, \reload\_roberta\) along with corresponding data loaders and criteria. Specifically, it fixes a bug in the XLM-Align data loading process (referenced as \#404) by implementing the \XlmAlignTask\ and its associated dataset utilities (\xlm\_align.py\ in data and criterions), ensuring correct handling of word alignment and offset data during training.
infoxlm/src-infoxlm · high confidence
Dependencies
Add dependency manifests for new model implementations
This change introduces requirements.txt and pyproject.toml files for several new model implementations, including PFPO, YOCO, BEiT v2/v3, Kosmos-2.5, InfoXLM, and MWPBench. These files define the specific Python package versions and build dependencies required to run these models, such as PyTorch, Transformers, and DeepSpeed, ensuring that users can install the correct environment for each new component.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Baseline
- First survey — no prior run to compare against. CAI 59.
Lenses
- Code Health 78
- Architecture 86
- Maturity 70
- Readiness 40
- Security 74
- Domain Modelling 100
Changes since last survey
- 300 commits — 283 feature/other, 17 fixes
By area
- (repo) — 53 commits
- (root) — 36 commits
- textdiffuser/README.md — 24 commits
- textdiffuser-2/README.md — 13 commits
- kosmos-2/README.md — 12 commits
- retnet/README.md — 10 commits
- kosmos-2.5/README.md — 9 commits
- Diff-Transformer/README.md — 7 commits
- dit/README.md — 7 commits
- bitnet/README.md — 5 commits
- kosmos-g/README.md — 5 commits
- Diff-Transformer/imgs — 4 commits
- e5/README.md — 4 commits
- layoutreader/README.md — 4 commits
- mathscale/README.md — 4 commits
- textdiffuser/data — 4 commits
- Diff-Transformer/multihead_flashdiff_1.py — 3 commits
- beit/README.md — 3 commits
- e5/utils.py — 3 commits
- kosmos-2/evaluation — 3 commits
Notable commits
- fix: Fix a model download link
- fix: Fix simlm training error with newer versions of transformers
- fix: Readme fix
- fix: fix Azure links for BEATs
- fix: fix Azure links for SpeechLM
- fix: fix Azure links for WavLM
- fix: fix beit-1
- fix: fix beit-2
- fix: fix beit-3
- fix: fix bug
- fix: fix import flex_head_fa
- fix: fix link
- fix: fix paper released year
- fix: fix setup
- fix: fix sparse indices
- fix: fix swa
- fix: fix urls
- change: remove unused params and add comments
- change: Add C-MTEB eval instructions
- change: Add MWPBench
- …and 280 more
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
microsoft/unilm was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 19 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 50224e387211f15ac6a3b2685730b9a0c850f145 — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-13a154b7f5d1.