Skip to content
CAI
Software that uses CAICheck a score

microsoft/VibeVoice

46.6

Weak · 2 August 2026

13.8k

lines of production code

Python

primary language

3

measurements over time

CAI band scale
CAI trend line
CAI lens gauges

What this system is

VibeVoice is a modular speech processing system that provides both text-to-speech (TTS) and automatic speech recognition (ASR) capabilities. It features a streaming TTS engine for real-time audio generation and an ASR module for transcribing audio and video files. The system supports efficient model adaptation through LoRA fine-tuning and offers high-performance inference via a vLLM plugin, all wrapped in web-based demos and deployment scripts.

How it got here

2025 — VibeVoice-Realtime feature expansion

7 changes.

This period focused on the modular restructuring of the VibeVoice model to support real-time streaming text-to-speech and automatic speech recognition. The work introduced new architecture classes, processors, and scheduling components to enable low-latency inference. Additionally, a comprehensive web-based demo and associated utility scripts were added to showcase these new capabilities.

2026 — VibeVoice ASR integration

5 changes.

This period focused on integrating the VibeVoice ASR model into the vLLM serving infrastructure, enabling high-performance inference with support for audio and video inputs. The work included adding LoRA fine-tuning scripts, generating tokenizer files, and implementing a Gradio-based demo for transcription tasks.

Features

Add VibeVoice ASR LoRA fine-tuning scripts and documentation

Added new scripts and documentation for performing LoRA (Low-Rank Adaptation) fine-tuning on the VibeVoice ASR model. The \lora\_finetune.py\ script enables efficient model adaptation using the PEFT library, while \inference\_lora.py\ provides a dedicated inference workflow for fine-tuned models. The accompanying \README.md\ details the required environment setup, data formatting (JSON labels with audio segments), and training parameters. A \toy\_dataset\ directory with sample JSON files is included to demonstrate the expected data structure for training.

finetuning-asr · high confidence

Add VibeVoice-Realtime demo and experimental voice downloads

Introduces a new VibeVoice-Realtime demo application (vibevoice\_realtime\_demo.py) that serves the model via Uvicorn, along with a Colab notebook (vibevoice\_realtime\_colab.ipynb) for quickstart on T4 GPUs. Adds a shell script (download\_experimental\_voices.sh) to download and extract experimental voice presets for multiple languages, and a Python script (realtime\_model\_inference\_from\_file.py) for file-based inference. Also includes new ASR demo scripts (vibevoice\_asr\_gradio\_demo.py, vibevoice\_asr\_inference\_from\_file.py, vibevoice\_asr\_inference\_from\_file.py) supporting MPS/Apple Silicon and CPU, and a batch inference script (vibevoice\_asr\_inference\_from\_file.py) for ASR.

demo · high confidence

Add vLLM plugin for VibeVoice ASR serving

Introduces a new vLLM plugin that registers the VibeVoice model architecture, configuration, tokenizer, and processor with vLLM and Hugging Face Transformers. The plugin enables high-performance ASR (Automatic Speech Recognition) inference by mapping audio inputs to tensors and handling audio loading via FFmpeg. It also includes a configurable maximum audio duration limit (default 3660 seconds) to prevent out-of-memory errors.

_vllm\plugin · high confidence

Added VibeVoice-Realtime scheduling components

Added new files to the vibevoice/schedule module, including a DPMSolverMultistepScheduler implementation and custom timestep samplers (UniformSampler and LogitNormalSampler). These additions support the VibeVoice-Realtime feature by providing specialized diffusion scheduling and sampling logic.

vibevoice/schedule · medium confidence

Added vLLM plugin tool for generating VibeVoice tokenizer files

A new standalone Python script, generate\_tokenizer\_files.py, has been added to the vllm\_plugin/tools directory. This tool downloads the base Qwen2.5 tokenizer files and patches them with VibeVoice-specific audio tokens and chat template modifications, enabling high-performance ASR (Automatic Speech Recognition) serving via the vLLM plugin.

_vllm\plugin/tools · medium confidence

Introduce VibeVoice-Realtime streaming TTS demo

A new web-based demo for VibeVoice-Realtime streaming TTS has been added to the project. The demo provides a user interface for text-to-speech generation, supporting streaming audio output and voice preset selection. The backend service handles model loading with safe deserialization for voice presets, ensuring secure loading of model weights.

demo/web · high confidence

Introduce modular VibeVoice architecture with streaming and ASR capabilities

The VibeVoice model is restructured into a modular architecture, introducing separate model classes for streaming inference and automatic speech recognition (ASR). The new \VibeVoiceStreamingForConditionalGenerationInference\ class enables real-time text-to-speech generation, while \VibeVoiceASRForConditionalGeneration\ handles speech-to-text tasks. Configuration classes (\VibeVoiceStreamingConfig\, \VibeVoiceConfig\) and model implementations (\modeling\_vibevoice\_streaming.py\, \modeling\_vibevoice\asr.py\) are added to support these distinct workflows. Additionally, \\\init\\_.py\ files are added to enable direct imports of these new components.

vibevoice/modular · high confidence

New VibeVoice processor components for audio and ASR tasks

The \vibevoice/processor\ package now exposes a public API via a new \\_\init\\_.py\ that exports \VibeVoiceProcessor\, \VibeVoiceStreamingProcessor\, \VibeVoiceTokenizerProcessor\, and \AudioNormalizer\. This enables direct imports from the \vibevoice.processor\ namespace. The package also introduces \audio\_utils.py\ for FFmpeg-based audio loading and normalization, \vibevoice\_asr\_processor.py\ for ASR-specific processing, and \vibevoice\_streaming\_processor.py\ for streaming workflows. These additions provide the underlying audio handling and tokenization logic required for VibeVoice's speech and text processing capabilities.

vibevoice/processor · high confidence

One-click deployment and Gradio demo for ASR with video support

Users can now deploy the VibeVoice ASR server with a single command via the new start\_server.py script, which handles system dependencies, model downloads, and nginx-based data parallel load balancing. A new Gradio-based demo (gradio\_asr\_demo\_api\_video.py) provides a web interface for audio and video transcription, supporting concurrent requests and streaming output.

_vllm\plugin/scripts · high confidence

Behavioural changes

2 commits (0 fixes) modifying demo/voices

A change to existing behaviour in demo/voices — 2 commits, 25 files.

demo/voices · medium confidence · unverified

Test coverage

Added automated tests for vLLM ASR plugin

Added new test scripts in vllm\_plugin/tests to validate the vLLM plugin's ASR serving capabilities. test\_api.py covers standard transcription and optional hotwords support, while test\_api\_auto\_recover.py introduces testing for streaming output, video file handling, and automatic recovery from model repetition loops.

_vllm\plugin/tests · high confidence

Dependencies

Introduce pyproject.toml for VibeVoice package management

The project now uses a pyproject.toml file to define the VibeVoice package, specifying a minimum Python version of 3.10 and listing core dependencies such as torch, transformers (constrained to \<5.0.0), and various AI and web frameworks. It also registers a vLLM plugin entry point for high-performance ASR serving and defines optional dependencies for streaming TTS.

(dependencies) · medium confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

Score

  • CAI 46 → 47 (+0.5)

Lenses

  • Code Health 61 → 61 (-0.2)
  • Architecture 100 → 100 (+0.0)
  • Maturity 52 → 52 (-0.2)
  • Readiness 26 → 26 (+0.0)
  • Security 93 → 100 (+7.0)

Resolved (42)

  • Duplicated block (10 lines × 3) (vibevoice/modular/modeling_vibevoice.py)
  • Duplicated block (11 lines × 2) (vibevoice/modular/modular_vibevoice_tokenizer.py)
  • Duplicated block (11 lines × 2) (vllm_plugin/scripts/gradio_asr_demo_api_video.py)
  • Duplicated block (11 lines × 2) (vllm_plugin/scripts/gradio_asr_demo_api_video.py)
  • Duplicated block (12 lines × 2) (vibevoice/modular/modeling_vibevoice_asr.py)
  • Duplicated block (12 lines × 2) (vllm_plugin/model.py)
  • Duplicated block (12 lines × 2) (vllm_plugin/scripts/gradio_asr_demo_api_video.py)
  • Duplicated block (12 lines × 3) (vibevoice/modular/modular_vibevoice_tokenizer.py)
  • Duplicated block (13 lines × 2) (demo/vibevoice_asr_gradio_demo.py)
  • Duplicated block (13 lines × 2) (vibevoice/modular/configuration_vibevoice.py)
  • Duplicated block (14 lines × 2) (vibevoice/modular/modeling_vibevoice.py)
  • Duplicated block (14 lines × 2) (vibevoice/modular/modular_vibevoice_tokenizer.py)
  • Duplicated block (14 lines × 3) (vibevoice/processor/vibevoice_asr_processor.py)
  • Duplicated block (14 lines × 4) (vibevoice/schedule/dpm_solver.py)
  • Duplicated block (15 lines × 2) (demo/vibevoice_asr_gradio_demo.py)
  • Duplicated block (15 lines × 2) (vibevoice/modular/modeling_vibevoice_streaming_inference.py)
  • Duplicated block (15 lines × 2) (vibevoice/modular/modular_vibevoice_tokenizer.py)
  • Duplicated block (16 lines × 2) (vibevoice/processor/vibevoice_asr_processor.py)
  • Duplicated block (16 lines × 2) (vibevoice/processor/vibevoice_processor.py)
  • Duplicated block (17 lines × 2) (vibevoice/modular/configuration_vibevoice.py)
  • …and 22 more

New (39)

  • Duplicated block (10 lines × 2) (vibevoice/modular/modeling_vibevoice.py)
  • Duplicated block (10 lines × 2) (vibevoice/modular/modular_vibevoice_tokenizer.py)
  • Duplicated block (10 lines × 2) (vibevoice/modular/modular_vibevoice_tokenizer.py)
  • Duplicated block (10 lines × 2) (vllm_plugin/model.py)
  • Duplicated block (11 lines × 2) (vibevoice/modular/modeling_vibevoice_asr.py)
  • Duplicated block (11 lines × 2) (vllm_plugin/scripts/gradio_asr_demo_api_video.py)
  • Duplicated block (11 lines × 3) (vibevoice/modular/modular_vibevoice_tokenizer.py)
  • Duplicated block (12 lines × 2) (demo/vibevoice_asr_gradio_demo.py)
  • Duplicated block (12 lines × 2) (demo/vibevoice_asr_gradio_demo.py)
  • Duplicated block (12 lines × 2) (vibevoice/modular/configuration_vibevoice.py)
  • Duplicated block (12 lines × 2) (vibevoice/modular/modular_vibevoice_tokenizer.py)
  • Duplicated block (12 lines × 2) (vibevoice/processor/vibevoice_processor.py)
  • Duplicated block (13 lines × 2) (vibevoice/modular/modeling_vibevoice.py)
  • Duplicated block (13 lines × 2) (vibevoice/modular/modular_vibevoice_tokenizer.py)
  • Duplicated block (13 lines × 3) (vibevoice/processor/vibevoice_asr_processor.py)
  • Duplicated block (13 lines × 4) (vibevoice/schedule/dpm_solver.py)
  • Duplicated block (14 lines × 2) (vibevoice/modular/modeling_vibevoice_streaming_inference.py)
  • Duplicated block (14 lines × 2) (vibevoice/modular/modular_vibevoice_tokenizer.py)
  • Duplicated block (14 lines × 2) (vibevoice/processor/vibevoice_asr_processor.py)
  • Duplicated block (15 lines × 2) (vibevoice/processor/vibevoice_processor.py)
  • …and 19 more

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

microsoft/VibeVoice was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 2 August 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit 94da20d98b2fa7688e9cbfaf7692ddb4954f7600 — the exact code this score is about.
  • Scored under rubric-2026.08.18 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer latest.