microsoft/VibeVoice
46.6
Weak · 2 August 2026
13.8k
lines of production code
Python
primary language
3
measurements over time
What this system is
VibeVoice is a modular speech processing system that provides both text-to-speech (TTS) and automatic speech recognition (ASR) capabilities. It features a streaming TTS engine for real-time audio generation and an ASR module for transcribing audio and video files. The system supports efficient model adaptation through LoRA fine-tuning and offers high-performance inference via a vLLM plugin, all wrapped in web-based demos and deployment scripts.
How it got here
2025 — VibeVoice-Realtime feature expansion
7 changes.
This period focused on the modular restructuring of the VibeVoice model to support real-time streaming text-to-speech and automatic speech recognition. The work introduced new architecture classes, processors, and scheduling components to enable low-latency inference. Additionally, a comprehensive web-based demo and associated utility scripts were added to showcase these new capabilities.
2026 — VibeVoice ASR integration
5 changes.
This period focused on integrating the VibeVoice ASR model into the vLLM serving infrastructure, enabling high-performance inference with support for audio and video inputs. The work included adding LoRA fine-tuning scripts, generating tokenizer files, and implementing a Gradio-based demo for transcription tasks.
Features
Add VibeVoice ASR LoRA fine-tuning scripts and documentation
Added new scripts and documentation for performing LoRA (Low-Rank Adaptation) fine-tuning on the VibeVoice ASR model. The \lora\_finetune.py\ script enables efficient model adaptation using the PEFT library, while \inference\_lora.py\ provides a dedicated inference workflow for fine-tuned models. The accompanying \README.md\ details the required environment setup, data formatting (JSON labels with audio segments), and training parameters. A \toy\_dataset\ directory with sample JSON files is included to demonstrate the expected data structure for training.
finetuning-asr · high confidence
Add VibeVoice-Realtime demo and experimental voice downloads
Introduces a new VibeVoice-Realtime demo application (vibevoice\_realtime\_demo.py) that serves the model via Uvicorn, along with a Colab notebook (vibevoice\_realtime\_colab.ipynb) for quickstart on T4 GPUs. Adds a shell script (download\_experimental\_voices.sh) to download and extract experimental voice presets for multiple languages, and a Python script (realtime\_model\_inference\_from\_file.py) for file-based inference. Also includes new ASR demo scripts (vibevoice\_asr\_gradio\_demo.py, vibevoice\_asr\_inference\_from\_file.py, vibevoice\_asr\_inference\_from\_file.py) supporting MPS/Apple Silicon and CPU, and a batch inference script (vibevoice\_asr\_inference\_from\_file.py) for ASR.
demo · high confidence
Add vLLM plugin for VibeVoice ASR serving
Introduces a new vLLM plugin that registers the VibeVoice model architecture, configuration, tokenizer, and processor with vLLM and Hugging Face Transformers. The plugin enables high-performance ASR (Automatic Speech Recognition) inference by mapping audio inputs to tensors and handling audio loading via FFmpeg. It also includes a configurable maximum audio duration limit (default 3660 seconds) to prevent out-of-memory errors.
_vllm\plugin · high confidence
Added VibeVoice-Realtime scheduling components
Added new files to the vibevoice/schedule module, including a DPMSolverMultistepScheduler implementation and custom timestep samplers (UniformSampler and LogitNormalSampler). These additions support the VibeVoice-Realtime feature by providing specialized diffusion scheduling and sampling logic.
vibevoice/schedule · medium confidence
Added vLLM plugin tool for generating VibeVoice tokenizer files
A new standalone Python script, generate\_tokenizer\_files.py, has been added to the vllm\_plugin/tools directory. This tool downloads the base Qwen2.5 tokenizer files and patches them with VibeVoice-specific audio tokens and chat template modifications, enabling high-performance ASR (Automatic Speech Recognition) serving via the vLLM plugin.
_vllm\plugin/tools · medium confidence
Introduce VibeVoice-Realtime streaming TTS demo
A new web-based demo for VibeVoice-Realtime streaming TTS has been added to the project. The demo provides a user interface for text-to-speech generation, supporting streaming audio output and voice preset selection. The backend service handles model loading with safe deserialization for voice presets, ensuring secure loading of model weights.
demo/web · high confidence
Introduce modular VibeVoice architecture with streaming and ASR capabilities
The VibeVoice model is restructured into a modular architecture, introducing separate model classes for streaming inference and automatic speech recognition (ASR). The new \VibeVoiceStreamingForConditionalGenerationInference\ class enables real-time text-to-speech generation, while \VibeVoiceASRForConditionalGeneration\ handles speech-to-text tasks. Configuration classes (\VibeVoiceStreamingConfig\, \VibeVoiceConfig\) and model implementations (\modeling\_vibevoice\_streaming.py\, \modeling\_vibevoice\asr.py\) are added to support these distinct workflows. Additionally, \\\init\\_.py\ files are added to enable direct imports of these new components.
vibevoice/modular · high confidence
New VibeVoice processor components for audio and ASR tasks
The \vibevoice/processor\ package now exposes a public API via a new \\_\init\\_.py\ that exports \VibeVoiceProcessor\, \VibeVoiceStreamingProcessor\, \VibeVoiceTokenizerProcessor\, and \AudioNormalizer\. This enables direct imports from the \vibevoice.processor\ namespace. The package also introduces \audio\_utils.py\ for FFmpeg-based audio loading and normalization, \vibevoice\_asr\_processor.py\ for ASR-specific processing, and \vibevoice\_streaming\_processor.py\ for streaming workflows. These additions provide the underlying audio handling and tokenization logic required for VibeVoice's speech and text processing capabilities.
vibevoice/processor · high confidence
One-click deployment and Gradio demo for ASR with video support
Users can now deploy the VibeVoice ASR server with a single command via the new start\_server.py script, which handles system dependencies, model downloads, and nginx-based data parallel load balancing. A new Gradio-based demo (gradio\_asr\_demo\_api\_video.py) provides a web interface for audio and video transcription, supporting concurrent requests and streaming output.
_vllm\plugin/scripts · high confidence
Behavioural changes
2 commits (0 fixes) modifying demo/voices
A change to existing behaviour in demo/voices — 2 commits, 25 files.
demo/voices · medium confidence · unverified
Test coverage
Added automated tests for vLLM ASR plugin
Added new test scripts in vllm\_plugin/tests to validate the vLLM plugin's ASR serving capabilities. test\_api.py covers standard transcription and optional hotwords support, while test\_api\_auto\_recover.py introduces testing for streaming output, video file handling, and automatic recovery from model repetition loops.
_vllm\plugin/tests · high confidence
Dependencies
Introduce pyproject.toml for VibeVoice package management
The project now uses a pyproject.toml file to define the VibeVoice package, specifying a minimum Python version of 3.10 and listing core dependencies such as torch, transformers (constrained to \<5.0.0), and various AI and web frameworks. It also registers a vLLM plugin entry point for high-performance ASR serving and defines optional dependencies for streaming TTS.
(dependencies) · medium confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Score
- CAI 46 → 47 (+0.5)
Lenses
- Code Health 61 → 61 (-0.2)
- Architecture 100 → 100 (+0.0)
- Maturity 52 → 52 (-0.2)
- Readiness 26 → 26 (+0.0)
- Security 93 → 100 (+7.0)
Resolved (42)
- Duplicated block (10 lines × 3) (vibevoice/modular/modeling_vibevoice.py)
- Duplicated block (11 lines × 2) (vibevoice/modular/modular_vibevoice_tokenizer.py)
- Duplicated block (11 lines × 2) (vllm_plugin/scripts/gradio_asr_demo_api_video.py)
- Duplicated block (11 lines × 2) (vllm_plugin/scripts/gradio_asr_demo_api_video.py)
- Duplicated block (12 lines × 2) (vibevoice/modular/modeling_vibevoice_asr.py)
- Duplicated block (12 lines × 2) (vllm_plugin/model.py)
- Duplicated block (12 lines × 2) (vllm_plugin/scripts/gradio_asr_demo_api_video.py)
- Duplicated block (12 lines × 3) (vibevoice/modular/modular_vibevoice_tokenizer.py)
- Duplicated block (13 lines × 2) (demo/vibevoice_asr_gradio_demo.py)
- Duplicated block (13 lines × 2) (vibevoice/modular/configuration_vibevoice.py)
- Duplicated block (14 lines × 2) (vibevoice/modular/modeling_vibevoice.py)
- Duplicated block (14 lines × 2) (vibevoice/modular/modular_vibevoice_tokenizer.py)
- Duplicated block (14 lines × 3) (vibevoice/processor/vibevoice_asr_processor.py)
- Duplicated block (14 lines × 4) (vibevoice/schedule/dpm_solver.py)
- Duplicated block (15 lines × 2) (demo/vibevoice_asr_gradio_demo.py)
- Duplicated block (15 lines × 2) (vibevoice/modular/modeling_vibevoice_streaming_inference.py)
- Duplicated block (15 lines × 2) (vibevoice/modular/modular_vibevoice_tokenizer.py)
- Duplicated block (16 lines × 2) (vibevoice/processor/vibevoice_asr_processor.py)
- Duplicated block (16 lines × 2) (vibevoice/processor/vibevoice_processor.py)
- Duplicated block (17 lines × 2) (vibevoice/modular/configuration_vibevoice.py)
- …and 22 more
New (39)
- Duplicated block (10 lines × 2) (vibevoice/modular/modeling_vibevoice.py)
- Duplicated block (10 lines × 2) (vibevoice/modular/modular_vibevoice_tokenizer.py)
- Duplicated block (10 lines × 2) (vibevoice/modular/modular_vibevoice_tokenizer.py)
- Duplicated block (10 lines × 2) (vllm_plugin/model.py)
- Duplicated block (11 lines × 2) (vibevoice/modular/modeling_vibevoice_asr.py)
- Duplicated block (11 lines × 2) (vllm_plugin/scripts/gradio_asr_demo_api_video.py)
- Duplicated block (11 lines × 3) (vibevoice/modular/modular_vibevoice_tokenizer.py)
- Duplicated block (12 lines × 2) (demo/vibevoice_asr_gradio_demo.py)
- Duplicated block (12 lines × 2) (demo/vibevoice_asr_gradio_demo.py)
- Duplicated block (12 lines × 2) (vibevoice/modular/configuration_vibevoice.py)
- Duplicated block (12 lines × 2) (vibevoice/modular/modular_vibevoice_tokenizer.py)
- Duplicated block (12 lines × 2) (vibevoice/processor/vibevoice_processor.py)
- Duplicated block (13 lines × 2) (vibevoice/modular/modeling_vibevoice.py)
- Duplicated block (13 lines × 2) (vibevoice/modular/modular_vibevoice_tokenizer.py)
- Duplicated block (13 lines × 3) (vibevoice/processor/vibevoice_asr_processor.py)
- Duplicated block (13 lines × 4) (vibevoice/schedule/dpm_solver.py)
- Duplicated block (14 lines × 2) (vibevoice/modular/modeling_vibevoice_streaming_inference.py)
- Duplicated block (14 lines × 2) (vibevoice/modular/modular_vibevoice_tokenizer.py)
- Duplicated block (14 lines × 2) (vibevoice/processor/vibevoice_asr_processor.py)
- Duplicated block (15 lines × 2) (vibevoice/processor/vibevoice_processor.py)
- …and 19 more
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
microsoft/VibeVoice was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 2 August 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 94da20d98b2fa7688e9cbfaf7692ddb4954f7600 — the exact code this score is about.
- Scored under rubric-2026.08.18 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer latest.