QwenAudio/CosyVoice
55.1
Adequate · 19 September 2026
19k
lines of production code
Python
primary language
1
measurement over time
What this system is
CosyVoice is an open-source text-to-speech system that synthesizes speech from text using a multi-stage pipeline comprising a language model, a flow-matching acoustic model, and a vocoder. It supports advanced capabilities such as zero-shot voice cloning, cross-lingual synthesis, and streaming inference, with models available in versions 1.0 through 3.0. The system provides comprehensive tooling for training, fine-tuning (including DPO and GRPO), and deployment via Web UIs, REST/gRPC APIs, and high-performance inference servers using vLLM and TensorRT-LLM.
How it got here
2024 — Initial CosyVoice release and infrastructure
18 changes.
This period marks the initial release of the CosyVoice text-to-speech system, establishing the core codebase with transformer-based models, flow-matching architectures, and LLM-driven token generation. It introduces comprehensive training pipelines, dataset processing tools, and deployment utilities including gRPC and FastAPI inference servers. The work also integrates third-party components like Matcha-TTS and sets up examples for fine-tuning on LibriTTS and MagicData-Read datasets.
2025 — CosyVoice2 deployment and training expansion
13 changes.
This period focused on expanding the CosyVoice2 ecosystem by introducing comprehensive training and fine-tuning recipes, including DPO and GRPO reinforcement learning methods. Significant effort was directed toward productionizing the model through high-performance inference solutions using vLLM, Triton Inference Server, and TensorRT-LLM, with support for streaming and disaggregated deployments. The work also extended the architecture with a new DiT-based audio generation backbone and added specialized models for audio tokenization and speaker embedding.
Features
Add CosyVoice2 GRPO reinforcement learning recipe
Introduces a new example recipe for fine-tuning the CosyVoice2 large language model using the GRPO reinforcement learning algorithm via the veRL framework. The addition includes a Dockerfile and setup scripts for the required environment, data preparation utilities to convert datasets into the veRL parquet format, and a custom reward function that leverages a Triton-based ASR server to compute pinyin-level error rates. It also provides scripts for model conversion between Hugging Face and CosyVoice native formats, as well as inference and evaluation pipelines to measure character error rate improvements.
examples/grpo · high confidence
Add CosyVoice2 model support for vLLM
Users can now run CosyVoice2 models using the vLLM inference engine. This change introduces a new model implementation that adapts the Qwen2 architecture for CosyVoice2, ensuring compatibility with HuggingFace weights. It includes specific handling for vLLM version 0.11.0 and above, which exclusively uses the V1 engine, by conditionally importing the correct sampling metadata and adjusting the logits computation logic to match the engine's requirements.
cosyvoice/vllm · high confidence
Add FastAPI-based CosyVoice2 inference server and client
Introduces a new FastAPI server (server.py) and a CLI client (client.py) for the CosyVoice2 model. The server exposes inference endpoints for standard, zero-shot, cross-lingual, and instruct modes, supporting both GET and POST requests to resolve browser compatibility issues, and binds to 0.0.0.0 by default to facilitate Docker port mapping.
runtime/python/fastapi · high confidence
Add LibriTTS CosyVoice training and inference example
Added a new example directory for fine-tuning the CosyVoice model on the LibriTTS dataset. This includes a shell script to download and prepare the LibriTTS data, extract speaker embeddings and speech tokens, and train the LLM, Flow, and HiFi-GAN components using PyTorch DDP. The entry also provides a sample JSON file for text-to-speech inference.
examples/libritts/cosyvoice · high confidence
Add LibriTTS data preparation and download scripts for CosyVoice
This change introduces the local utility scripts required to set up the LibriTTS dataset for the CosyVoice example. It adds a download script that handles fetching and extracting corpus parts (such as dev-clean, train-clean-100, etc.) with archive size validation and completion tracking to prevent redundant downloads. Additionally, it provides a data preparation script that converts raw audio and text files into standard Kaldi-style format (wav.scp, text, utt2spk, spk2utt) and supports an optional instruction field for fine-tuning. A third script is included to generate rejection samples by running zero-shot inference on short audio clips using the CosyVoice2 model, filtering out samples longer than 30 seconds.
examples/libritts/cosyvoice/local · high confidence
Add MagicData-Read example for CosyVoice training and inference
This change introduces a new example workflow for the CosyVoice model using the MagicData-Read dataset. It provides scripts to download and prepare the dataset, extract speaker embeddings and speech tokens, and train the LLM, flow, and HiFi-GAN models. The example also includes a sample TTS text file with Chinese content for testing inference.
examples/magicdata-read · high confidence
Add audio tokenizer model for extracting semantic tokens
A new Triton model repository for audio tokenization has been added, implementing a Python-based backend that accepts reference audio waveforms and lengths as input. The model utilizes the s3tokenizer library with an ONNX speech tokenizer to extract semantic tokens, which are then returned as integer outputs. This enables downstream components to process audio input by converting it into discrete semantic representations.
_runtime/triton\_trtllm/model\_repo/audio\tokenizer · high confidence
Add speaker embedding model for Triton Inference Server
Introduces a new Triton Python backend model that extracts speaker embeddings from reference audio. The model accepts FP32 waveform input and returns FP16 embeddings, utilizing TensorRT for inference when available (falling back to ONNX Runtime on CPU) to optimize performance for speaker identification tasks.
_runtime/triton\_trtllm/model\_repo/speaker\embedding · high confidence
Added Triton Inference Server model configurations for CosyVoice2 and Token2Wav
New model repository configurations have been added for the CosyVoice2 and Token2Wav models, enabling them to run on the Triton Inference Server. The CosyVoice2 model is configured as a Python backend that accepts optional reference audio and text inputs alongside required target text to produce waveform output. The Token2Wav model, also a Python backend, accepts speech tokens and optional prompt features (speech tokens, features, and speaker embedding) to generate waveform output. Both models are configured with dynamic batching and CPU instance groups, exposing the necessary input/output tensors for integration into the streaming TTS pipeline.
_runtime/triton\_trtllm/model\repo/token2wav · high confidence
Added multilingual character-level tokenizer with 25Hz text tokenization
The tokenizer module now includes a new 25Hz text tokenizer implementation, featuring a multilingual character-level vocabulary (assets/multilingual\_zh\_ja\_yue\_char\_del.tiktoken) and updated tokenizer logic (tokenizer.py) that supports a wide range of languages and special tokens for TTS and ASR tasks.
cosyvoice/tokenizer · high confidence
Initial addition of CosyVoice transformer backbone components
This change introduces the core transformer architecture for the CosyVoice speech synthesis model, adding key modules such as the encoder, decoder, attention mechanisms, convolution modules, and positional encodings. These components form the foundational neural network structure used for processing and generating speech data within the CosyVoice system.
cosyvoice/transformer · high confidence
Initial release of CosyVoice TTS with Web UI and vLLM support
The repository is initialized with the CosyVoice text-to-speech system, providing inference examples for CosyVoice 1.0, 2.0, and 3.0 models via \example.py\. A Gradio-based web interface (\webui.py\) is included for interactive synthesis, and \vllm\_example.py\ demonstrates integration with vLLM for accelerated inference. The project also includes a \third\_party/Matcha-TTS\ submodule, a Code of Conduct, and an FAQ to assist with setup.
(repo-wide) · high confidence
Initial release of CosyVoice training and export utilities
This change introduces the core command-line tools for the CosyVoice project, including scripts for training models (supporting DDP, DeepSpeed, and Direct Preference Optimization), averaging model checkpoints, and exporting models to TorchScript and ONNX formats for deployment.
cosyvoice/bin · high confidence
Initial release of CosyVoice utility modules
This change introduces the foundational utility infrastructure for the CosyVoice project. It adds a central registry in \class\_utils.py\ that maps configuration keys to specific model components (such as \CosyVoiceModel\, \CosyVoice2Model\, and \CosyVoice3Model\) and their internal parts (activations, subsamplings, embeddings, and attention mechanisms). The \common.py\ module provides core tensor operations, including padding, accuracy calculation, and repetition-aware sampling, alongside a list of predefined voice control instructions. Training workflows are supported by \executor.py\ (handling DDP and GAN training loops) and \train\_utils.py\ (initializing distributed environments, datasets, and optimizers). Frontend text processing is handled by \frontend\_utils.py\, which includes logic for splitting paragraphs, normalizing numbers, and detecting punctuation. Additional utilities cover file I/O (\file\_utils.py\), ONNX/TensorRT conversion (\onnx.py\), learning rate scheduling (\scheduler.py\), and mask generation for attention (\mask.py\).
cosyvoice/utils · high confidence
Introduce CosyVoice CLI with streaming, multi-model, and text-frontend support
The \cosyvoice/cli\ package now provides the primary interface for text-to-speech, voice conversion, and speaker management. Users can initialize \CosyVoice\, \CosyVoice2\, or \CosyVoice3\ models with optional JIT, TensorRT, or vLLM acceleration. The CLI supports streaming output, variable playback speed, and zero-shot/cross-lingual voice cloning. Text normalization now prioritizes \ttsfrd\ with a fallback to \wetext\ (or manual processing if neither is available), and inputs containing only punctuation or whitespace are filtered out to prevent crashes. Additionally, users can register custom zero-shot speakers via \add\_zero\_shot\_spk\ and manage them with \save\_spkinfo\.
cosyvoice/cli · high confidence
Introduce CosyVoice dataset processing pipeline
Adds the core dataset infrastructure for CosyVoice, including a new \Dataset\ class that manages data loading, sharding, and distributed sampling, along with a \Processor\ class for chaining data transformations. The \processor.py\ module implements specific data handling functions such as filtering samples by length, resampling audio to 22050 Hz, truncating or padding waveforms, and extracting fbank features using both standard and Whisper-based extractors. This change establishes the data pipeline foundation required for training and inference workflows within the CosyVoice system.
cosyvoice/dataset · high confidence
Introduce CosyVoice2 Triton model for streaming text-to-speech
Adds a new Triton Python backend model (\model.py\) for the \token2wav\ repository, implementing the CosyVoice2 architecture to convert tokens into audio waveforms. The model supports streaming TTS with mel and source caching to handle chunked inference, utilizes TensorRT for the flow decoder estimator, and includes fade-in/out logic for seamless audio stitching. It initializes with FP16 precision and loads speaker information from \spk2info.pt\.
_runtime/triton\_trtllm/model\repo/token2wav/1 · high confidence
Introduce CosyVoice2 streaming TTS model with speaker support
This change adds the Triton Python backend implementation for the CosyVoice2 model, enabling streaming text-to-speech generation. The model orchestrates an end-to-end pipeline by coordinating an audio tokenizer, a TensorRT-LLM component, and a vocoder. It supports dynamic chunking strategies (exponential or time-based) for streaming inference and includes built-in speaker information loading from a local \spk2info.pt\ file to allow for speaker-specific voice synthesis.
_runtime/triton\_trtllm/model\repo/cosyvoice2/1 · high confidence
Introduce DiT-based audio generation backbone
Added the core components for a Diffusion Transformer (DiT) model in the \cosyvoice/flow/DiT\ directory, including \dit.py\ and \modules.py\. This introduces a new neural network architecture featuring text and speaker embedding layers, a transformer backbone with rotary positional embeddings, and causal convolutional position embeddings, enabling the system to generate audio conditioned on text and speaker identity.
cosyvoice/flow/DiT · high confidence
Introduce HiFi-GAN vocoder training infrastructure
Added the core components for training a HiFi-GAN vocoder within the CosyVoice pipeline, including the generator model with residual blocks and sine-based pitch generation, a multi-resolution discriminator for audio quality assessment, an F0 predictor for pitch estimation, and a unified training module that orchestrates generator and discriminator loss calculations (including mel-spectral, feature matching, and TPR losses).
cosyvoice/hifigan · high confidence
Introduce LLM-based speech token generation model
Adds a new \TransformerLM\ class in \cosyvoice/llm/llm.py\ that implements a language model for generating speech tokens from text and speaker embeddings. This component handles text encoding, speaker embedding projection, and sequence generation via inference methods, forming the core logic for the LLM-driven text-to-speech pipeline.
cosyvoice/llm · high confidence
New CosyVoice acceleration solutions with Triton and TensorRT-LLM
This location introduces three new acceleration solutions for CosyVoice text-to-speech models using NVIDIA Triton Inference Server and TensorRT-LLM. The package includes a Dockerfile and Compose files for launching CosyVoice3, CosyVoice2 with UNet Token2Wav, and CosyVoice2 with DiT-based Token2Wav (from Step-Audio2). It provides Triton model repository configurations and Python backend scripts for these pipelines, along with gRPC and HTTP client utilities for benchmarking and inference. Documentation details performance metrics for streaming and offline modes, including support for disaggregated deployment where LLM and Token2Wav components run on separate GPUs.
_runtime/triton\trtllm · high confidence
New CosyVoice gRPC service and client for text-to-speech inference
This change introduces a new gRPC-based runtime for the CosyVoice text-to-speech model, providing both a server implementation and a Python client. The server exposes a streaming \Inference\ RPC that supports four distinct modes: standard fine-tuned (SFT), zero-shot, cross-lingual, and instruction-based synthesis. The accompanying client script allows users to connect to the service and send requests for any of these modes, handling the conversion of audio prompts and saving the resulting synthesized speech to a WAV file.
runtime/python/grpc · high confidence
New LibriTTS training and DPO fine-tuning recipes for CosyVoice2
This change introduces new example scripts for the CosyVoice2 model, enabling users to train on the LibriTTS dataset and perform Direct Preference Optimization (DPO). The \run.sh\ script provides a pipeline for data preparation, speaker embedding extraction, and training the LLM, flow, and HiFi-GAN components. Additionally, \run\_dpo.sh\ adds support for DPO fine-tuning on the LLM component, including steps to generate negative samples using a reference model and prepare DPO-specific data formats.
examples/libritts/cosyvoice2 · high confidence
New LibriTTS training pipeline for CosyVoice3
This change introduces a new example directory for training the CosyVoice3 model on the LibriTTS dataset. It provides a complete end-to-end workflow including data preparation, offline embedding and token extraction (with a note that online feature extraction is also supported), and model training for the LLM, flow, and HiFi-GAN components. The pipeline supports model averaging and exports models to JIT and ONNX formats for inference.
examples/libritts/cosyvoice3 · high confidence
New model conversion and testing utilities for TensorRT-LLM
Added four new Python scripts to the runtime/triton\_trtllm/scripts directory to support model preparation and validation. convert\_checkpoint.py provides a command-line tool to convert HuggingFace checkpoints into TensorRT-LLM format, supporting advanced quantization options like SmoothQuant, weight-only quantization (INT4/INT8), and parallelism configurations. convert\_cosyvoice3\_to\_hf.py introduces a specialized converter for the CosyVoice3 model, handling the merging of speech embeddings and tokenizer extensions required for TRT-LLM integration. Additionally, fill\_template.py offers a utility for substituting variables in .pbtxt configuration files, and test\_llm.py provides a reference implementation for running inference tests using the TensorRT-LLM C++ runtime.
_runtime/triton\trtllm/scripts · high confidence
New token2wav\_dit model for streaming audio synthesis
Added a new Triton model repository for \token2wav\_dit\, which converts speech tokens into audio waveforms using a CosyVoice2-based architecture. The implementation supports streaming inference with dynamic batching and optional TensorRT acceleration for the flow decoder and speaker embedding models. Users can provide reference audio for speaker cloning and target text tokens to generate continuous speech output.
_runtime/triton\_trtllm/model\_repo/token2wav\dit · high confidence
New tools for extracting embeddings, speech tokens, and building parquet datasets
Added three new scripts to the tools directory: \extract\_embedding.py\ extracts speaker embeddings from audio using an ONNX model with multi-threading; \extract\_speech\_token.py\ extracts speech tokens via Whisper and an ONNX model, including logic to convert audio to mono and handle sample rate resampling; and \make\_parquet\_list.py\ aggregates audio data, text, and optional embeddings/tokens into Parquet files, supporting Direct Preference Optimization (DPO) data structures and multi-process parallelization.
tools · high confidence
Behavioural changes
New CosyVoice flow-matching TTS model with streaming and caching support
The \cosyvoice/flow\ module now implements a new flow-matching based text-to-speech model (replacing the previous masked-diffusion approach). This change introduces a new \ConditionalDecoder\ and \ConditionalCFM\ architecture that supports classifier-free guidance and flow matching. Key user-facing capabilities include streaming inference via a \flow\_cache\ mechanism that maintains state between chunks, causal convolution blocks for real-time generation, and integration with TensorRT for optimized inference. The model also features a new \InterpolateRegulator\ for length alignment and supports online feature extraction for zero-shot voice cloning scenarios.
cosyvoice/flow · high confidence
Dependencies
Added Matcha-TTS as a third-party submodule
The project now includes the Matcha-TTS repository as a git submodule under the third\_party directory, pinning it to commit dd9105b. This integrates the external text-to-speech library into the codebase for future use.
_third\party · high confidence
Updated dependency manifests for CosyVoice2 and Triton runtime
The project's dependency files have been updated to support the CosyVoice2 example and the Triton TRT-LLM runtime. The main requirements.txt now pins specific versions for core libraries including PyTorch 2.3.1, Transformers 4.51.3, Gradio 5.4.0, and FastAPI 0.115.6, while adding platform-specific constraints for DeepSpeed (Linux only) and TensorRT (Linux only). It also introduces new dependencies such as x-transformers, uvicorn, wetext, and pyarrow. A new requirements.txt for the CosyVoice2 example includes diffusers 0.29.0 and WeTextProcessing 1.0.3, and the Triton runtime manifest lists tritonclient and other inference-related packages.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Baseline
- First survey — no prior run to compare against. CAI 55.
Lenses
- Code Health 84
- Architecture 99
- Maturity 52
- Readiness 44
- Security 68
Changes since last survey
- 300 commits — 219 feature/other, 81 fixes
By area
- (repo) — 72 commits
- (root) — 63 commits
- cosyvoice/cli — 40 commits
- runtime/triton_trtllm — 39 commits
- examples/libritts — 16 commits
- cosyvoice/bin — 13 commits
- cosyvoice/llm — 11 commits
- cosyvoice/flow — 10 commits
- cosyvoice/dataset — 6 commits
- cosyvoice/utils — 6 commits
- cosyvoice/hifigan — 4 commits
- docker/Dockerfile — 4 commits
- examples/grpo — 4 commits
- runtime/python — 4 commits
- cosyvoice/transformer — 2 commits
- asset/cross_lingual_prompt.wav — 1 commit
- asset/dingding.png — 1 commit
- asset/zero_shot_prompt.wav — 1 commit
- examples/magicdata-read — 1 commit
- tools/extract_speech_token.py — 1 commit
Notable commits
- fix: Fix CosyVoice3 config error
- fix: Fix diffusers / huggingface_hub compatibility in requirements.txt
- fix: Fix: Use wetext cache
- fix: Fix: Use wetext replace WeTextProcessing
- fix: Fix: generate token2wav_request_id from cosyvoice2
- fix: Fix: remove full_to_half
- fix: Fix: remove overwrite_cache
- fix: Fixed an issue where onnxruntime would not install on Windows
- fix: Merge pull request #1100 from jingfelix/fix/dockerfile-dependency
- fix: Merge pull request #1622 from GoyoUijin/Fix/token2wav-cache-thread-unsafe
- fix: Merge pull request #1722 from majiayu000/fix/issue-1683-ja-language-tag
- fix: Merge pull request #1758 from orbisai0security/fix/V-005-pickle-deserialization
- fix: Revert "fix triton token2wav model cache thread unsafety"
- fix: [BUG FIX] 使用 float64 避免精度误差问题,弃用 CPU 计算,避免拖累性能
- fix: fix bistream bug
- fix: fix bistream extra token
- fix: fix bug
- fix: fix bug
- fix: fix bug
- fix: fix bug
- …and 280 more
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
QwenAudio/CosyVoice was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 19 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 074ca6dc9e80a2f424f1f74b48bdd7d3fea531cc — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-5d04157a340d.