Skip to content
CAI
Software that uses CAICheck a score

coqui-ai/TTS

48.4

Weak · 26 September 2026

57.1k

lines of production code

Python

primary language

4

measurements over time

CAI band scale
CAI trend line
CAI lens gauges

What this system is

Coqui TTS is a comprehensive Python library for text-to-speech synthesis, speaker encoding, and voice conversion. It provides a unified framework for training, fine-tuning, and inferring a wide variety of neural models, including XTTS, VITS, Bark, and Tortoise, across multiple languages and speakers. The system supports end-to-end workflows from dataset preparation and model training to deployment via a local web server or Python API.

How it got here

2018–2020 — Coqui rebranding and architectural overhaul

27 changes.

This period marked the project's rebranding to Coqui TTS and a comprehensive architectural refactoring, replacing legacy modules with unified base classes for models, vocoders, and text processing. The work established a modern, modular codebase with a new Python API, CLI tools, and a demo server, while significantly expanding test coverage and documentation to support the new structure.

2021 — VITS integration and configuration standardization

34 changes.

This period focused on introducing the VITS model architecture and its associated training recipes, alongside expanding support for other models like Align-TTS and SpeedySpeech. The codebase underwent significant refactoring to standardize configuration management using Coqpit and dataclasses, ensuring type safety and consistency across all TTS and vocoder models. Comprehensive test suites were added to validate the new components, text processing utilities, and end-to-end training workflows.

2022–2023 — multilingual expansion and advanced model integration

32 changes.

This period focused on significantly expanding language support through new phonemizers and text normalization for Belarusian, Bangla, Korean, and Chinese, alongside adding training recipes for Thorsten German and VCTK. It also introduced a wide array of advanced model architectures, including Neural HMM, FreeVC, Tortoise, Bark, Delightful-TTS, and XTTS v1/v2, complete with corresponding training infrastructure and fine-tuning recipes.

Features

Add Align-TTS duration predictor and MDN layer components

New files have been added to the Align-TTS layer module to support specific model architecture components. The duration predictor now utilizes a transformer-based block with positional encoding to process character embeddings, while a Mixture of Density Network (MDN) block has been introduced to output mean and log-variance parameters for alignment modeling.

_TTS/tts/layers/align\tts · high confidence

Add AlignTTS training recipe for LJSpeech

A new training script for the AlignTTS model has been added to the LJSpeech recipe directory. This entry point allows users to train an AlignTTS text-to-speech model using the LJSpeech dataset, utilizing the updated Trainer API and supporting evaluation splits for model validation during training.

_recipes/ljspeech/align\tts · high confidence

Add Bark text-to-speech model implementation

Added the core implementation for the Bark text-to-speech model, including the GPT-based architecture for text-to-semantic and coarse audio generation, the FineGPT model for fine audio code generation, and inference utilities for voice cloning and semantic token generation.

TTS/tts/layers/bark · high confidence

Add Belarusian TTS training recipe

Added a new recipe for training a Belarusian text-to-speech model, including configuration scripts for GlowTTS and HiFiGAN vocoder, a Jupyter notebook for speaker selection from the Common Voice corpus, and Docker support for the preparation environment.

recipes/bel-alex73 · high confidence

Add Capacitron-enabled Tacotron2 training recipe

A new training script for the LJSpeech dataset has been added that enables the Capacitron VAE architecture within Tacotron2. This configuration allows users to train models with speaker embedding-based voice cloning capabilities by utilizing the Capacitron optimizer, dynamic convolution attention, and specific loss weighting parameters tailored for the VAE component.

recipes/ljspeech/tacotron2-Capacitron · high confidence

Add Chinese Mandarin and French text normalization utilities

New text processing modules have been added for Chinese Mandarin and French. The Chinese Mandarin module includes a phonemizer that converts text to phonemes using pinyin and tone information, a number-to-character converter for handling numeric values in speech synthesis, and a mapping dictionary for pinyin to phoneme conversion. The French module introduces an abbreviation expansion system that replaces common French abbreviations (such as 'M.', 'Mme', 'N.B') with their full forms to improve text normalization before synthesis.

_TTS/tts/utils/text/chinese\mandarin · high confidence

Add Delightful-TTS acoustic model layers

The TTS engine now includes the neural network layers required for the Delightful-TTS model. This change adds the \AcousticModel\ which integrates a Conformer-based encoder and decoder, along with dedicated modules for predicting and adapting pitch, energy, and phoneme-level prosody. It also introduces supporting components such as the \ReferenceEncoder\ for style token extraction, \AlignmentNetwork\ for duration prediction, and various convolutional utilities like \ConvNorm\ with optional weight normalization.

_TTS/tts/layers/delightful\tts · high confidence

Add FastPitch and FastSpeech LJSpeech training recipes

New training scripts for FastPitch and FastSpeech on the LJSpeech dataset have been added to the recipes directory. These scripts provide a complete, runnable configuration for training these specific TTS models, including dataset setup, audio processing parameters, and integration with the Trainer framework.

_recipes/ljspeech/fast\_pitch, recipes/vctk/fast\_pitch, recipes/vctk/fast\speech · high confidence

Add LJSpeech SpeedySpeech training recipe

A new training script for the SpeedySpeech model on the LJSpeech dataset has been added. This recipe configures the ForwardTTS model with specific audio settings (22050 Hz sample rate, silence trimming) and dataset parameters, allowing users to train a text-to-speech model using the SpeedySpeech architecture on the LJSpeech corpus.

_recipes/ljspeech/speedy\speech · high confidence

Add Neural HMM TTS training recipe for LJSpeech

A new training script has been added for the Neural HMM TTS model, specifically configured for the LJSpeech dataset. This recipe allows users to train a Neural HMM TTS model using phoneme-based text cleaning, mixed precision training, and specific audio preprocessing settings (22050 Hz sample rate, silence trimming).

_recipes/ljspeech/neuralhmm\tts · high confidence

Add Thorsten German (de-DE) training recipes

New training scripts and configuration files have been added for the Thorsten German dataset, enabling users to train text-to-speech and vocoder models in German. The update includes recipes for AlignTTS, GlowTTS, SpeedySpeech, and Tacotron2 for synthesis, alongside HiFi-GAN, MultiBand-MelGAN, UnivNet, WaveGrad, and WaveRNN for vocoding. Each script is pre-configured with German phoneme settings and includes logic to automatically download the Thorsten-Voice dataset if it is not already present.

_recipes/thorsten\DE · high confidence

Add UnivNet vocoder training recipe for LJSpeech

Users can now train the UnivNet vocoder on the LJSpeech dataset using the provided training script. This new recipe initializes the GAN model with specific audio processing and data loading configurations, and utilizes the updated TrainerArgs and Trainer API to manage the training loop.

recipes/ljspeech/univnet · high confidence

Add VCTK speaker encoder training recipe

A new training script for the ResNet-based speaker encoder has been added, providing a ready-to-use configuration for training on the VCTK dataset. The recipe includes specific settings for audio augmentation (such as RIR simulation and additive noise), model parameters (ResNet architecture with 512-dimensional embeddings), and training hyperparameters (including batch composition and loss function). Users can now leverage this script to generate d-vectors for speaker identification tasks using the VCTK corpus.

(repo-wide) · high confidence

Add VCTK training recipes for SpeedySpeech and Tacotron-DDC

New training scripts have been added for the VCTK multi-speaker dataset, enabling users to train SpeedySpeech and Tacotron-DDC models. These recipes configure the models for multi-speaker synthesis using speaker embeddings, set up the VCTK dataset formatter, and utilize the updated TrainerArgs API for initialization.

_recipes/vctk/speedy\speech, recipes/vctk/tacotron-DDC · high confidence

Add VITS model recipe for LJSpeech

A new training script for the VITS text-to-speech model on the LJSpeech dataset has been added. This recipe configures the VitsConfig with specific audio parameters (22050 Hz sample rate, 80 mel bands) and dataset settings, initializes the audio processor and tokenizer, loads training and evaluation samples with a configurable split, and launches the Trainer to fit the model.

_recipes/ljspeech/vits\tts · high confidence

Add XTTS v1 fine-tuning recipe for LJSpeech

A new training script has been added to fine-tune the XTTS v1 model on the LJSpeech dataset. This recipe handles the automatic download of necessary XTTS v1.1.2 checkpoint files (model weights, vocabulary, DVAE, and mel stats) and configures the GPTTrainer with specific hyperparameters, including a batch size of 3, gradient accumulation steps, and a learning rate of 5e-06, to facilitate voice cloning and transfer learning.

_recipes/ljspeech/xtts\v1 · high confidence

Add YourTTS training recipe for CML-TTS dataset

A new training script for the YourTTS model has been added to the multilingual recipes. This recipe enables training on the CML-TTS dataset across multiple languages (Portuguese, Polish, Italian, French, Dutch, German, Spanish) combined with English from LibriTTS. It automatically handles dataset downloading, speaker embedding extraction using a pre-trained speaker encoder, and configures the VITS-based YourTTS architecture with external speaker embeddings.

_recipes/multilingual/cml\yourtts · high confidence

Add multilingual VITS TTS training recipes

New training scripts are provided for the \recipes/multilingual/vits\_tts\ directory, enabling users to train VITS-based text-to-speech models on multi-lingual datasets. The \train\_vits\_tts.py\ script demonstrates training with character-based input, while \train\_vits\_tts\_phonemes.py\ offers a phoneme-based alternative. Both recipes are configured to use the Mailabs dataset, support multi-speaker and multi-language capabilities via dedicated managers, and include example test sentences in English, French, German, and Russian.

_recipes/multilingual/vits\tts · high confidence

Added Bangla text preprocessing and phonemization support

Users can now process Bangla text for TTS with built-in normalization, including English-to-Bangla digit conversion, numerization, and expansion of common religious attributions. The new phonemizer handles whitespace collapsing and sentence segmentation to prepare text for synthesis.

TTS/tts/utils/text/bangla · high confidence

Added Capacitron training recipes for Blizzard 2013 dataset

New training scripts and documentation have been added for the Blizzard 2013 dataset, enabling users to train Tacotron and Tacotron2 models with the Capacitron VAE extension for improved prosody modeling. The entry includes a README with instructions for obtaining the dataset and preprocessing steps, alongside \train\_capacitron\_t1.py\ and \train\_capacitron\_t2.py\ which configure the specific hyperparameters, optimizers, and audio processing settings required for these models.

recipes/blizzard2013 · high confidence

Added English text normalization utilities for abbreviations, numbers, and time

New modules have been added to the English text processing pipeline to normalize spoken output. The \abbreviations\ module expands common titles (e.g., 'mr' to 'mister'). The \number\_norm\ module handles the expansion of numeric values, including commas, decimals, ordinals, and currencies (USD, EUR, GBP, JPY). The \time\_norm\ module converts time formats (e.g., '10:30 am') into spoken words. These utilities are now available for use in the TTS text preprocessing stage.

TTS/tts/utils/text/english · high confidence

Added FastSpeech2 training recipe for LJSpeech

A new training script for the FastSpeech2 model on the LJSpeech dataset has been added to the recipes directory. This entry point configures the model with specific audio settings (22050 Hz sample rate, silence trimming) and enables auxiliary feature computation for pitch (f0) and energy, which are cached to disk. The script also handles the automatic download and use of a Tacotron2-DCA model to pre-compute attention masks required for training, streamlining the setup process for users wishing to train this specific TTS architecture.

_recipes/ljspeech/delightful\tts, recipes/ljspeech/fastspeech2 · high confidence

Added Kokoro speech dataset training recipe

Users can now train a Tacotron2 model with Double Decoder Consistency (DDC) specifically for the Kokoro speech dataset. This change introduces a new recipe directory containing a shell script to handle dataset preparation (splitting metadata and computing statistics) and a JSON configuration file defining the model architecture, audio parameters (22kHz sample rate, 80 mel bands), and training hyperparameters tailored for Japanese phoneme processing.

recipes/kokoro · high confidence

Added LJSpeech dataset download and training recipe scripts

Users can now easily download and prepare the LJSpeech dataset using the new \download\_ljspeech.sh\ script, which handles downloading, extraction, and train/validation split creation. A \README.md\ has been added to guide users through running the training templates for various models, noting that these are starting points rather than optimized configurations.

recipes/ljspeech · high confidence

Added LJSpeech training recipe for the Overflow model

Users can now train the Overflow TTS model on the LJSpeech dataset using the new \recipes/ljspeech/overflow/train\_overflow.py\ script. This entry point configures the \Overflow\ model with specific audio settings (22050 Hz sample rate, silence trimming) and text processing (US English phonemes), allowing for direct model training and evaluation on this standard dataset.

recipes/ljspeech/overflow · high confidence

Added VCTK dataset download script

A new shell script (download\_vctk.sh) has been added to the VCTK recipe directory to automate the acquisition of the VCTK dataset. The script downloads the VCTK-Corpus-0.92.zip archive, extracts it, and organizes the data into the appropriate directory structure for training and validation splits.

recipes/vctk · high confidence

Added VCTK training recipe for DelightfulTTS

A new training script has been added for the DelightfulTTS model, specifically configured to train on the VCTK dataset. This recipe sets up the necessary data loading, audio processing, and speaker management to facilitate model training using the VCTK corpus.

_recipes/vctk/delightful\tts · high confidence

Added XTTS v2.0 fine-tuning recipe for LJSpeech

A new training script (train\_gpt\_xtts.py) is now available in the LJSpeech recipe directory, enabling users to fine-tune the XTTS v2.0 model. The recipe automatically downloads the required XTTS v2.0 checkpoint, tokenizer, and DVAE files, and configures the GPTTrainer with specific parameters such as a 22050 Hz sample rate, AdamW optimizer, and MultiStepLR scheduler for efficient training on the LJSpeech dataset.

_recipes/ljspeech/xtts\v2 · high confidence

Automated synchronization of TTS CLI help text in README

A new script, scripts/sync\_readme.py, has been added to automatically keep the TTS command-line usage instructions in the project README.md in sync with the actual help text generated by the TTS CLI tool. This ensures that documentation remains accurate without manual updates, reducing the risk of outdated or misleading usage information for users.

scripts · high confidence

Initial implementation of Tortoise TTS model layers

Added the core neural network components for the Tortoise text-to-speech model, including the autoregressive GPT-2 inference model, diffusion-based audio decoder, and supporting utilities for audio loading, voice management, and tokenization. This introduces the architectural foundation for Tortoise-style voice cloning and generation within the TTS module.

TTS/tts/layers/tortoise · high confidence

Initial release of the TTS demo server

Introduces a new Flask-based demo server for Text-to-Speech synthesis. Users can now run a local web interface to generate audio from text, supporting both pre-trained models (via the ModelManager) and custom model paths. The server exposes an API endpoint at /api/tts and includes a UI for selecting speakers and languages, with configuration managed through command-line arguments or a conf.json file.

TTS/server · high confidence

Introduce FreeVC voice conversion model

A new FreeVC voice conversion model is now available in the TTS library. This change adds the model implementation, its specific configuration classes (FreeVCConfig, FreeVCArgs, FreeVCAudioConfig), and the necessary base classes and module structures to support it. Users can now utilize the FreeVC architecture for voice conversion tasks by selecting 'freevc' as the model type in their configuration.

TTS/vc · high confidence

Introduce FreeVC voice conversion module with WavLM and speaker encoder support

Added the FreeVC voice conversion module under TTS/vc/modules/freevc, providing a new capability for converting speech while preserving speaker identity. This implementation includes core neural network components (modules.py) using weight-normalized WaveNet-style layers, audio processing utilities (mel\_processing.py, commons.py) for spectrogram generation and tensor manipulation, and a dedicated speaker encoder (speaker\_encoder/) that extracts embeddings from audio to guide the conversion. It also integrates a WavLM feature extractor (wavlm/) with automatic model downloading and configuration, enabling robust content representation for the voice conversion pipeline.

TTS/vc/modules/freevc · high confidence

Introduce Neural HMM-based TTS architecture with Overflow layers

Adds a new TTS model implementation based on the Neural HMM approach, replacing non-monotonic attention with an autoregressive left-right hidden Markov model for monotonic alignment. This change introduces a new \overflow\ module containing the \NeuralHMM\ core layer, an \Encoder\ that expands input length by states per phone, an \Outputnet\ for emission parameters, a \Decoder\ wrapping Glow-TTS components, and utility scripts for plotting transition probabilities. Users can now utilize this attention-free model which supports deterministic duration generation and gradient checkpointing to improve training stability and synthesis quality.

TTS/tts/layers/overflow · high confidence

Introduce TTS v0.22.0 with new Python API and model registry

This release updates the TTS library to version 0.22.0 and introduces a new high-level Python API (\TTS.api.TTS\) that simplifies loading and using text-to-speech models. The API allows users to initialize models by name or path, automatically handling downloads and synthesizer setup, while exposing properties for listing available models, speakers, and languages. A new \.models.json\ file serves as a central registry for all supported TTS models (including XTTS v2, Bark, and various single/multi-lingual models), providing metadata like download URLs, hashes, and licenses. Additionally, a \BaseTrainerModel\ abstraction is added to standardize model initialization and inference interfaces for future model implementations.

TTS · high confidence

Introduce VITS model architecture components

This change adds the core neural network layers required for the VITS (Variational Inference with adversarial learning over end-to-end Text-to-Speech) model. The new files in \TTS/tts/layers/vits\ implement the \VitsDiscriminator\ (combining HiFiGAN-style scale and period discriminators), the \TextEncoder\ (supporting optional language embeddings), \ResidualCouplingBlock\ for the flow-based generator, \StochasticDurationPredictor\ (with spline flows and language conditioning), and the underlying \piecewise\_rational\_quadratic\_transform\ utility for normalizing flows. These components enable the VITS architecture for high-quality, multilingual text-to-speech synthesis.

TTS/tts/layers/vits · high confidence

Introduce dedicated NumPy and PyTorch audio transform modules

The \TTS/utils/audio\ package now includes separate \numpy\_transforms.py\ and \torch\_transforms.py\ modules to handle audio processing. The NumPy module provides standalone functions for spectrogram operations (STFT, mel conversion, pre-emphasis, etc.), while the PyTorch module introduces a \TorchSTFT\ neural network layer for efficient batch processing on GPUs. The existing \AudioProcessor\ class has been updated to import and utilize these new transform functions, enabling the TTS system to leverage PyTorch-based audio computations where applicable.

TTS/utils/audio · high confidence

Introduce new TTS utility modules for data, alignment, and speaker management

This change adds a comprehensive set of new utility modules to the TTS package. It introduces \SpeakerManager\ and \LanguageManager\ classes to handle multi-speaker and multi-lingual configurations, including loading embeddings and managing ID mappings. New data utilities provide functions for padding inputs, balancing dataset samples by audio length and language, and handling random audio segments. Alignment and visualization helpers are added, including a Cython-optimized monotonic alignment algorithm, SSIM loss calculation, and plotting tools for alignments, spectrograms, pitch, and energy. Additionally, a new synthesis module standardizes the inference interface, and a Fairseq checkpoint rehashing utility is provided to support model migration.

TTS/tts/utils · high confidence

Introduces dedicated phonemizers for Belarusian, Bangla, Japanese, Korean, and Chinese

The text processing pipeline now includes specific phonemizer implementations for Belarusian (using the Fanetyka Java library via jpype), Bangla, Japanese (Japanese phonemizer), Korean (KO\_KR\_Phonemizer), and Chinese (ZH\_CN\_Phonemizer). These new backends are registered in the central phonemizer registry and set as the defaults for their respective language codes (be, bn, ja-jp, ko-kr, zh-cn), allowing users to generate phonemes for these languages with improved accuracy compared to generic fallbacks. A new \MultiPhonemizer\ class and \BasePhonemizer\ abstract class were also added to manage these language-specific backends uniformly.

TTS/tts/utils/text/phonemizers · high confidence

Introduction of XTTS neural network layers

Added the core neural network components for the XTTS text-to-speech model, including the GPT-based language model, HiFi-GAN vocoder, DVAE quantizer, and Perceiver encoder. This commit also introduces streaming inference capabilities and a comprehensive tokenizer with support for multiple languages and text normalization.

TTS/tts/layers/xtts · high confidence

Introduction of XTTS v2.0 GPT training infrastructure

This change introduces the core training components for the XTTS v2.0 model, specifically adding a new \XTTSDataset\ class and a \GPTTrainer\ implementation within the \TTS/tts/layers/xtts/trainer\ directory. The dataset module provides logic for loading and conditioning audio samples, including support for language-based sampling, reproducible evaluation modes, and optional masking of ground-truth prompts. The trainer module handles the initialization of the XTTS model, GPT encoder, and Discrete VAE, while also managing checkpoint loading and mel spectrogram extraction for style encoding. This establishes the foundation for fine-tuning and training the XTTS v2.0 architecture.

TTS/tts/layers/xtts/trainer · high confidence

Introduction of experimental vocoder module

An experimental vocoder component has been added to the TTS package, providing implementations for Melgan, MultiBand-Melgan, ParallelWaveGAN, and GAN-TTS (Discriminator Only). This module allows users to combine these vocoder models with existing TTS models, offering a flexible framework for training, fine-tuning, and restoring models via command-line tools and TensorBoard integration.

TTS/vocoder · high confidence

New HiFi-GAN training recipe for LJSpeech

A new training script has been added for the HiFi-GAN vocoder using the LJSpeech dataset. This recipe configures the model with specific hyperparameters (such as batch size, sequence length, and learning rates) and utilizes the updated TrainerArgs and GAN model classes to initialize and fit the training process.

recipes/ljspeech/hifigan · high confidence

New Korean text normalization and phonemization support

Added a new Korean text processing module that normalizes input text by converting English acronyms and special characters into their Korean phonetic equivalents, and converts the resulting text into Hangeul Jamo or ASCII phonemes using the g2pkk library.

TTS/tts/utils/text/korean · high confidence

New TTS CLI and training scripts

The TTS/bin directory now includes a comprehensive set of command-line tools for the TTS workflow. Users can synthesize speech using the new \synthesize.py\ script, train TTS and vocoder models via \train\_tts.py\ and \train\_vocoder.py\, and train speaker encoders with \train\_encoder.py\. Additional utility scripts have been added for dataset preparation and analysis, including \compute\_embeddings.py\ for speaker embeddings, \extract\_tts\_spectrograms.py\ for feature extraction, \compute\_attention\_masks.py\ for attention visualization, \compute\_statistics.py\ for normalization stats, \find\_unique\_chars.py\ and \find\_unique\_phonemes.py\ for character set analysis, \remove\_silence\_using\_vad.py\ for audio cleaning, \resample.py\ for sample rate conversion, and \collect\_env\_info.py\ for environment diagnostics.

TTS/bin · high confidence

New TTS utility modules for training, inference, and model management

This change introduces a comprehensive set of new utility modules in TTS/utils to support the updated training and inference workflows. The new TrainerCallback class provides lifecycle hooks (init, epoch, step, interrupt) for models, criteria, and optimizers, enabling custom training logic. A new Synthesizer class offers a unified Python API for inference, handling TTS, vocoder, and voice conversion models with automatic sentence segmentation and speaker/language management. The ModelManager class simplifies model discovery and downloading from a central registry, while new download utilities support resuming, progress bars, and archive extraction. Additional modules include a CapacitronOptimizer for dual-optimizer training, distributed training helpers, dataset downloaders for common corpora (LJSpeech, VCTK, LibriTTS, etc.), and improved I/O with fsspec support for cloud storage and pickle renaming for compatibility.

TTS/utils · high confidence

New Tacotron2-DDC training recipe for LJSpeech

A new training script for the Tacotron2-DDC model on the LJSpeech dataset has been added. This recipe demonstrates the updated API usage, including the migration to \TrainerArgs\ for configuration, the use of \TTSTokenizer\ for text processing, and the integration of explicit \eval\_split\ and \eval\_split\_size\ parameters when loading dataset samples.

_recipes/ljspeech/multiband\melgan, recipes/ljspeech/tacotron2-DDC · high confidence

New VCTK GlowTTS training recipe

A new training script for the GlowTTS model on the VCTK dataset has been added, enabling multi-speaker text-to-speech training. The recipe automatically downloads the VCTK dataset if missing, configures audio processing with external resampling, and sets up a speaker manager to handle multiple speakers. It utilizes the updated TrainerArgs API and includes specific evaluation split configurations to support model validation during training.

_recipes/ljspeech/glow\_tts, recipes/vctk/glow\tts · high confidence

New XTTS Fine-tuning Gradio Demo

A new interactive Gradio-based demo has been added to the TTS demos directory, enabling users to fine-tune the XTTS model. This tool provides a workflow for data processing (including audio formatting and transcription via Whisper), model training configuration, and inference, allowing users to customize and train voice clones directly through a web interface.

TTS/demos · high confidence

New and updated TTS tutorial and utility notebooks

The notebooks directory now includes several new and updated Jupyter notebooks to guide users through various TTS workflows. A new Tortoise notebook demonstrates inference using the Tortoise model with preset quality modes and voice cloning capabilities. Tutorial notebooks have been added and refreshed: Tutorial 1 covers quick inference with pre-trained models (including multi-speaker selection), while Tutorial 2 provides a complete workflow for training a GlowTTS model on the LJSpeech dataset. Additionally, new utility notebooks allow users to extract mel-spectrograms from trained TTS models for vocoder training, test attention alignment performance on specific sentences, and interactively plot speaker embeddings from the LibriTTS corpus using UMAP and Bokeh.

notebooks · high confidence

New dataset analysis notebooks for spectrogram, SNR, and pitch checks

The \notebooks/dataset\_analysis\ directory now includes three new interactive notebooks: \AnalyzeDataset.ipynb\ for inspecting dataset metadata and audio file integrity, \CheckDatasetSNR.ipynb\ for computing average Signal-to-Noise Ratio using WADA and FFMPEG, and \CheckPitch.ipynb\ for visualizing pitch contours alongside mel-spectrograms. These tools provide users with dedicated workflows for validating audio quality and dataset structure before model training.

_notebooks/dataset\analysis · high confidence

New development Dockerfile for local TTS setup

A new Dockerfile.dev has been added to streamline local development. It is based on nvidia/cuda:11.8.0 and installs OS dependencies, Python tools, and major Python packages like torch and torchaudio. It also copies and installs project-specific requirements and the TTS package itself, providing a ready-to-use environment for developers.

dockerfiles · high confidence

New encoder training utilities and data preparation tools

This change introduces a new set of utility modules for the speaker encoder subsystem. It adds \generic\_utils.py\ to handle audio augmentation (additive noise and reverberation) and model instantiation for LSTM and ResNet architectures, \training.py\ to streamline the training loop setup and argument processing, \visual.py\ to provide UMAP embedding visualization, and \prepare\_voxceleb.py\ to automate the downloading, decoding, and CSV generation for the VoxCeleb dataset.

TTS/encoder/utils · high confidence

New feed-forward encoder and decoder layers for Speedy Speech and AlignTTS models

Added a new \feed\_forward\ module containing specialized encoder and decoder layers to support Speedy Speech and AlignTTS architectures. This includes factory classes for \Encoder\ and \Decoder\ that allow selecting between residual convolutional blocks (\residual\_conv\_bn\) and transformer-based variants (\relative\_position\_transformer\ and \fftransformer\). A new \DurationPredictor\ layer is also included to predict phoneme durations from encoder outputs, enabling more efficient text-to-speech synthesis pipelines.

_TTS/tts/layers/feed\forward · high confidence

New generic neural network layers for TTS models

The TTS library now includes a new \TTS/tts/layers/generic\ module providing reusable building blocks for speech synthesis models. This addition introduces an Alignment Network for learning text-to-speech alignment, various normalization layers (LayerNorm, ActNorm, TemporalBatchNorm1d), and positional encoding. It also adds specialized convolutional blocks including GatedConv, ResidualConv1dBN, and TimeDepthSeparableConv, alongside a full WaveNet implementation with weight normalization support and a Transformer-based FFTransformer block with duration prediction capabilities.

TTS/tts/layers/generic · high confidence

New speaker encoder models (LSTM and ResNet) with unified base class

The speaker encoder module now includes new \LSTMSpeakerEncoder\ and \ResNetSpeakerEncoder\ implementations, both inheriting from a newly introduced \BaseEncoder\. This base class standardizes common functionality such as audio preprocessing (including a switch to \torchaudio.transforms.MelSpectrogram\), embedding computation, and model checkpoint loading. The new encoder architectures provide alternative backends for generating speaker embeddings, with the ResNet model supporting both SAP and ASP pooling methods.

TTS/encoder/models · high confidence

New utility modules for distribution sampling and spectrogram interpolation

Added new utility modules in the vocoder utils package: distribution.py provides functions for Gaussian loss, sampling from Gaussian distributions, numerically stable log-sum-exp, and discretized mix-logistic loss/sampling (used by models like WaveRNN and GlowTTS); generic\_utils.py adds interpolate\_vocoder\_input to match TTS and vocoder sample rates via spectrogram interpolation, and plot\_results for visualizing predicted vs. real waveforms and spectrograms.

TTS/vocoder/utils · high confidence

New vocoder layer implementations for HiFi-GAN, WaveGrad, MelGAN, and Parallel WaveGAN

The \TTS/vocoder/layers\ module has been populated with new PyTorch layer definitions for several state-of-the-art vocoder architectures. This includes the residual stack and multi-resolution frequency blocks for HiFi-GAN, the U-Net/D-Net blocks with FiLM conditioning for WaveGrad, the residual stacks for MelGAN, and the causal residual blocks for Parallel WaveGAN. Additionally, the module introduces location-variable convolution blocks (LVC) for LVC-GAN, multi-scale STFT and spectral loss functions for training, and utility layers for polyphase quadrature filter banks (PQMF) and upsampling networks. These components provide the foundational building blocks for training and inference in these specific vocoder models.

TTS/vocoder/layers · high confidence

Removals

Removal of LJSpeech dataset implementation

The LJSpeech dataset class has been removed from the codebase, eliminating the previous implementation that relied on pandas for CSV parsing and direct audio loading. This change removes the specific data handling logic for the LJSpeech corpus, likely as part of a broader refactoring or migration to a different data loading strategy.

datasets · high confidence

Removal of legacy audio, data, and utility modules

The \utils\ package has removed four files: \audio.py\ (containing STFT/mel-spectrogram processing and Griffin-Lim synthesis), \data.py\ (data padding and preparation helpers), \generic\_utils.py\ (experiment folder management, checkpoint saving, and progress bar logging), and \model.py\ (parameter counting). These legacy utilities are no longer part of the codebase, implying that audio processing, data handling, and training utilities have been moved to other modules or replaced by new implementations.

utils · high confidence

Removal of text processing and normalization utilities

The text processing module has been removed, deleting the \text\ package and its submodules (\\_\init\\_.py\, \cleaners.py\, \cmudict.py\, \numbers.py\, \symbols.py\). This eliminates the functionality for normalizing input text (including number and abbreviation expansion), transliterating non-English text to ASCII, and converting text to/from symbol ID sequences using the CMU Pronouncing Dictionary. Any components relying on these text cleaning and symbol mapping utilities will no longer function.

text · high confidence

Architecture

New TTS model abstraction and unified model setup

The TTS models directory has been refactored to introduce a unified \BaseTTS\ class that standardizes initialization, multi-speaker handling, and batch formatting for all TTS models. A new \setup\model\ function in \models/\\init\\_.py\ dynamically loads model implementations based on configuration, supporting a \base\_model\ field for inheritance. This change affects the core architecture of models like AlignTTS, Bark, BaseTacotron, ForwardTTS, GlowTTS, and NeuralHMM TTS, ensuring consistent interfaces for training and inference while simplifying model registration and loading.

TTS/tts/models · high confidence

Behavioural changes

GlowTTS layers refactored into modular components with new encoder and duration predictor options

The GlowTTS layer implementation has been restructured into distinct modules (decoder, encoder, duration predictor, and transformer) to improve maintainability and extensibility. The encoder now supports multiple architectures, including a gated convolution, a residual convolution with batch normalization, and a time-depth separable convolution, in addition to the existing relative position transformer. The duration predictor has been updated to accept an optional language embedding dimension, enabling better multilingual support. Additionally, the decoder's invertible convolution initialization now uses \torch.linalg.qr\ for PyTorch versions 1.9 and above, ensuring compatibility and numerical stability.

_TTS/tts/layers/glow\tts · high confidence

Introduce centralized loss functions module with masking and normalization support

The TTS training pipeline now uses a dedicated \TTS.tts.layers.losses\ module for all loss computations, replacing scattered implementations. This change introduces masked versions of L1, MSE, and BCE losses that correctly handle variable sequence lengths via \sequence\_mask\, ensuring that padding tokens do not contribute to the gradient. It also adds SSIM loss with sample-wise min-max normalization for audio quality metrics, attention entropy loss for alignment regularization, and utility functions like \sample\_wise\_min\_max\. This centralization ensures consistent loss masking behavior across models like GlowTTS, Tacotron, and VITS, fixing previous issues where loss computation ignored sequence lengths or produced NaN values due to unmasked padding.

TTS/tts/layers · high confidence

New text processing pipeline with tokenizer and cleaner modules

The text processing subsystem has been restructured into dedicated modules: \tokenizer.py\ introduces the \TTSTokenizer\ class for converting text to token IDs and back, supporting optional phonemization and blank/BOS/EOS padding; \characters.py\ defines \BaseVocabulary\ and \BaseCharacters\ classes to manage character sets, phonemes, and special tokens with configurable initialization from config; \cleaners.py\ provides language-specific text cleaning pipelines (English, French, Portuguese, Chinese Mandarin, multilingual, etc.) using \anyascii\ for transliteration; \punctuation.py\ adds a \Punctuation\ class to strip and restore punctuation marks during processing; and \cmudict.py\ integrates a CMU Pronouncing Dictionary wrapper for ARPAbet phoneme lookups. These changes collectively replace previous ad-hoc text handling with a modular, configurable, and extensible text processing stack.

TTS/tts/utils/text · high confidence

Project rebranding and governance overhaul

The project has been rebranded to Coqui TTS, introducing a formal governance structure via CODE\_OWNERS.rst and a Contributor Covenant Code of Conduct. This change includes the addition of a CITATION.cff file for academic referencing, a new Dockerfile for containerized development, and a Makefile to streamline testing, linting, and installation. The repository also adopts pre-commit hooks for code formatting (black, isort) and linting (pylint), while removing legacy files like module.py and updating the setup configuration to enforce Python 3.9+ compatibility.

(repo-wide) · high confidence

Refactored dataset loading and introduced new dataset formatters

The dataset loading pipeline has been refactored to support multiple datasets in a single configuration, allowing users to merge and train on combined data sources. The \load\_meta\_data\ function has been renamed to \load\_tts\_samples\ and now accepts a list of dataset dictionaries, applying specific formatters to each. New formatters have been added for the CML-TTS, Coqui, and M-AILABS datasets, enabling support for these specific data structures. Additionally, the \TTSDataset\ class now supports computing and caching f0 and energy features, and allows ignoring specific speakers during training.

TTS/tts/datasets · high confidence

Replace fairseq dependency with a custom HuBERT implementation for Bark TTS

The Bark text-to-speech module no longer requires the heavy fairseq library for HuBERT processing. Instead, it uses a new, self-contained implementation in \TTS/tts/layers/bark/hubert\ that leverages \transformers\ to load the HuBERT model, along with custom modules for tokenization and K-means clustering. This change simplifies the dependency chain and improves the ease of installation for users generating voice clones with Bark.

TTS/tts/layers/bark/hubert · high confidence

Standardized TTS model configuration files

The TTS configuration system has been refactored to use dedicated, structured config classes for each supported model (including AlignTTS, Bark, DelightfulTTS, FastPitch, FastSpeech, FastSpeech2, GlowTTS, NeuralHMM, and Overflow). These new dataclass-based configs replace the previous flat or inconsistent structures, providing explicit type hints, default values, and dedicated sections for model-specific parameters, multi-speaker settings, optimizer/learning rate schedules, loss weights, and testing sentences. This change ensures that users can now instantiate and modify model settings through a consistent, type-safe API, while also enabling better validation and documentation of available configuration options for each model.

TTS/tts/configs · high confidence

TTS server UI now supports multi-speaker, multi-language, and model details views

The TTS server's web interface has been updated to allow users to select specific speakers and languages via dropdown menus when the underlying model supports them. Additionally, a new 'Model Details' page is available (accessible via a button on the main page when enabled) that displays CLI arguments, model configuration, and vocoder configuration for transparency. The main page also includes support for style WAV inputs for GST-based models.

TTS/server/templates · high confidence

Tacotron models refactored to PyTorch with new attention and style layers

The Tacotron layer implementations have been rewritten in PyTorch, introducing support for Graves Attention and Global Style Tokens (GST) for prosody control. The decoder now includes a configurable \max\_decoder\_steps\ limit to prevent infinite generation loops, and the prenet module allows optional dropout during inference to improve quality. Additionally, the \CapacitronVAE\ layer has been added to enable variational embedding for prosody transfer.

TTS/tts/layers/tacotron · high confidence

Unified configuration system with Coqpit and backward-compatible loading

The TTS configuration system has been refactored to use the Coqpit library for structured, validated config classes (e.g., BaseAudioConfig, BaseDatasetConfig) and introduces a centralized loading mechanism in TTS/config/\_\init\\_.py. Users can now load configs via load\_config(), which automatically detects the model type and instantiates the correct Coqpit class, supporting both JSON and YAML files. To ensure backward compatibility, the system includes read\_json\_with\_comments() to handle legacy JSON files containing comments, and helper functions like get\_from\_config\_or\_model\_args() to bridge differences between models using model\_args versus direct config fields.

TTS/config · high confidence

Unified vocoder dataset infrastructure with model-specific loaders

The \TTS/vocoder/datasets\ module has been restructured to provide a centralized \setup\_dataset\ factory that instantiates the correct data loader based on the configured model type. This change introduces dedicated dataset classes for GAN-based vocoders (\GANDataset\), WaveGrad (\WaveGradDataset\), and WaveRNN (\WaveRNNDataset\), each handling specific requirements such as segment selection, feature caching, and noise augmentation. A new \preprocess\ module supports WaveRNN by enabling on-the-fly or precomputed feature extraction (mel spectrograms and quantized signals). This refactoring ensures that each vocoder architecture receives data in the format it expects, improving training stability and configuration clarity.

TTS/vocoder/datasets · high confidence

Updated Japanese phonemizer implementation

The Japanese text processing module has been updated with a new phonemizer implementation that converts Japanese text to phonemes compatible with the Julius segmentation kit. This change includes updated conversion rules for various Japanese character combinations and ensures proper handling of MeCab dependencies for text analysis.

TTS/tts/utils/text/japanese · high confidence

Updated LJSpeech Tacotron2-DCA recipe with new training arguments and evaluation split handling

The LJSpeech Tacotron2-DCA training script has been updated to use the new \TrainerArgs\ class instead of the previous \TrainingArgs\, and now explicitly passes \eval\_split\ and \eval\_split\_size\ parameters when loading TTS samples to control evaluation data partitioning.

recipes/ljspeech/tacotron2-DCA · high confidence

Vocoder configurations migrated to structured dataclasses

The TTS vocoder configuration system has been refactored to use Python dataclasses, replacing the previous configuration format. This change introduces dedicated configuration classes for each supported vocoder model—including HiFi-GAN, MelGAN, MultiBand MelGAN, Parallel WaveGAN, UnivNet, WaveGrad, and WaveRNN—along with shared base classes for common parameters. Users can now instantiate and inspect vocoder settings using standard Python objects (e.g., \HifiganConfig()\), which provides better type safety, clearer documentation of default values, and easier customization of training and inference parameters.

TTS/vocoder/configs · high confidence

Vocoder model architecture refactoring and new model support

The TTS vocoder models have been restructured to support a unified GAN-based training framework and new model architectures. A new \GAN\ base class now wraps generator and discriminator networks, allowing for flexible mixing of components like HiFi-GAN, MelGAN, and Parallel WaveGAN. The diff introduces implementations for HiFi-GAN (generator and discriminator), Fullband MelGAN, Multiband MelGAN, and Parallel WaveGAN (generator and discriminators), along with a \BaseVocoder\ class for standardizing model arguments. Model instantiation is now handled via \setup\_model\, \setup\_generator\, and \setup\discriminator\ functions in \\\init\\_.py\, which dynamically load classes based on configuration.

TTS/vocoder/models · high confidence

Test coverage

2 commits adding/updating tests in tests/data/ljspeech/f0\_cache; Add comprehensive test suite for TTS data loading and sampling; Add small LJSpeech test fixtures for multi-format and multi-speaker validation; Added auxiliary test suite for audio, embeddings, and speaker models; Added bash-based tests for compute statistics and demo server; Added comprehensive test suite for vocoder training and components; Added dummy speaker data for testing; Added inference tests for CLI synthesis and Synthesizer sentence splitting; Added integration and unit tests for TTS models; Added test coverage for text processing components; Added test infrastructure and utility helpers; Added test input fixtures for Common Voice and model configurations; Added tests for FreeVC, XTTS streaming, and model zoo integration; Added unit and integration tests for TTS models and utilities; Added unit tests for XTTS GPT training workflows.

Dependencies

Modernize core dependencies and introduce modular requirement files

The project has significantly updated its core dependencies to support modern Python versions (up to 3.11) and newer library versions, including upgrading NumPy, SciPy, Librosa, and PyTorch (\>=2.1). To improve maintainability and reduce installation overhead, dependencies are now split into modular files: \requirements.dev.txt\ for development tools, \requirements.ja.txt\ for Japanese language support, \requirements.notebooks.txt\ for visualization, and specific files for the XTTS fine-tuning demo and encoder. The main \requirements.txt\ now explicitly lists version constraints for core libraries like \torch\, \torchaudio\, and \gruut\, while removing older or unnecessary constraints.

(dependencies) · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

Score

  • CAI 43 → 48 (+5.3)
  • Rubric changed (rubric-2026.08.15 → rubric-2026.09.15) — scores are not directly comparable.

Lenses

  • Code Health 84 → 82 (-1.4)
  • Architecture 94 → 96 (+1.5)
  • Maturity 58 → 58 (+0.1)
  • Readiness 26 → 40 (+14.0)
  • Security 46 → 65 (+18.3)
  • Accessibility 44 (new)

Resolved (89)

  • Coverage not measured — test suite did not build
  • Dimension evaluation failed
  • Duplicated block (10 lines × 2) (TTS/tts/layers/generic/aligner.py)
  • Duplicated block (10 lines × 2) (tests/aux_tests/test_audio_processor.py)
  • Duplicated block (10 lines × 2) (tests/tts_tests/test_vits.py)
  • Duplicated block (11 lines × 2) (TTS/vc/modules/freevc/wavlm/modules.py)
  • Duplicated block (12 lines × 2) (TTS/tts/layers/tortoise/xtransformers.py)
  • Duplicated block (12 lines × 2) (TTS/tts/models/vits.py)
  • Duplicated block (13 lines × 2) (TTS/tts/layers/tortoise/diffusion.py)
  • Duplicated block (13 lines × 2) (TTS/utils/manage.py)
  • Duplicated block (13 lines × 2) (TTS/utils/synthesizer.py)
  • Duplicated block (13 lines × 2) (tests/data_tests/test_loader.py)
  • Duplicated block (13 lines × 4) (tests/tts_tests/test_tacotron_model.py)
  • Duplicated block (14 lines × 2) (TTS/tts/layers/delightful_tts/encoders.py)
  • Duplicated block (14 lines × 2) (TTS/tts/models/xtts.py)
  • Duplicated block (14 lines × 2) (tests/tts_tests/test_tacotron2_model.py)
  • Duplicated block (15 lines × 2) (TTS/tts/layers/xtts/stream_generator.py)
  • Duplicated block (16 lines × 3) (TTS/tts/layers/xtts/stream_generator.py)
  • Duplicated block (20 lines × 2) (TTS/tts/layers/xtts/stream_generator.py)
  • Duplicated block (27 lines × 2) (TTS/tts/models/glow_tts.py)
  • …and 69 more

New (558)

  • Attention.forward (cognitive 29) (TTS/tts/layers/tortoise/xtransformers.py)
  • Attention.forward (cyclomatic 27) (TTS/tts/layers/tortoise/xtransformers.py)
  • AttentionLayers.init (cognitive 53) (TTS/tts/layers/tortoise/xtransformers.py)
  • AttentionLayers.init (cyclomatic 42) (TTS/tts/layers/tortoise/xtransformers.py)
  • AttentionLayers.forward (cognitive 50) (TTS/tts/layers/tortoise/xtransformers.py)
  • AttentionLayers.forward (cyclomatic 34) (TTS/tts/layers/tortoise/xtransformers.py)
  • AudioProcessor.denormalize (cognitive 17) (TTS/utils/audio/processor.py)
  • AudioProcessor.normalize (cognitive 17) (TTS/utils/audio/processor.py)
  • AugmentWAV.init (cognitive 24) (TTS/encoder/utils/generic_utils.py)
  • Banned license: mutagen
  • Banned license: unidecode
  • BaseTTS.get_aux_input_from_test_sentences (cognitive 19) (TTS/tts/models/base_tts.py)
  • BaseTTS.get_data_loader (cognitive 51) (TTS/tts/models/base_tts.py)
  • BaseTTS.get_data_loader (cyclomatic 21) (TTS/tts/models/base_tts.py)
  • BaseTTS.get_sampler (cognitive 17) (TTS/tts/models/base_tts.py)
  • BaseVC.get_aux_input_from_test_sentences (cognitive 19) (TTS/vc/models/base_vc.py)
  • BaseVC.get_data_loader (cognitive 51) (TTS/vc/models/base_vc.py)
  • BaseVC.get_data_loader (cyclomatic 21) (TTS/vc/models/base_vc.py)
  • BaseVC.get_sampler (cognitive 17) (TTS/vc/models/base_vc.py)
  • BaseVocoder._set_model_args (cognitive 22) (TTS/vocoder/models/base_vocoder.py)
  • …and 538 more

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

coqui-ai/TTS was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 26 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit dbf1a08a0d4e47fdad6172e433eeb34bc6b13b4e — the exact code this score is about.
  • Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-09659c52afae.