Skip to content
CAI
Software that uses CAICheck a score

RVC-Boss/GPT-SoVITS

41.2

Weak · 26 September 2026

36.7k

lines of production code

Python

primary language

3

measurements over time

CAI band scale
CAI trend line
CAI lens gauges

What this system is

This system is a comprehensive text-to-speech and voice synthesis platform, primarily built around the GPT-SoVITS architecture with support for multiple model variants including v2, F5-TTS, and BigVGAN. It provides a full pipeline for high-fidelity audio generation, featuring advanced text processing with multi-lingual support (Chinese, English, Japanese, Korean, Cantonese) and robust data preparation tools. The system also integrates auxiliary capabilities such as automatic speech recognition, vocal separation, and audio super-resolution to enhance input quality and output fidelity.

How it got here

2024 — GPT-SoVITS v2 architecture and tooling

18 changes.

This period focused on establishing the foundational infrastructure for GPT-SoVITS v2, including a complete rewrite of the model architecture, data preparation pipelines, and text processing modules to support new languages and improved pronunciation accuracy. It also introduced comprehensive tooling for audio processing, vocal separation, and multi-backend ASR, alongside deployment optimizations like ONNX export and CUDA Graph acceleration.

2025 — F5-TTS and BigVGAN integration

11 changes.

This period focused on integrating the F5-TTS model architecture and the BigVGAN v2 neural vocoder into the GPT-SoVITS framework, introducing new DiT backbones and high-fidelity audio synthesis capabilities. The work also expanded the project's tooling with audio super-resolution models and improved text processing for mixed-language inputs, while updating the Docker environment to support modern Python and CUDA versions.

Features

Add Bs\_Roformer and MelBand\_Roformer models to UVR5

New source files for the Bs\_Roformer and MelBand\_Roformer models have been added to the UVR5 module, introducing transformer-based architectures with flash attention support and rotary embeddings for audio separation.

_tools/uvr5/bs\roformer · high confidence

Add Chinese text normalization for non-standard words

The \GPT\_SoVITS/text/zh\_normalization\ module has been added to handle Chinese Non-Standard Word (NSW) normalization. This feature converts numeric and symbolic representations into their spoken Chinese equivalents, including dates, times, phone numbers, temperatures, fractions, percentages, and monetary values. It also includes utilities for converting between traditional and simplified Chinese characters and handling full-width/half-width character conversions, ensuring the text input is properly verbalized for the speech synthesis pipeline.

_GPT\_SoVITS/text/zh\normalization · high confidence

Add DiT backbone with gradient checkpointing and caching support

The model backbones now include a new DiT (Diffusion Transformer) implementation that supports gradient checkpointing to reduce VRAM usage during training and inference. This backbone introduces text and audio conditioning embeddings, rotary position embeddings, and optional long-skip connections. It also adds support for caching text and timestep embeddings during inference to improve efficiency, alongside configurable ConvNeXtV2 blocks for text processing.

_GPT\_SoVITS/f5\tts/model/backbones · high confidence

Add English text normalization for numbers, time, and currency

Introduces a new English text normalization module (expend.py) that converts written-out numeric formats into spoken-style text. This includes handling of measurements (e.g., km, °C), time (24-hour to 12-hour format), currency (dollars, pounds), decimals, fractions, and ordinals, enabling more accurate text-to-speech processing for these specific patterns.

_GPT\_SoVITS/text/en\normalization · high confidence

Added BigVGAN v2 neural vocoder submodule

The GPT\_SoVITS/BigVGAN directory now contains the NVIDIA BigVGAN v2 source code, including the main model implementation (bigvgan.py), Snake/SnakeBeta activation functions, and utility modules. This adds a new neural vocoding capability to the project, allowing for high-fidelity audio synthesis using the BigVGAN architecture with support for custom CUDA kernels and Hugging Face Hub integration.

_GPT\SoVITS/BigVGAN · high confidence

Added ERes2NetV2 speech feature extraction components

The eres2net module now includes the ERes2NetV2 architecture (ERes2NetV2.py) and its supporting fusion logic (fusion.py), introducing an improved model with reduced parameters and computational cost for short-duration feature extraction. Additionally, new utility modules were added: kaldi.py provides Kaldi-compatible audio preprocessing (windowing, mel filters, fbank), pooling\_layers.py implements temporal pooling strategies (TAP, TSDP, TSTP, ASTP) for feature aggregation, and the existing install script was fixed to reduce log noise and improve error reporting.

_GPT\SoVITS/eres2net · high confidence

Added alias-free activation components for BigVGAN

The BigVGAN module now includes the underlying PyTorch implementation for alias-free activation, providing 1D upsampling, downsampling, and low-pass filtering capabilities. This addition enables the model to perform activation operations with anti-aliasing, which helps reduce spectral artifacts in generated audio.

_GPT\_SoVITS/BigVGAN/alias\_free\activation · high confidence

CUDA Graph acceleration and ONNX export support for text-to-semantic model

The text-to-semantic inference pipeline now supports CUDA Graph acceleration, which doubles inference speed without changing output quality, and introduces ONNX module support for broader deployment compatibility. These changes are implemented through new CUDA-graph-optimized embedding and attention structures, alongside dedicated ONNX model variants and their corresponding Lightning training modules, ensuring the core model logic remains consistent across both accelerated and exportable formats.

_GPT\SoVITS/AR/models · high confidence

Initial project scaffolding and deployment configuration

This change introduces the foundational configuration files required to build, run, and maintain the GPT-SoVITS application. It adds \.gitignore\ and \.dockerignore\ to exclude build artifacts, caches, and large model weights from version control and container images. A \Dockerfile\ and \docker-compose.yaml\ are provided to containerize the application with support for CUDA 12.6/12.8 and a 'Lite' mode that excludes heavy ASR/UVR5 models. Cross-platform installation is enabled via \install.sh\ (Linux/macOS) and \install.ps1\ (Windows), alongside convenience launch scripts (\go-webui.bat\, \go-webui.ps1\). The repository also includes \api.py\ and \api\_v2.py\ for programmatic text-to-speech access, \config.py\ for managing model paths and device detection, and Colab notebooks (\Colab-Inference.ipynb\, \Colab-WebUI.ipynb\) for cloud-based usage.

(repo-wide) · high confidence

Introduce APNet BWE model for 24kHz to 48kHz audio super-resolution

The \tools/AP\_BWE\main/models\ directory now contains the implementation of the APNet BWE model, which enables up-sampling audio from 24kHz to 48kHz. This change adds the core model architecture (\model.py\), including ConvNeXt blocks for feature extraction and multi-period discriminators for adversarial training, alongside the module initialization file (\\\init\\_.py\).

_tools/AP\_BWE\main/models · high confidence

Introduce G2PW-based polyphonic character pronunciation for Chinese text

Added a new G2PW (Grapheme-to-Phoneme with Polyphonic characters) module to the text processing pipeline to improve the accuracy of Chinese pronunciation, specifically for polyphonic characters (characters with multiple pronunciations depending on context). This change integrates an ONNX-based inference model (G2PWOnnxConverter) that automatically downloads and caches the model, handles context-aware pronunciation selection, and falls back to standard pypinyin logic when the model does not support a specific character. It also includes a polyphonic dictionary (polyphonic.rep and polyphonic-fix.rep) for correcting specific words and idioms, and optimizes inference input construction to reduce computational overhead for long sentences.

_GPT\SoVITS/text/g2pw · high confidence

Introduces AR modules with ONNX inference support

Adds a new set of modules in the AR (Autoregressive) component, including attention, embedding, transformer, and optimizer implementations. Crucially, this change introduces ONNX-specific variants (e.g., \activation\_onnx.py\, \transformer\_onnx.py\) alongside standard PyTorch versions, enabling the model to be exported and run via ONNX runtime for improved inference compatibility and performance.

_GPT\SoVITS/AR/modules · high confidence

Introduction of F5-TTS model architecture components

The GPT\_SoVITS module now includes the foundational building blocks for the F5-TTS model. This change adds the \DiT\ backbone import and defines core neural network layers in \modules.py\, including sinusoidal and convolutional position embeddings, rotary positional embeddings, Global Response Normalization (GRN), ConvNeXt-V2 blocks, and AdaLayerNormZero variants. These components provide the necessary infrastructure for the new text-to-speech model implementation.

_GPT\_SoVITS/f5\tts/model · high confidence

Introduction of modular TTS inference pipeline with configurable text segmentation

The \GPT\_SoVITS/TTS\_infer\_pack\ directory now contains a new, structured TTS inference implementation (\TTS.py\) alongside a dedicated text preprocessor (\TextPreprocessor.py\) and a pluggable text segmentation module (\text\_segmentation\_method.py\). This change introduces a registry-based system for text splitting strategies (such as 'cut0' through 'cut5'), allowing users to select different algorithms for segmenting input text before synthesis. The new architecture also includes specific mel-spectrogram processing functions for different model versions (v2 vs v4) and integrates language detection via \LangSegmenter\ to handle multi-lingual inputs more robustly.

_GPT\_SoVITS/TTS\_infer\pack · high confidence

New ASR tooling with multi-backend support and automatic model downloading

The \tools/asr\ directory now includes a new configuration and execution framework for Automatic Speech Recognition. Users can choose between three backends: Fun-ASR-Nano, SenseVoice, and Faster Whisper, each supporting different language sets (including Cantonese) and precision levels. The system automatically downloads the required models from HuggingFace or ModelScope on first use, and Faster Whisper includes a fallback to CPU if CUDA compilation fails.

tools/asr · high confidence

New audio processing and utility tools

The \tools\ directory now includes a suite of new modules for audio handling and dataset management. \audio\_sr.py\ adds 24kHz-to-48kHz audio super-resolution capabilities, while \cmd-denoise.py\ provides a command-line interface for noise suppression using a ModelScope pipeline. \slicer2.py\ and \slice\_audio.py\ introduce a new audio slicing engine that segments audio based on silence detection and volume thresholds. Additionally, \subfix\_webui.py\ offers a Gradio-based interface for batch editing and splitting audio-text pairs, and \my\_utils.py\ centralizes helper functions for path cleaning, file existence checks, and platform-specific CUDA library loading.

tools · high confidence

New dataset module supporting 24kHz to 48kHz audio super-resolution

A new dataset module has been added to the \tools/AP\_BWE\_main/datasets1\ package, introducing a PyTorch \Dataset\ class designed for audio super-resolution tasks. This implementation enables processing audio at a 24kHz sampling rate and upscaling it to 48kHz, handling both high-resolution (HR) and low-resolution (LR) audio segments. The module includes utilities for STFT/ISTFT transformations and manages file lists for training and validation, allowing users to load and preprocess audio pairs for super-resolution models.

_tools/AP\_BWE\main/datasets1 · high confidence

New feature extractor module for CN-HuBERT and Whisper

A new \feature\_extractor\ package has been added to GPT\_SoVITS, providing a unified interface for extracting audio features using either the CN-HuBERT or Whisper models. The module exposes a \content\_module\_map\ to select between the two backends, with \cnhubert.py\ implementing a wrapper around Hugging Face's \HubertModel\ and \Wav2Vec2FeatureExtractor\, and \whisper\_enc.py\ implementing a wrapper around the Whisper encoder to process 16kHz audio inputs.

_GPT\_SoVITS/feature\extractor · high confidence

New i18n scanning and management tooling

A new i18n tooling suite has been added to the \tools/i18n\ directory, introducing \i18n.py\ for runtime language loading and fallback, and \scan\_i18n.py\ for static analysis. The scanner extracts internationalization keys from Python source files using AST parsing and automatically updates JSON locale files by adding missing keys (marked with \\#!\ for untranslated strings) and removing unused ones, while also detecting duplicate translation values.

tools/i18n · high confidence

New inference entry points and model export capabilities

The GPT\_SoVITS module now includes dedicated scripts for command-line and GUI-based inference, alongside new tools to export models to TorchScript and ONNX formats. The new \inference\_cli.py\ provides a programmatic interface for synthesis, while \inference\_gui.py\ offers a standalone PyQt5 application for users who prefer not to use the web UI. Additionally, \export\_torch\_script.py\ and \export\_torch\_script\_v3v4.py\ enable exporting models for use in non-Python environments, and \onnx\_export.py\ facilitates ONNX conversion. These changes expand how users can integrate and deploy the voice synthesis engine outside the standard web interface.

_GPT\SoVITS · high confidence

UVR5 module adds BS-Roformer and Mel-Band-Roformer vocal separation models

The UVR5 tool now supports vocal separation using the BS-Roformer and Mel-Band-Roformer architectures. This change introduces new model implementations (\bsroformer.py\, \nets\_new.py\, \layers\_new.py\) and configuration files (\modelparams/4band\_v2.json\, \4band\_v3.json\) alongside the existing MDXNet and VR models. Users can now select these newer models in the UVR5 web UI to separate vocals from music, with the system automatically handling model loading and configuration.

tools/uvr5 · high confidence

Behavioural changes

Added third-party license files for BigVGAN dependencies

The GPT\_SoVITS/BigVGAN/incl\_licenses directory now includes eight new license files (LICENSE\_1 through LICENSE\_8) covering MIT, Apache 2.0, and BSD 3-Clause terms for various contributors and projects (including Jungil Kong, Edward Dixon, Seungwon Park, Alexandre Défossez, Descript, Charactr Inc., and Amphion). This ensures compliance with the licensing requirements of the underlying BigVGAN components and their dependencies.

_GPT\_SoVITS/BigVGAN/incl\licenses · high confidence

Docker environment now uses Miniforge and installs Python 3.12 with CUDA 12.6/12.8 support

The Docker build process has been updated to use Miniforge instead of the previous conda distribution. This change introduces support for Python 3.12 and allows the environment to be built for either CUDA 12.6 or 12.8 based on the build configuration. The installation scripts now handle platform-specific sysroot requirements (linux/amd64 vs linux/arm64) and install additional dependencies such as gcc, ffmpeg, cmake, and specific Python packages like torch, torchcodec, and flash-attn.

Docker · high confidence

Improved language segmentation for mixed CJK and short English text

The LangSegmenter module has been updated to better handle mixed-language inputs, specifically addressing issues with short English segments being misidentified and CJK characters being incorrectly grouped. The new implementation uses a custom regex-based approach to detect full English text and CJK characters, ensuring that short English phrases are correctly tagged as 'en' and CJK text is properly separated. This results in more accurate language detection for Chinese, Japanese, Korean, and English text, especially in mixed-language scenarios.

_GPT\SoVITS/text/LangSegmenter · high confidence

Introduces GPT-SoVITS v2 module architecture with v2Pro/v2ProPlus support

The \GPT\_SoVITS/module\ package has been replaced with a new implementation supporting the GPT-SoVITS v2 model family, including v2Pro and v2ProPlus variants. This update introduces new attention mechanisms (\attentions.py\, \attentions\_onnx.py\) and a data loading pipeline (\data\_utils.py\) that explicitly handles v2Pro-specific speaker embedding paths. The module also includes a new vector quantization implementation (\core\_vq.py\) and distributed training utilities (\ddp\_utils.py\, \distrib.py\) to support multi-GPU training and ONNX export for the v2 architecture.

_GPT\SoVITS/module · high confidence

New AR data pipeline with length-aware batching and DPO batch-size adjustment

The AR data module now uses a new \Text2SemanticDataModule\ that loads phoneme and semantic data via a \Text2SemanticDataset\ with validation (file existence checks, length/phoneme-ratio filtering, and error handling for missing entries). Training batches are distributed across GPUs using a new \DistributedBucketSampler\ that groups samples by input length to reduce padding. When DPO training is enabled (via the \if\_dpo\ config flag), the training batch size is automatically halved to accommodate the larger memory requirements of DPO.

_GPT\SoVITS/AR/data · high confidence

New dataset preparation scripts for GPT-SoVITS v2

The \GPT\_SoVITS/prepare\_datasets\ directory now contains four new Python scripts (\1-get-text.py\, \2-get-hubert-wav32k.py\, \2-get-sv.py\, \3-get-semantic.py\) that implement the data processing pipeline for GPT-SoVITS v2. These scripts handle text cleaning and BERT feature extraction, HuBERT audio feature extraction with NaN filtering, speaker embedding computation using ERes2NetV2, and semantic token extraction via the v2/v3 synthesizer model, replacing the previous v1 preparation workflow.

_GPT\_SoVITS/prepare\datasets · high confidence

Text processing module restructured for GPT-SoVITS v2 with Cantonese and Korean support

The text processing module has been reorganized to support GPT-SoVITS v2, introducing a version-aware pipeline that switches between v1 and v2 symbol sets and language handlers. This update adds native text-to-phoneme support for Cantonese (using ToJyutping) and Korean (using g2pk2), alongside an updated Chinese handler (chinese2) that integrates G2PW for improved polyphone resolution. The cleaner module now routes languages to their specific normalizers and G2P engines, and new dictionary files (CMUdict, engdict-hot) and a Japanese user dictionary are included to enhance pronunciation accuracy.

_GPT\SoVITS/text · high confidence

Dependencies

Update Python dependencies in requirements.txt

The project's Python dependencies have been updated to specific versions to ensure compatibility and stability. Key changes include pinning librosa to 0.10.2, restricting numpy to versions below 2.0, and setting gradio to versions below 5. New or updated packages include funasr (\>=1.3.7), fast\_langdetect (\>=0.3.1), fastapi\[standard\] (\>=0.115.2), and ctranslate2 (\>=4.0,\<5), alongside various other libraries like peft, transformers, and pydantic with specific version constraints.

(dependencies) · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

Score

  • CAI 41 → 41 (-0.2)
  • Rubric changed (rubric-2026.08.15 → rubric-2026.09.15) — scores are not directly comparable.

Lenses

  • Code Health 83 → 75 (-7.8)
  • Architecture 94 → 97 (+3.0)
  • Maturity 51 → 50 (-1.5)
  • Readiness 22 → 17 (-5.3)
  • Security 49 → 70 (+20.4)

Resolved (65)

  • Dimension evaluation failed
  • Duplicated block (10 lines × 2) (GPT_SoVITS/AR/models/t2s_model.py)
  • Duplicated block (11 lines × 2) (GPT_SoVITS/AR/models/t2s_model.py)
  • Duplicated block (11 lines × 3) (tools/uvr5/vr.py)
  • Duplicated block (13 lines × 2) (GPT_SoVITS/module/data_utils.py)
  • Duplicated block (13 lines × 2) (tools/uvr5/bsroformer.py)
  • Duplicated block (16 lines × 2) (GPT_SoVITS/AR/modules/activation.py)
  • Duplicated block (17 lines × 2) (GPT_SoVITS/module/models.py)
  • Duplicated block (17 lines × 2) (GPT_SoVITS/module/models_onnx.py)
  • Duplicated block (18 lines × 2) (GPT_SoVITS/BigVGAN/bigvgan.py)
  • Duplicated block (31 lines × 2) (GPT_SoVITS/text/tone_sandhi.py)
  • Duplicated block (31 lines × 2) (GPT_SoVITS/text/tone_sandhi.py)
  • Duplicated block (5 lines × 2) (GPT_SoVITS/eres2net/ERes2NetV2.py)
  • Duplicated block (5 lines × 2) (GPT_SoVITS/module/attentions.py)
  • Duplicated block (5 lines × 2) (GPT_SoVITS/module/models_onnx.py)
  • Duplicated block (6 lines × 2) (GPT_SoVITS/AR/models/t2s_model_onnx.py)
  • Duplicated block (6 lines × 2) (GPT_SoVITS/TTS_infer_pack/TTS.py)
  • Duplicated block (7 lines × 2) (GPT_SoVITS/AR/models/t2s_model.py)
  • Duplicated block (7 lines × 2) (GPT_SoVITS/export_torch_script_v3v4.py)
  • Duplicated block (7 lines × 2) (GPT_SoVITS/text/zh_normalization/text_normlization.py)
  • …and 45 more

New (458)

  • AudioPre._path_audio_ (cognitive 60) (tools/uvr5/vr.py)
  • AudioPre._path_audio_ (cyclomatic 26) (tools/uvr5/vr.py)
  • AudioPreDeEcho._path_audio_ (cognitive 53) (tools/uvr5/vr.py)
  • AudioPreDeEcho._path_audio_ (cyclomatic 23) (tools/uvr5/vr.py)
  • BSRoformer.forward (cognitive 47) (tools/uvr5/bs_roformer/bs_roformer.py)
  • BSRoformer.forward (cyclomatic 29) (tools/uvr5/bs_roformer/bs_roformer.py)
  • Banned license: frozendict
  • CUDAGraphRunner._handle_request (cognitive 40) (GPT_SoVITS/AR/models/t2s_model_cudagraph.py)
  • CUDAGraphRunner._handle_request (cyclomatic 20) (GPT_SoVITS/AR/models/t2s_model_cudagraph.py)
  • DDP.forward (cognitive 57) (GPT_SoVITS/module/ddp_utils.py)
  • DDP.forward (cyclomatic 29) (GPT_SoVITS/module/ddp_utils.py)
  • DistributedBucketSampler.init (cognitive 17) (GPT_SoVITS/AR/data/bucket_sampler.py)
  • Duplicated block (10 lines × 2) (GPT_SoVITS/AR/models/t2s_model_onnx.py)
  • Duplicated block (10 lines × 2) (GPT_SoVITS/inference_webui.py)
  • Duplicated block (10 lines × 2) (GPT_SoVITS/inference_webui_fast.py)
  • Duplicated block (10 lines × 2) (tools/uvr5/bs_roformer/bs_roformer.py)
  • Duplicated block (10 lines × 2) (tools/uvr5/bs_roformer/bs_roformer.py)
  • Duplicated block (10 lines × 3) (GPT_SoVITS/s2_train.py)
  • Duplicated block (10 lines × 4) (GPT_SoVITS/AR/models/t2s_model.py)
  • Duplicated block (10 lines × 4) (GPT_SoVITS/TTS_infer_pack/TTS.py)
  • …and 438 more

Changes since last survey

  • 1 commits — 0 feature/other, 1 fixes

By area

  • (root) — 1 commit

Notable commits

  • fix: Fix Fun-ASR-Nano Transformers requirement (#2824)

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

RVC-Boss/GPT-SoVITS was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 26 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit 48b1a0169a28582a8984402f82cf438d3bfa6aca — the exact code this score is about.
  • Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-09659c52afae.