Skip to content
CAI
Software that uses CAICheck a score

modelscope/FunASR

44.7

Weak · 4 October 2026

157.4k

lines of production code

Python

with C

1

measurement over time

CAI band scale
CAI lens gauges

What this system is

FunASR is an open-source Automatic Speech Recognition toolkit that provides a comprehensive suite of models for transcription, speaker verification, emotion recognition, and language identification. It supports diverse inference modes, including offline batch processing, real-time streaming, and large language model-based generation, with extensive runtime deployment options across Python, C++, ONNX, and various hardware accelerators. The system also includes utilities for text normalization, data loading, and evaluation, alongside a unified API for integrating speech capabilities into applications.

How it got here

2022–2023 — Runtime expansion and model architecture growth

79 changes.

This period focused on establishing the FunASR project's public structure and significantly expanding its runtime ecosystem with C++, Python, and mobile clients for diverse deployment scenarios. It also introduced a wide array of new model architectures, including Paraformer, Conformer, and various streaming and diarization models, alongside comprehensive training and evaluation utilities.

2024 — Model expansion and runtime integration

56 changes.

This period focused on significantly expanding the FunASR model library with new architectures for streaming ASR, keyword spotting, emotion recognition, and LLM-based speech understanding. It also introduced a unified AutoModel API and added runtime support through new C++, Go, and Python WebSocket clients to facilitate deployment and integration.

2025–2026 — Fun-ASR-Nano launch and ecosystem expansion

32 changes.

This period focused on introducing the high-throughput Fun-ASR-Nano model with vLLM support, alongside adding adapters for Qwen3-ASR, GLM-ASR, and Silero VAD. The team significantly expanded the project's ecosystem by releasing a standalone llama.cpp runtime, OpenAI-compatible and MCP server examples, and a comprehensive product website with bilingual documentation and deployment guides.

Features

Add ASR utility modules for the FunASR PyTorch runtime

The \runtime/python/libtorch/funasr\_torch/utils\ package now includes a suite of new utility modules to support the PyTorch-based FunASR runtime. These additions provide core functionality for audio processing (\frontend.py\ with Kaldi-based Fbank and LFR/CMVN), text tokenization (\sentencepiece\_tokenizer.py\, \utils.py\ with CharTokenizer and TokenIDConverter), post-processing and timestamping (\postprocess\_utils.py\, \timestamp\_utils.py\), and evaluation metrics (\compute\_wer.py\).

_runtime/python/libtorch/funasr\torch/utils · high confidence

Add AliFsmnVadSharp voice activity detection library and example

Introduces the AliFsmnVadSharp library for C\# (.NET 6.0), enabling users to detect speech segments in audio using the Alibaba FSMN-Monophone VAD ONNX model. The change includes the core implementation classes (AliFsmnVad, E2EVadModel, WavFrontend) and model configuration entities, along with a console example application that demonstrates loading WAV files, running inference, and outputting segment timestamps and real-time factor metrics.

runtime/csharp · high confidence

Add BAT model and ESPnetSV speaker verification model

Introduces the BAT (Boundary-Aware Transducer) model for low-latency streaming ASR with explicit boundary detection, and the ESPnetSV model for speaker verification using a CTC-attention hybrid architecture. The BAT model is registered under the 'BAT' key in the model registry and inherits from the Transducer base class, while the ESPnetSV model supports configurable CTC, Transducer, and attention decoder branches for speaker verification tasks.

funasr/models/bat · high confidence

Add Branchformer and E-Branchformer speech recognition models

Introduces two new parallel-branch encoder architectures for speech recognition: Branchformer and E-Branchformer. Branchformer combines self-attention and a Convolutional Gating MLP (cgMLP) using concatenation or learned averaging to capture local and global context, while E-Branchformer enhances this with an element-wise merging strategy via depth-wise convolution fusion for improved parameter efficiency. Both models are registered in the framework, include full encoder implementations with FastSelfAttention support, and come with configuration templates for training.

_funasr/models/branchformer, funasr/models/e\branchformer · high confidence

Add Branchformer and E-Branchformer training and inference recipes for AISHELL-1

New end-to-end recipes for Branchformer and E-Branchformer models on the AISHELL-1 dataset have been added to the examples/aishell directory. These include run.sh scripts for data preparation, feature extraction, dictionary creation, model training, and inference, along with README files documenting the training configurations (e.g., 80-dim fbank, speed perturbation, 180 epochs) and reported Character Error Rates (Branchformer: 4.51% test CER; E-Branchformer: 4.52% test CER). Demo scripts and utility symlinks are also provided to facilitate quick start and inference.

_examples/aishell/branchformer, examples/aishell/e\branchformer · high confidence

Add C++ HTTP server for offline ASR file transcription

Introduces a new C++ HTTP server (\funasr-http-server\) in \runtime/http\ that enables offline speech-to-text transcription via HTTP POST requests. The server accepts audio files (e.g., WAV) and returns JSON results containing the transcribed text, timestamps, and sentence-level details. It is built using Boost.Asio for networking and nlohmann/json for data handling, and supports configuration of ASR, VAD, and punctuation models via command-line arguments.

runtime/http · high confidence

Add CR-CTC loss and new loss modules to funasr.losses

The funasr.losses package now includes a new Consistency-Regularized CTC (CR-CTC) loss, which applies consistency regularization by computing the symmetric KL divergence between CTC outputs from augmented and clean encoder runs. Additionally, the label\_smoothing\_loss module introduces a SequenceBinaryCrossEntropy class for sequence-level binary classification tasks, alongside existing LabelSmoothingLoss and NllLoss implementations.

funasr/losses · high confidence

Add CT-Transformer punctuation restoration model with ONNX export support

Introduces the CT-Transformer model, a punctuation restoration component that adds commas, periods, and question marks to unpunctuated text (supporting Chinese and English) within the ASR pipeline. The implementation includes the core model class, utility functions for text splitting and language detection, and an export module that enables converting the model to ONNX format for deployment.

_funasr/models/ct\transformer · high confidence

Add Conformer ASR example for AIShell-1

Introduces a new Conformer-based Automatic Speech Recognition example for the AIShell-1 dataset. This includes a training and inference pipeline (run.sh) supporting multi-GPU training, feature extraction, and batch inference, along with documentation of the model configuration (80-dim fbank, speed perturbation, SpecAugment) and reported Character Error Rates (4.42% on dev, 4.87% on test). Demo scripts are provided as symlinks to the Paraformer equivalents.

examples/aishell/conformer, examples/aishell/paraformer · high confidence

Add Conformer model and normalization utilities

This change introduces the Conformer speech recognition model to the FunASR library, including its encoder architecture, main model class, and configuration template. It also adds GlobalMVN and UtteranceMVN normalization classes to handle feature scaling, enabling users to configure and run Conformer-based ASR pipelines with standard mean-variance normalization options.

funasr/models/conformer · high confidence

Add Conformer-based ASR example for AIShell-1

Introduces a new example recipe for training and inferring a Conformer model on the AIShell-1 dataset. This includes a run script for the full pipeline (data preparation, feature extraction, training, and inference), a configuration file for a 46M parameter model, and a README documenting the training setup and resulting Character Error Rates (4.97% on dev, 5.37% on test). Demo scripts are provided as symlinks to the Paraformer equivalents.

examples/aishell/transformer · high confidence

Add Data2Vec self-supervised speech model implementation

Introduces the Data2Vec model architecture to FunASR, enabling self-supervised speech representation learning. This change adds the core model components, including the \Data2VecPretrainModel\ and \Data2VecEncoder\, along with supporting utilities for feature extraction (\ConvFeatureExtractionModel\), multi-head attention, exponential moving average (EMA) tracking, gradient scaling, and mask index computation. The implementation integrates with FunASR's frontend and encoder interfaces to allow pre-training on unlabeled audio data.

funasr/models/data2vec · high confidence

Add Dockerfile and entrypoint for online CPU inference runtime

Users can now build and run a Docker container for online CPU-based speech recognition. This change introduces a new Dockerfile (runtime/dockerfile/Dockerfile.online.cpu) that sets up the environment using a specific FunASR base image, installs dependencies, and builds the websocket server binary for both amd64 and arm64 architectures. It also includes a corresponding .dockerignore file to optimize build context and an entrypoint script (online-cpu-entrypoint.sh) that configures and launches the FunASR websocket server on port 10095 with default model paths and thread settings.

runtime/dockerfile · high confidence

Add E-Paraformer speech recognition model

Introduces the E-Paraformer model, an enhanced Paraformer variant supporting both offline and streaming modes through dynamic masking in the encoder. This change adds the core model implementation (\model.py\), a specialized decoder with SANM layers (\decoder.py\), a PIF predictor for token alignment (\pif\_predictor.py\), beam search logic (\search.py\), and ONNX export utilities (\export\_meta.py\). Users can now utilize this model for Mandarin speech recognition tasks, including streaming inference and export to ONNX format.

_funasr/models/e\paraformer · high confidence

Add EEND-OLA end-to-end diarization model and supporting utilities

Introduces the EEND-OLA (End-to-End Neural Diarization with Overlap-aware Latent Attractors) model to the FunASR framework. This change adds the core \DiarEENDOLAModel\ class, which uses a transformer encoder and an encoder-decoder attractor mechanism to perform speaker diarization, including handling overlapping speech. It also includes the \EENDOLATransformerEncoder\, \EncoderDecoderAttractor\, and associated data loading utilities (\EENDOLADataset\, \EENDOLADataLoader\) for Kaldi-style data directories. Supporting modules for feature extraction (STFT, log-mel transforms), loss functions (standard, PIT, power loss), and evaluation metrics (DER, SAD, speaker errors) are also added to enable training and inference for this specific diarization capability.

funasr/models/eend · high confidence

Add ERes2NetV2 speaker verification model

The ERes2NetV2 speaker verification model is now available in FunASR, providing 192-dimensional speaker embeddings for speaker verification and diarization. This new model, registered under the class name 'ERes2NetV2' and the alias 'iic/speech\_eres2netv2\_sv\_zh-cn\_16k-common', is designed to outperform CAM++ for short-duration audio (less than 3 seconds). It integrates with the standard FunASR inference pipeline, automatically loading pretrained weights from 'pretrained\_eres2netv2.ckpt' when available.

funasr/models/eres2net · high confidence

Add Emotion2vec speech emotion representation model

Introduces the Emotion2vec model for self-supervised speech emotion representation. This change adds the core model implementation, including audio encoding, transformer blocks, and ALiBi positional biases, along with utilities for ONNX export and configuration templates, registering the model under the 'Emotion2vec' name in the FunASR framework.

funasr/models/emotion2vec · high confidence

Add FunASR ONNX runtime utility modules

The \runtime/python/onnxruntime/funasr\_onnx/utils\ package now includes core utility modules for the FunASR ONNX runtime: \e2e\_vad.py\ provides an end-to-end Voice Activity Detection state machine and window detector; \frontend.py\ implements audio frontends (\WavFrontend\, \WavFrontendOnline\) using \kaldi\_native\_fbank\ for feature extraction with LFR and CMVN; \postprocess\_utils.py\ handles text post-processing including abbreviation handling and sentence formatting; \sentencepiece\_tokenizer.py\ wraps the SentencePiece library for tokenization; \timestamp\_utils.py\ calculates word-level timestamps from model outputs; and \utils.py\ contains foundational classes like \OrtInferSession\ for ONNX Runtime session management, \TokenIDConverter\, and \CharTokenizer\.

_runtime/python/onnxruntime/funasr\onnx/utils · high confidence

Add FunASR Web Demo with Real-Time and Offline Speech Recognition

Introduces a new client-side demo application for FunASR in the runtime/html5/static directory. The interface allows users to connect to a WebSocket ASR server and perform speech recognition via microphone or uploaded audio files. It supports switching between online, offline, and 2-pass recognition modes, enables Inverse Text Normalization (ITN), and allows configuration of hotwords. The implementation includes the core recording logic (recorder-core.js), audio format encoders for PCM and WAV (pcm.js, wav.js), and WebSocket communication handling (wsconnecter.js).

runtime/html5/static · high confidence

Add FunASR subtitle generator example for SRT/VTT creation

Introduces a new example script (\generate\_subtitle.py\) and documentation that allows users to generate SRT or VTT subtitle files from audio or video inputs using the FunASR library. The tool supports configurable output formats, speaker labeling, language hints, and different segmentation modes (readable grouping vs. raw sentence boundaries), with options to adjust VAD segment limits for memory-constrained environments.

examples/subtitle · high confidence

Add GLM-ASR model support with vLLM inference engine

Users can now run GLM-ASR and GLM-ASR-Nano models for speech recognition. The update introduces a new vLLM-based inference engine (GLMASRVLLMEngine) that offloads the language model to vLLM for high-throughput decoding while keeping the audio tower in PyTorch. It also adds a standard FunASR-compatible model class (GLMASR) for general inference. Key behavioral improvements include deduplication of result keys when input files share the same basename, automatic fallback of repetition\_penalty to 1.0 to prevent crashes in vLLM prompt-embeds mode, and a warning when using fp16 precision due to potential transcription degradation.

_funasr/models/glm\asr · high confidence

Add German text normalization support

Introduces a new German (de) text normalization module, including data files for currencies, measurements, dates, and numbers, along with finite-state transducers to classify and normalize German text entities.

_fun\_text\_processing/text\normalization · high confidence

Add HTML5 demo server for Web Speech Recognition access

A new Flask-based server (h5Server.py) is introduced in the runtime/html5 directory to host the HTML5 client interface for the FunASR speech recognition service. This server serves static web assets and supports WebSocket connections, allowing users to access the demo via a browser on desktop or mobile devices. It includes command-line arguments to configure the host, port (default 1337), and SSL certificate paths, enabling secure HTTPS access. Documentation (readme.md and readme\_zh.md) has been added to guide users on starting the service and connecting clients.

runtime/html5 · high confidence

Add Java WebSocket client for FunASR with offline connection cleanup

A new Java WebSocket client (FunasrWsClient) is introduced in the runtime/java directory, allowing users to connect to the FunASR server for speech recognition in offline, online, and 2pass modes. The client includes a Makefile for building and running the client, along with a test suite (FunasrWsClientTest) that verifies correct connection cleanup behavior, specifically ensuring the connection closes after receiving a final response in offline mode.

runtime/java · high confidence

Add LCB-NET multimodal ASR example and evaluation utilities

The examples/industrial\_data\_pretraining/lcbnet directory now includes a complete demonstration and evaluation suite for the LCB-NET model. This adds a Python demo script for inference using the funasr AutoModel, a shell script for batched inference on GPU, and new utility scripts (compute\_wer\_details.py, run\_bwer\_recall.sh) to calculate Word Error Rate (WER) and BWER/UWER metrics. A symlink to shared utils is also provided.

_examples/industrial\_data\pretraining/lcbnet · high confidence

Add MFCCA and SOND model implementations

Introduces two new model architectures to the transformer module: the MFCCA (Multi-Frame Cross-Channel Attention) model for multi-speaker ASR in multi-party meeting scenarios, and the SOND (Speaker Overlap-aware Neural Diarization) model for speaker diarization. These additions include the full model definitions, encoder layers, attention mechanisms, and supporting components such as label aggregation and decoders specific to these architectures.

funasr/models/transformer · high confidence

Add MOSS-Transcribe-Diarize model adapter

Users can now use the MOSS-Transcribe-Diarize model for audio transcription with speaker diarization. This new adapter integrates the OpenMOSS model, supporting both standard text output and structured diarized responses from vLLM and SGLang backends, enabling precise timestamp and speaker identification in transcriptions.

_funasr/models/moss\_transcribe\diarize · high confidence

Add MonotonicAligner model for forced alignment and timestamp prediction

A new MonotonicAligner model is introduced to provide character-level forced alignment and timestamp prediction for audio inputs. This component uses a SANMEncoder and CifPredictorV3 to compute alignments, allowing users to obtain precise start and end times for each token in the recognized text. The implementation includes the model class, a YAML configuration template for training and inference, and registration within the FunASR framework.

_funasr/models/monotonic\aligner · high confidence

Add Python FunASR API with WebSocket support

Introduces a new Python client library for the FunASR speech recognition engine, enabling both offline and streaming recognition via WebSocket connections. Users can now instantiate a recognizer to process audio files or buffers directly, or create streaming sessions to feed audio chunks in real-time and receive intermediate results via callbacks.

_runtime/funasr\api · high confidence

Add Python bindings for Paraformer and SenseVoice speech recognition models

The \runtime/python/libtorch/funasr\_torch\ module now provides Python classes (\Paraformer\ and \SenseVoiceSmall\) for running speech recognition models using PyTorch. These bindings support loading TorchScript models, handling audio preprocessing, and performing inference with features like batch processing, language identification (for SenseVoice), and timestamp plotting.

_runtime/python/libtorch/funasr\torch · high confidence

Add Qwen-Audio and Qwen-Audio-Chat model wrappers

New model wrappers for Qwen-Audio and Qwen-Audio-Chat have been added to the FunASR framework, enabling universal audio understanding and interactive multi-turn audio chat. The \QwenAudioWarp\ class supports transcription tasks via a unified audio-language model interface, while \QwenAudioChatWarp\ facilitates conversational interactions with audio inputs. These models are registered in the FunASR model registry under multiple aliases (e.g., \Qwen/Qwen-Audio\, \QwenAudioChat\) and integrate with the existing frontend and tokenizer infrastructure, including support for loading models from Hugging Face Transformers with remote code execution.

_funasr/models/qwen\audio · high confidence

Add Qwen3-ASR model support with dependency validation and language code mapping

Users can now use the Qwen3-ASR model within FunASR via the \Qwen3ASR\ class, which wraps the \qwen-asr\ package and supports downloading models from both HuggingFace and ModelScope. The implementation includes a dependency guard that validates the presence and version compatibility of the \qwen-asr\ and \transformers\ packages, providing clear installation instructions if mismatches are detected. Additionally, the model automatically maps short or ISO language codes (e.g., 'zh', 'en') to the canonical names required by the underlying \qwen-asr\ library, allowing users to specify languages using standard codes without manual translation.

_funasr/models/qwen3\asr · high confidence

Add Qwen3-ASR offline and streaming WebSocket examples

This location introduces runnable examples for the Qwen3-ASR model, including \demo.py\ and \demo\_ms.py\ for basic offline transcription (supporting Chinese, English, and auto-detection via Hugging Face or ModelScope), \transcribe\_vllm\_offline.py\ for long-form audio processing using native vLLM acceleration with chunking and optional MOSS diarization, and \serve\_qwen3\_asr\_ws.py\ for real-time streaming transcription over a WebSocket interface. These examples provide users with concrete implementations for both batch and real-time speech recognition workflows using the Qwen3-ASR-1.7B model.

_examples/industrial\_data\_pretraining/qwen3\asr · high confidence

Add RWKV-based encoder with CUDA-accelerated linear attention

Introduces a new RWKV encoder (\RWKVEncoder\) for the FunASR speech recognition pipeline, featuring a linear attention mechanism backed by custom CUDA kernels (\wkv\_cuda.cu\) and PyTorch C++ extensions (\wkv\_op.cpp\) for both encoder and decoder paths. The implementation includes Python modules for RWKV blocks, self-attention (\EncoderSelfAttention\, \DecoderSelfAttention\), feed-forward layers, and conv-based subsampling, registered under \encoder\_classes\ as \RWKVEncoder\. A configuration template (\template.yaml\) is provided demonstrating usage with a Transducer model, specifying encoder parameters such as 18 blocks, 512 output size, and time reduction.

_funasr/models/rwkv\bat · high confidence

Add SOND encoder and pooling components

Added new encoder implementations for the SOND (Speaker and Noise Diarization) module, including ConvEncoder, ECAPA-TDNN, ResNet34, FSMN, and Self-Attention encoders, along with corresponding scoring mechanisms (DotScorer, CosScorer) and pooling layers (TAP, TSDP, TSTP, ASTP, StatisticPooling) to support speaker embedding extraction and diarization tasks.

funasr/models/sond/encoder · high confidence

Add SenseVoice Whisper compatibility library

The \funasr/models/sense\_voice/whisper\_lib\ package has been added, providing a self-contained implementation of the Whisper ASR engine (including model loading, audio processing, decoding, tokenization, and text normalization) to support SenseVoice features. This library enables language detection, transcription, and translation, and includes specific adaptations for SenseVoice such as handling special tokens for audio events (Speech, BGM, Laughter, Applause) and emotions, as well as fixes to prevent batch crashes when processing special-token strings.

_funasr/models/sense\_voice/whisper\lib · high confidence

Add Silero VAD adapter for FunASR

Users can now integrate Silero Voice Activity Detection into the FunASR AutoModel pipeline by specifying \vad\_model='silero-vad'\. This new adapter handles audio loading, speech timestamp detection, and segment formatting compatible with FunASR, supporting both standard PyTorch and ONNX inference modes with configurable sampling rates and thresholds.

_funasr/models/silero\vad · high confidence

Add Transducer (RNN-T) speech recognition model

Introduces a new Transducer model for streaming automatic speech recognition, including the main model class, RNN-based decoders (LSTM/GRU), a joint network, and beam search inference logic. This enables users to perform low-latency, chunk-by-chunk transcription using the encoder-predictor-joint architecture.

funasr/models/transducer · high confidence

Add Triton GPU deployment support for SenseVoice and Paraformer ASR models

This change introduces new Dockerfiles, model repositories, and client scripts to deploy SenseVoice and Paraformer large ASR models via NVIDIA Triton Inference Server. It includes a Dockerfile for SenseVoice (based on tritonserver 24.05) that installs dependencies like kaldifeat and downloads the model, and a generic server Dockerfile (based on tritonserver 23.01) for Paraformer. The update also provides comprehensive README documentation for building TensorRT engines for SenseVoice, launching the servers, and running benchmarks, along with client utilities for offline and streaming inference against these new model endpoints.

_runtime/triton\gpu · high confidence

Add WenetSpeech Conformer ASR example

Introduces a new Conformer-based automatic speech recognition example for the WenetSpeech dataset. This includes a training configuration file (conformer\_12e\_6d\_2048\_512.yaml) defining the model architecture, a main run script for data preparation, feature generation, training, and inference, along with helper scripts for data preparation and downloading. The example also provides a README with training details and reported Character Error Rates (CER) of 4.42% on dev and 4.87% on test sets.

examples/wenetspeech · high confidence

Add Whisper model integration with fine-tuning support

The FunASR library now includes a new Whisper model wrapper (\WhisperWarp\) that integrates OpenAI's Whisper models for multilingual speech recognition and translation. This addition supports model variants from \whisper-tiny\ through \whisper-large-v3-turbo\ and registers them in the model registry. A key capability is the implementation of a \forward()\ method that computes cross-entropy loss, enabling users to fine-tune the model on custom datasets. The integration handles audio preprocessing via \WhisperFrontend\ and supports both inference and training workflows.

funasr/models/whisper · high confidence

Add Whisper-LID language identification model

Introduces a new language identification model (OpenAIWhisperModel) that detects spoken language from audio input across 99 languages. The implementation wraps OpenAI's Whisper encoder and decoder components (OpenAIWhisperEncoderWarp and OpenAIWhisperDecoderWarp) and integrates an ERes2Net-based language ID predictor (LidPredictor) with various pooling layers (TAP, TSDP, TSTP, ASTP) for feature extraction.

_funasr/models/whisper\lid · high confidence

Add custom optimizer implementations and registry

The \funasr/optimizers\ module now includes a registry of available optimizers and two custom implementations: \FairseqAdam\, which adapts the Adam algorithm with fixed weight decay regularization and memory-efficient FP16 support, and \SGD\, a thin wrapper around PyTorch's SGD that provides a default learning rate to ensure compatibility with the framework's task invocation requirements.

funasr/optimizers · high confidence

Add demo scripts for multilingual language identification using Whisper

New demo scripts have been added to the \examples/common\_voice/whisper\_lid\ directory to demonstrate language identification capabilities. The \demo\_funasr.py\ script showcases integration with the FunASR library, while \demo\_modelscope.py\ provides an example using the ModelScope pipeline. Both scripts illustrate how to process multilingual audio samples (Chinese, English, Japanese, and Korean) to detect the language.

_examples/common\voice · high confidence

Add e-Paraformer recipe for AISHELL-1

Introduces a new end-to-end training and inference recipe for the e-Paraformer model on the AISHELL-1 dataset. This includes a configuration file defining the EParaformer architecture (ConformerEncoder, ParaformerSANDecoder, PifPredictor), data preparation and download scripts, and utility tools for text tokenization, normalization, and WER calculation.

_examples/aishell/e\paraformer · high confidence

Add iOS Paraformer online streaming demo

Adds an iOS demo application for Paraformer online streaming speech recognition. The project includes an Xcode workspace with C++ ONNX runtime integration, audio capture via AVFoundation, and a simple UI for recording and displaying real-time transcription results.

(repo-wide) · high confidence

Add migration benchmark example for comparing FunASR with other ASR providers

Users can now run a benchmark script in the examples/migration directory to evaluate FunASR performance against baselines like Whisper or cloud ASR APIs. The tool processes a set of audio files, outputs machine-readable results (results.jsonl) and a Markdown summary (summary.md) containing metrics such as throughput, error rates, and per-file transcription previews, helping users assess migration readiness.

examples/migration · high confidence

Add offline and streaming SANM KWS examples for Xiaoyun commands

New example directories for the SANM keyword spotting (KWS) model are added for both offline and streaming inference modes, targeting the 'Xiaoyun' command set. The offline examples include a configuration file, a Python demo using the FunASR AutoModel API, and shell scripts for fine-tuning, inference, and ONNX export. The streaming examples provide equivalent resources, with configurations and inference scripts supporting chunk-based processing (e.g., chunk sizes \[4, 8, 4\] and \[5, 10, 5\]) to enable real-time keyword detection.

_examples/industrial\_data\_pretraining/sanm\_kws, examples/industrial\_data\_pretraining/sanm\_kws\streaming · high confidence

Add self-hosted FunASR realtime transcription plugin

The integrations/openclaw directory now includes a new OpenClaw plugin that connects OpenClaw Talk and Voice Call to a self-hosted FunASR WebSocket server, keeping audio on infrastructure you control. Users can configure the provider via the voice-call streaming provider map (or Talk transcription sessions) with options for the WebSocket endpoint, authentication, recognition mode (online, offline, or 2pass), hotwords, and inverse text normalization. The plugin handles audio conversion from OpenClaw's 8 kHz G.711 mu-law to 16 kHz PCM, sends 60 ms frames over the binary WebSocket subprotocol, and returns partial and final transcripts with bounded queues and transcript retention. It requires OpenClaw \>=2026.7.2 and is enabled by default.

integrations, integrations/openclaw · high confidence

Added AISHELL data preparation and download scripts for Conformer and Transformer examples

New shell scripts have been added to the \examples/aishell/conformer/local\ and \examples/aishell/transformer/local\ directories to support the AISHELL dataset. The \download\_and\_untar.sh\ script handles downloading and extracting the \data\_aishell\ and \resource\_aishell\ corpus parts from OpenSLR, including verification of archive sizes and optional cleanup. The \aishell\_data\_prep.sh\ script prepares the dataset by splitting audio files into train, dev, and test sets, generating utterance lists, and creating standard Kaldi-style data directories (wav.scp, text) for both model architectures.

_examples/aishell/branchformer/local, examples/aishell/e\branchformer/local, examples/aishell/conformer/local, examples/aishell/transformer/local · high confidence

Added AISHELL-1 example scripts and HF model stub

Added shell scripts for downloading and preparing the AISHELL-1 dataset for the Paraformer example, and created an empty initialization file for the Hugging Face model integration.

_examples/aishell/paraformer/local, funasr/models/model\hf · high confidence

Added ASR evaluation and performance benchmarking utilities

New scripts have been added to the runtime utilities to support speech recognition model evaluation and performance testing. This includes \test\_cer.sh\ and \test\_cer.py\ for calculating Word Error Rate (WER) and Character Error Rate (CER) against reference transcripts, \test\_rtf.sh\ and \test\_rtf.py\ for measuring Real-Time Factor (RTF) inference speed on CPU, and \test\_rtf\_gpu.py\ for GPU-based RTF benchmarking with batch size support. Supporting tools include \compute\_wer.py\ for detailed error analysis, \proce\_text.py\ for text normalization, and \split\_scp.pl\ for parallelizing data processing across multiple jobs.

runtime/python/utils · high confidence

Added CTC inference demo and shell script

Added a Python demo script and a shell script to the CTC example directory, enabling users to run inference on local audio files using a pre-trained CTC model via the FunASR library.

_examples/industrial\_data\pretraining/ctc · high confidence

Added Chinese industrial pretraining example with FunASR 1.0

Added a new example directory for the \paraformer-zh-spk\ model, providing Python and shell scripts to demonstrate inference using the FunASR 1.0 API. The example showcases a pipeline combining speech recognition, voice activity detection (VAD), punctuation restoration, and speaker diarization, including support for hotword injection and streaming configurations.

_examples/industrial\_data\pretraining/paraformer-zh-spk · high confidence

Added Conformer ASR demo scripts

New demo files (demo.py and demo.sh) have been added to the Conformer example directory, providing users with ready-to-run examples for speech recognition using the FunASR library. The Python script demonstrates loading a pre-trained Conformer model via AutoModel and generating transcription from an audio URL, while the shell script shows how to run inference using the FunASR inference tool with specific model and input configurations.

_examples/industrial\_data\pretraining/conformer · high confidence

Added FST-based decoding graph compilation tools

New scripts have been added to the runtime/tools/fst directory to support building Finite State Transducer (FST) decoding graphs for speech recognition. This includes utilities for lexicon processing (add\_lex\_disambig.pl, make\_lexicon\_fst.pl), token FST generation (ctc\_token\_fst.py), and end-to-end graph assembly (compile\_dict\_token.sh, make\_decode\_graph.sh, train\_lms.sh). These tools enable the creation of the final TLG.fst decoding graph from language models and lexicons, along with helper scripts for option parsing and resource collection.

runtime/tools · high confidence

Added FunASR offline and online CPU deployment scripts and Docker installer

New shell scripts have been added to the deploy tools to streamline the setup of FunASR runtime environments. This includes \funasr-runtime-deploy-offline-cpu-en.sh\ and \funasr-runtime-deploy-offline-cpu-zh.sh\ for offline English and Chinese speech recognition, \funasr-runtime-deploy-online-cpu-zh.sh\ for online Chinese recognition, and \install\_docker.sh\ to automatically install Docker on supported Linux distributions (Ubuntu, CentOS, Debian, AliOS, AliLinux). These scripts handle downloading models, configuring Docker containers, and managing progress, enabling users to quickly deploy FunASR services on CPU-only systems.

_runtime/deploy\tools · high confidence

Added FunASR-based speech timestamp prediction demo

The monotonic aligner example now includes a demo for speech timestamp prediction using the FunASR library. Users can run the provided Python script (demo.py) or shell script (demo.sh) to perform inference with the 'iic/speech\_timestamp\_prediction-v1-16k-offline' model, supporting both single and batch audio inputs to generate text with timestamps. A Chinese README (README\_zh.md) has also been added to document the usage and API parameters for the FunASR AutoModel.

_examples/industrial\_data\_pretraining/monotonic\aligner · high confidence

Added Golang WebSocket client for offline audio recognition

A new Golang-based WebSocket client implementation has been added to the runtime, providing a native alternative for connecting to the FunASR websocket backend. This client supports offline mode processing of WAV files, handling audio chunking, base64 encoding, and hotword configuration to interact with the C++ websocket server.

runtime/golang · high confidence

Added Kaldi base library and build configuration for Windows support

This change introduces the Kaldi base library components (including error handling, I/O functions, math utilities, and versioning scripts) and a CMake build configuration to the ONNX Runtime third-party dependencies. The CMakeLists.txt explicitly configures the build for Windows MSVC, linking against system libraries like \dl\ on non-Windows platforms and defining \GLOG\_NO\_ABBREVIATED\_SEVERITIES\ on Windows to resolve compilation errors. This enables the Kaldi speech recognition toolkit to be built and integrated into the ONNX Runtime on Windows.

_runtime/onnxruntime/third\_party/kaldi, runtime/onnxruntime/third\party/yaml-cpp · high confidence

Added Paraformer ASR example utility scripts

The \examples/aishell/paraformer/utils\ directory now includes a suite of helper scripts to support the Paraformer speech recognition example. This includes \compute\_wer.py\ for calculating Word Error Rate, \extract\_embeds.py\ for generating feature embeddings using Hugging Face Transformers, and \postprocess\_text\_zh.py\ for cleaning Chinese transcripts. Data management is supported by \fix\_data.sh\ and \fix\_data\_feat.sh\ to synchronize and clean data files, while \text\_tokenize.sh\ and \text\_tokenize.py\ handle Chinese word segmentation. Additional utilities include \text2token.py\ for general text tokenization, \textnorm\_zh.py\ for Chinese text normalization, and shell/Perl scripts (\filter\_scp.pl\, \shuffle\_list.pl\, \split\_scp.pl\, \parse\_options.sh\) for standard Kaldi-style data processing and argument parsing.

examples/aishell/paraformer/utils · high confidence

Added Qwen-Audio demo scripts for transcription and chat

Added four new Python demo scripts (\demo.py\, \demo\_chat.py\, \demo\_from\_local.py\, \demo\_chat\_from\_local.py\) in the \examples/industrial\_data\_pretraining/qwen\_audio\ directory. These examples demonstrate how to use the \funasr\ library's \AutoModel\ to perform audio transcription and multi-turn chat with the Qwen-Audio and Qwen-Audio-Chat models, supporting both remote audio URLs and local model paths.

_examples/industrial\_data\_pretraining/qwen\audio · high confidence

Added RNN and Transformer language model implementations

New language model components have been added to the \funasr/models/language\_model\ package, including \seq\_rnn\_lm.py\ for sequential RNN-based language modeling and \transformer\_lm.py\ for transformer-based language modeling. The update also introduces a full suite of RNN encoder, decoder, and attention modules under \funasr/models/language\_model/rnn/\, providing configurable architectures for LSTM, GRU, and various attention mechanisms to support speech recognition and translation tasks.

_funasr/models/language\model · high confidence

Added SCAMA streaming ASR demo scripts

New demo files (demo.py and demo.sh) have been added to the SCAMA example directory, providing users with ready-to-run examples for streaming speech recognition using the FunASR AutoModel interface and command-line inference tools.

_examples/industrial\_data\pretraining/scama · high confidence

Added Transducer demo script for speech recognition

A new demo script (demo.py) has been added to the industrial data pretraining transducer examples, demonstrating how to use the FunASR AutoModel API to perform speech recognition with the speech\_bat\_asr-zh-cn-16k-aishell1-vocab4234-pytorch model on a sample audio file.

_examples/industrial\_data\pretraining/transducer · high confidence

Added UniASR demo scripts for industrial data pretraining

New demo files (demo.py and demo.sh) have been added to the examples/industrial\_data\_pretraining/uniasr directory, providing users with ready-to-run examples for invoking the UniASR model via Python and shell scripts.

_examples/industrial\_data\pretraining/uniasr · high confidence

Added cppjieba Chinese text segmentation library

The runtime now includes the cppjieba library for Chinese text processing, providing capabilities for word segmentation, part-of-speech tagging, and keyword extraction. This addition introduces header files for core components including the main Jieba interface, dictionary tries, HMM models, and various segmentation strategies (Mix, Full, Query, and HMM segments), enabling the ONNX Runtime to handle Chinese language inputs effectively.

_runtime/onnxruntime/third\party/jieba · high confidence

Added demo and ONNX export scripts for CT-Transformer streaming punctuation model

New example files have been added to the ct\_transformer\_streaming directory to facilitate usage of the punctuation model. demo.py and demo.sh provide a Python script and shell command for running real-time streaming inference using the AutoModel interface, while export.py and export.sh offer methods to export the model to ONNX format (with optional quantization) from either the model hub or a local path.

_examples/industrial\_data\_pretraining/ct\_transformer\streaming · high confidence

Added demo and finetuning scripts for Seaco Paraformer

The examples/industrial\_data\_pretraining/seaco\_paraformer directory now includes a Python demo script (demo.py) and shell scripts (demo.sh, finetune.sh) to facilitate inference and model adaptation. The demo script provides a high-level interface using the FunASR AutoModel API for transcribing audio files, supporting both URL inputs and local tensor/numpy arrays, with optional hotword injection. The shell scripts offer a complete workflow for finetuning the 'iic/speech\_seaco\_paraformer\_large\_asr\_nat-zh-cn-16k-common-vocab8404-pytorch' model on custom datasets, including data conversion utilities (scp2jsonl) and distributed training configuration via torchrun.

_examples/industrial\_data\_pretraining/contextual\_paraformer, examples/industrial\_data\_pretraining/seaco\paraformer · high confidence

Added demo scripts for GLM-ASR-Nano model integration

New demo scripts (demo.py and demo\_ms.py) have been added to the GLM-ASR example directory, providing users with ready-to-run examples for the GLM-ASR-Nano model. These scripts demonstrate how to initialize the model via the FunASR AutoModel interface, supporting both Hugging Face and ModelScope hubs, and perform speech recognition on sample audio files in Chinese and English.

_examples/industrial\_data\_pretraining/glm\asr · high confidence

Added demo, export, and finetuning scripts for Paraformer and CT-Transformer

New example scripts have been added to the \examples/industrial\_data\_pretraining\ directory to facilitate usage of the Paraformer and CT-Transformer models. For the Paraformer model (\bicif\_paraformer\), the entry includes \demo.py\ and \demo.sh\ for inference with VAD and punctuation, \export.py\ and \export.sh\ for exporting to TorchScript and ONNX formats, and \finetune.sh\ for distributed training with DeepSpeed support. For the CT-Transformer punctuation model (\ct\_transformer\), corresponding \demo.py\, \demo.sh\, \export.py\, and \export.sh\ files are provided to demonstrate inference and ONNX export capabilities.

_examples/industrial\_data\_pretraining/bicif\paraformer · high confidence

Added emotion2vec demo script

A new demo script (demo.py) has been added to the emotion2vec example directory, demonstrating how to load the iic/emotion2vec\_plus\_large model via FunASR's AutoModel and generate utterance-level embeddings from a test audio file.

_examples/industrial\_data\pretraining/emotion2vec · high confidence

Added kaldi-native-fbank third-party library for Fbank feature extraction

Added the kaldi-native-fbank library (version 1.13) to the runtime, providing C++ source files for computing Fbank features (including windowing, FFT, and mel-bank computation) and a CMake build system that fetches dependencies like pybind11 and googletest. This enables the ONNX Runtime to build and link against this native feature extraction component.

_runtime/onnxruntime/third\party/kaldi-native-fbank · high confidence

Added speaker verification demo script

A new demo script (demo.py) has been added to the campplus\_sv example directory, demonstrating how to use the FunASR AutoModel API to perform speaker verification on a sample Chinese audio file.

_examples/industrial\_data\_pretraining/campplus\sv · high confidence

Added streaming VAD demo and ONNX export scripts

New example files have been added to the FSMN VAD streaming location to demonstrate usage and model deployment. The \demo.py\ and \demo.sh\ scripts provide a Python and shell-based workflow for running the \iic/speech\_fsmn\_vad\_zh-cn-16k-common-pytorch\ model in streaming mode, including chunked inference with cache management. Additionally, \export.py\ and \export.sh\ scripts enable exporting this model to the ONNX format, supporting both direct model hub loading and local path inference.

_examples/industrial\_data\_pretraining/fsmn\_vad\streaming · high confidence

Added training and validation data manifests for ASR model

New data list files have been added to the \data/list\ directory to support the training and validation of the automatic speech recognition (ASR) model. The \train\ set includes \train.jsonl\, \train\_wav.scp\, \train\_text.txt\, \train\_emo.txt\, \train\_event.txt\, and \train\_text\_language.txt\, providing audio sources, transcriptions, emotion labels, event types, and language tags for four sample entries. The \val\ set includes \val.jsonl\, \val\_wav.scp\, and \val\_text.txt\ for two validation samples. These files enable the trainer to ingest specific audio-text pairs along with metadata for supervised learning.

data · high confidence

Demo scripts for ERes2NetV2 speaker verification and diarization

Added two Python demo scripts in the ERes2NetV2 speaker verification example directory to demonstrate integration with the FunASR AutoModel API. The first script, demo.py, shows how to perform automatic speech recognition with speaker diarization by combining a Paraformer ASR model, a VAD model, a punctuation model, and the ERes2NetV2 speaker model, outputting transcribed text tagged with speaker IDs. The second script, demo\_sv.py, demonstrates standalone speaker verification by loading only the ERes2NetV2 model to extract speaker embeddings from audio input.

_examples/industrial\_data\_pretraining/eres2netv2\sv · high confidence

Deploy FunASR product hub with secure URL redirection

The FunASR website (funasr.com) is now served via a new Nginx configuration that enforces HTTPS and provides a structured redirection system for ecosystem resources. A new \conversion-map.conf\ defines specific paths (e.g., \/go/github\, \/go/fun-asr\) that redirect users to external repositories and documentation. The main server configuration includes security hardening with strict Content Security Policies, HSTS, and Permissions-Policy headers, while also proxying bot traffic to a local service.

web-pages/nginx · high confidence

Fun-ASR-Nano example directory added with vLLM-based streaming ASR and speaker diarization

The \examples/industrial\_data\_pretraining/fun\_asr\_nano\ directory has been added, providing the Fun-ASR-Nano speech recognition example. This location introduces a unified vLLM inference server (\serve\_vllm.py\) and a real-time WebSocket streaming service that supports dynamic Voice Activity Detection (VAD), speaker diarization (SPK), and hotword customization. The directory includes a browser-based demo (\client\_mic.html\) and Python clients (\client\_python.py\, \client\_test.py\) to interact with the streaming ASR endpoints, along with updated documentation (\README.md\, \README\_zh.md\) and an Apache 2.0 license file.

_examples/industrial\_data\_pretraining/fun\_asr\nano · high confidence

FunASR 1.4.16 release with CLI and lazy loading

This release updates the FunASR package to version 1.4.16. It introduces a new agent-friendly CLI (funasr/cli.py) that supports command-line speech recognition with JSON output and SRT subtitle formatting. The package now uses lazy loading for top-level exports like AutoModel and AutoFrontend to improve import diagnostics and reduce startup overhead. Additionally, the internal registration system (funasr/register.py) has been updated to handle duplicate registry keys more gracefully.

funasr · high confidence

Initial Android client release for FunASR

This change introduces the complete Android client application for FunASR, enabling users to perform real-time speech recognition on Android devices. The app features a main interface with a record button and an audio visualization view, connecting to an ASR server via WebSocket. It includes configuration options for the server URI and custom hotwords, which are persisted in shared preferences. The implementation handles microphone permissions, audio recording, and SSL socket connections (with a trust-all configuration for development flexibility).

runtime/android/AndroidClient · high confidence

Initial HTTP-to-WebSocket recognition client for FunASR

Added a new Spring Boot application that exposes HTTP endpoints (/recognition/testIO and /recognition/testNIO) to accept audio file uploads and forward them to a FunASR server via WebSocket. The service saves the uploaded file, sends recognition configuration and audio data over the WebSocket connection, and supports configurable model, hotwords, and server address via application.yml.

_runtime/java/java\_http2ws\src/http · high confidence

Initial release of FunTextProcessing toolkit

This change introduces the FunTextProcessing library, a Python toolkit for fundamental text processing tasks in Automatic Speech Recognition (ASR). It provides capabilities for Inverse Text Normalization (ITN), Text Normalization (TN), and number-to-words conversion across multiple languages. The release includes the core package structure, a README with usage examples, and an installation script to handle the pynini dependency.

_fun\_text\processing · high confidence

Initial release of SenseVoice model with ONNX export support

The SenseVoice speech recognition model is now available in the FunASR library. This release includes the core model implementation (model.py) and utility functions for CTC alignment. A key addition is the export\_meta.py module, which enables exporting the SenseVoice model to ONNX format, allowing for deployment in environments that support ONNX runtimes. The model supports inputs for speech, language, and text normalization, and outputs CTC logits and encoder output lengths.

_funasr/models/sense\voice · high confidence

Initial repository scaffolding and documentation structure

The repository has been initialized with core documentation and configuration files, including localized READMEs (English, Chinese, Japanese, Korean), a detailed CONTRIBUTING guide, a model license, and a security policy. This establishes the project's public-facing structure, defining the scope of the FunASR toolkit versus model-specific repositories, outlining installation and deployment paths (including native Transformers and vLLM), and setting community contribution standards.

(repo-wide) · high confidence

Initial setup of the FunASR Vue-based web interface

This change introduces the foundational structure for the FunASR website (www.funasr.com) using a Vue.js framework. It adds the necessary configuration files for the development environment, including ESLint rules, Babel presets for Ant Design Vue, and a Vue CLI configuration that sets up a local development server on port 9018 with API proxying. The entry also includes a README outlining the site's content strategy (covering FunASR features, Paraformer models, and offline/online transcription services) and mock data handlers to support initial frontend development.

web-pages · high confidence

Initial web application shell and core UI components

This change introduces the foundational structure for the web application, including the main entry point, routing, and global state management. It adds a responsive app layout with a custom scrollbar component and integrates a video player component (Jessibuca) for streaming. The update also establishes a design system using SCSS with a px-to-rem conversion utility, applies browser style normalization, and configures global Axios interceptors for request handling and authentication.

web-pages/src · high confidence

Introduce BiCifParaformer model with bidirectional CIF timestamp prediction

Adds the BiCifParaformer model, an extension of Paraformer that integrates a bidirectional Contextual Information Flow (CIF) predictor to provide accurate character-level timestamp predictions alongside automatic speech recognition. This change introduces the core model implementation in \model.py\, the \CifPredictorV3\ predictor logic in \cif\_predictor.py\, and ONNX export support via \export\_meta.py\, allowing users to export the model for deployment while retaining timestamp output capabilities.

_funasr/models/bicif\paraformer · high confidence

Introduce CT-Transformer Streaming model for online punctuation restoration

Adds the CTTransformerStreaming model and its supporting components (encoder, attention, export utilities) to FunASR. This new capability enables online, incremental punctuation restoration for streaming ASR pipelines by processing text in sliding windows while maintaining context cache. The model supports VAD-aware punctuation decisions and includes ONNX export functionality for deployment.

_funasr/models/ct\_transformer\streaming · high confidence

Introduce ContextualParaformer model with hotword biasing and ONNX export support

Adds the ContextualParaformer model, an extension of Paraformer that incorporates a context encoder to boost recognition of user-defined hotwords or keywords. This new model component includes the core architecture (model.py, decoder.py), configuration templates (template.yaml), and dedicated ONNX export utilities (export\_meta.py) to enable deployment of the biasing capability in optimized formats.

_funasr/models/contextual\paraformer · high confidence

Introduce FSMN-KWS and SANM-KWS keyword spotting models

Added two new keyword spotting models, FSMN-KWS and SANM-KWS, to the FunASR framework. The FSMN-KWS model utilizes a Feedforward Sequential Memory Network architecture, while the SANM-KWS model employs a Self-Attention Neural Memory mechanism for improved context modeling. Both models are registered in the system, support CTC-based training and inference, and include specific encoder implementations (with Kaldi/PyTorch conversion utilities for FSMN) and export metadata for ONNX deployment.

_funasr/models/fsmn\kws · high confidence

Introduce Fun-ASR-Nano with vLLM inference and pipeline support

This change introduces the Fun-ASR-Nano model module, providing a new high-throughput inference engine powered by vLLM. Users can now leverage vLLM for batched and streaming ASR, significantly improving performance. The update also includes a new pipeline that integrates VAD, ASR, and Speaker Diarization, and adds support for LoRA fine-tuning of the underlying Qwen3 LLM.

_funasr/models/fun\_asr\nano · high confidence

Introduce FunASR Server and CLI entry points

The \funasr/bin\ package now includes a new \server.py\ CLI and \\_server\_app.py\ backend that provide an OpenAI-compatible REST API for speech recognition. This server supports automatic model selection (Fun-ASR-Nano on GPU, SenseVoice on CPU), custom model paths, HuggingFace/ModelScope hubs, CORS configuration, and speaker diarization. It also introduces a \realtime\_ws.py\ WebSocket server for streaming ASR with VAD, batching, and hotword support, alongside standard CLI tools for inference, training, model export, and text tokenization.

funasr/bin · high confidence

Introduce HTTP runtime for FunASR with multi-process support

Adds a new HTTP-based runtime for FunASR, providing a FastAPI server (server.py) and a sample client (client.py) to expose speech recognition via a REST API. The server supports configurable ASR, VAD, and punctuation models, hotword injection, and SSL. It includes an Nginx configuration (asr\_nginx.conf) and a startup script (start\_server.sh) to facilitate multi-process deployment for handling concurrent requests.

runtime/python/http · high confidence

Introduce LCBNet, a lightweight convolutional block network for ASR

Adds the LCBNet model to the FunASR framework, providing a new architecture optimized for low-resource deployment scenarios using depthwise separable convolutions. This new model integrates a text encoder, fusion encoder, and bias predictor, and is registered in the model registry for immediate use alongside existing models.

funasr/models/lcbnet · high confidence

Introduce LLM-based non-autoregressive ASR model

Added a new \LLMASRNAR\ model class and its supporting \Linear\ adaptor component within the \funasr/models/llm\_asr\_nar\ module. This feature enables speech recognition by integrating a pre-trained Large Language Model (defaulting to Vicuna) with an audio encoder, allowing users to leverage LLM capabilities for ASR tasks via the standard FunASR registration system.

_funasr/models/llm\_asr\nar · high confidence

Introduce MossFormer speech separation model

Added the MossFormer model to FunASR, enabling end-to-end speech separation. This new capability allows users to separate mixed audio input into distinct speaker streams using a architecture composed of a MossFormerEncoder, a MossFormer\_MaskNet for computing separation masks, and a MossFormerDecoder.

funasr/models/mossformer · high confidence

Introduce Paraformer v2 Community model variant

Adds a new Paraformer v2 Community speech recognition model to the FunASR library, including its core implementation (model, decoder, encoder layers), configuration template, and registration. This variant provides an alternative architecture for users to experiment with or deploy, distinct from the standard Paraformer models.

_funasr/models/paraformer\_v2\community · high confidence

Introduce ParaformerStreaming model with ONNX export support

Adds a new streaming (online) inference mode for the Paraformer model, allowing real-time, chunk-by-chunk audio transcription with state caching. This release includes the core model implementation in \model.py\, a configuration template (\template.yaml\) for streaming-specific encoder/decoder settings, and an \export\_meta.py\ module that enables exporting the model to ONNX format for deployment.

_funasr/models/paraformer\streaming · high confidence

Introduce SANM and CTC model implementations

Adds the SANM (Memory-equipped Self-Attention for Speech Recognition) model and a standalone CTC module to the FunASR framework. The SANM model (\funasr/models/sanm/model.py\) extends the Transformer architecture with specialized encoder and decoder layers (\funasr/models/sanm/encoder.py\, \funasr/models/sanm/decoder.py\) that utilize self-attention with memory buffers and FSMN blocks for efficient sequence modeling. The new CTC module (\funasr/models/ctc/ctc.py\) provides a hybrid CTC-attention model (\funasr/models/ctc/model.py\) supporting both builtin PyTorch and warpctc backends, including logic to handle NaN gradients. These components are registered in the system tables and include a configuration template (\funasr/models/sanm/template.yaml\) for training.

funasr/models/sanm · high confidence

Introduce SCAMA streaming ASR model

Adds the SCAMA (Streaming Chunk-Aware Multi-head Attention) model to FunASR, enabling low-latency, chunk-based speech recognition. This new model includes a dedicated encoder (SANMEncoderChunkOpt) and decoder (FsmnDecoderSCAMAOpt) that process audio in overlapping chunks to support online inference, along with a custom beam search implementation and chunking utilities. A configuration template (template.yaml) is provided to demonstrate the recommended architecture settings for training and inference.

funasr/models/scama · high confidence

Introduce SeacoParaformer model with ONNX export support

Adds the SeacoParaformer model, a Semantic-Aware Contextual Paraformer variant that integrates hotword boosting via a bias encoder and seaco decoder. This release includes the core model implementation, a YAML configuration template for training and inference, and an export module enabling conversion to ONNX format for deployment.

_funasr/models/seaco\paraformer · high confidence

Introduce UniASR unified streaming and non-streaming model

Adds the UniASR model component, a unified architecture that supports both streaming (online) and non-streaming (offline) speech recognition within a single model. This change introduces the core model implementation (\model.py\), a beam search decoding module (\beam\_search.py\), and a configuration template (\template.yaml\) that defines the encoder, decoder, predictor, and frontend settings required to run the model.

funasr/models/uniasr · high confidence

Introduce in-browser audio decoding via Jessibuca WASM

The public web interface now includes a JavaScript decoder module that embeds the Jessibuca WebAssembly runtime. This adds client-side audio decoding capabilities to the web page, allowing the application to process audio streams directly in the browser without relying on server-side processing for this step.

web-pages/public · high confidence

Introduce multi-task FSMN keyword spotting model

Added a new \FsmnKWSMT\ model and its \FSMNMT\ encoder in the \funasr/models/fsmn\_kws\_mt\ module. This multi-task model simultaneously performs keyword spotting and filler token classification using two CTC branches, improving detection robustness through auxiliary tasks.

_funasr/models/fsmn\_kws\mt · high confidence

Introduce new training utility modules for checkpoint averaging and validation metrics

The \funasr/train\_utils\ package now includes new modules to support advanced training workflows. \average\_nbest\_models.py\ provides functionality to average the last N best checkpoints based on validation metrics, with specific handling for DeepSpeed model states. \checkpoint\_metrics.py\ introduces a \ValidationMetrics\ class to track and compute distributed validation loss and accuracy, ensuring that only checkpoints with finite, valid metrics are considered for best-model tracking and averaging. Additionally, \add\_gradient\_noise.py\ adds a utility to inject noise into gradients during training, and \load\_pretrained\_model.py\ offers improved logic for loading and mapping pretrained model weights, including support for scope mapping and exclusion lists.

_funasr/train\utils · high confidence

Introduce standalone ONNX Runtime Python package for FunASR

The \funasr-onnx\ package is now available as a separate distribution from the main \funasr\ library, allowing users to install and upgrade ONNX wrappers independently. This release (v0.4.3) provides Python bindings for running ONNX models via ONNX Runtime, including demos and documentation for Paraformer (offline and online), FSMN-VAD, CT-Transformer punctuation restoration, Contextual Paraformer, Seaco Paraformer, and SenseVoice. It also includes a FastAPI-based HTTP server example for deploying ASR inference over HTTP.

runtime/python/onnxruntime · high confidence

Introduce standalone llama.cpp runtime for FunASR models

Adds a new C++ runtime in \runtime/llama.cpp\ that runs FunASR models (Fun-ASR-Nano, SenseVoiceSmall, Paraformer) on CPU and edge devices using the llama.cpp/ggml stack, eliminating the need for Python or PyTorch at runtime. The package includes a CMake build system, CLI binaries for each model, and Python scripts to convert and download pre-quantized GGUF weights. It supports built-in FSMN-VAD for audio segmentation, SRT subtitle output, and optional GPU acceleration via Windows CUDA and Linux/Windows Vulkan backends.

runtime/llama.cpp · high confidence

Introduce streaming FSMN-VAD model with dynamic silence thresholds

A new streaming Voice Activity Detection (VAD) implementation is added under \funasr/models/fsmn\_vad\_streaming\, providing a \DynamicStreamingVAD\ wrapper that adjusts silence cut-off thresholds based on the accumulated duration of the current speech segment (e.g., waiting longer for short utterances, cutting faster for long ones). The package includes the core \model.py\ defining the FSMN encoder and VAD state machine, \encoder.py\ for the neural network layers, and \export\_meta.py\ to support ONNX model export. This allows users to process audio in chunks with adaptive segmentation logic tailored for real-time streaming scenarios.

_funasr/models/fsmn\_vad\streaming · high confidence

Introduce streaming keyword spotting model with ONNX export support

Adds the \SanmKWSStreaming\ model class and associated utilities to the \funasr/models/sanm\_kws\_streaming\ package. This new component enables streaming-based keyword spotting inference, including chunk-based encoding and state caching for continuous audio processing. It also introduces an \export\_meta.py\ module that provides specific methods to export the model's encoder to ONNX format, facilitating deployment in environments requiring static graph execution.

_funasr/models/sanm\_kws\streaming · high confidence

Introduce unified AutoModel API with vLLM support and frontend separation

The \funasr/auto\ package now provides a unified entry point for inference via the new \AutoModel\ class, which consolidates model loading, VAD, punctuation, and speaker diarization into a single interface. A dedicated \AutoFrontend\ class handles audio preprocessing (loading, feature extraction) separately from the model logic. Additionally, a new \AutoModelVLLM\ class enables high-performance inference for LLM-based ASR models (such as Fun-ASR-Nano, LLMASR, and GLM-ASR) by automatically extracting LLM weights and leveraging the vLLM engine, while explicitly excluding non-autoregressive models like Paraformer and SenseVoice.

funasr/auto · high confidence

Introduces CAMPPlus speaker embedding model and LoRA fine-tuning support

This change adds the CAMPPlus model to the FunASR library, enabling 192-dimensional speaker embedding extraction for verification and diarization tasks. The implementation includes a new clustering backend (cluster\_backend.py) that supports both Spectral Clustering and UMAP+HDBSCAN algorithms for speaker diarization, along with utility functions for audio preprocessing, feature extraction, and post-processing. Additionally, the update introduces a Low-Rank Adaptation (LoRA) module (funasr/models/lora) with specific layer implementations for Embedding and Linear layers, allowing for efficient fine-tuning of pre-trained models by freezing original weights and training only the low-rank adapters.

funasr/models/campplus · high confidence

Introduces Python WebSocket ASR service with SSL, concurrency controls, and 2pass mode

Adds a complete Python-based WebSocket speech recognition service (server and client) supporting offline, online, and 2pass unifying modes. The server (funasr\_wss\_server.py) now supports SSL/TLS connections via certfile/keyfile arguments and introduces concurrency controls (worker\_threads, concurrent\_vad/asr/punc limits) to prevent event-loop blocking during inference. The client (funasr\_wss\_client.py) supports hotword files with UTF-8 encoding, configurable chunking, and explicit result timeouts to handle connection closure gracefully. A new funASR-Service.py demonstrates a Flask-based API for speaker registration and meeting transcription with speaker identification, while funasr\_client\_api.py provides a synchronous WebSocket client library for integration.

runtime/python/websocket · high confidence

Introduction of Paraformer and SA-ASR model architectures

The Paraformer non-autoregressive speech recognition model and the SA-ASR speaker-aware model are now available. This change introduces the core model implementations, including the CIF predictor for character-level timestamping, SANM-based encoder/decoder layers, and beam search logic. It also adds ONNX export support for Paraformer, allowing users to generate optimized inference graphs, and provides a configuration template for standard training setups.

funasr/models/paraformer · high confidence

Introduction of a modular tokenizer framework with multiple backend support

The \funasr/tokenizer\ package has been restructured to provide a unified, extensible tokenizer interface. A new abstract base class (\AbsTokenizer\) and a concrete \BaseTokenizer\ define the core API for text-to-token and token-to-text conversion, including vocabulary management and ID mapping. A factory function, \build\_tokenizer\, allows users to instantiate specific tokenizers by type (\bpe\, \word\, \char\, \phn\). The package now includes dedicated implementations for SentencePiece, character-level, word-level, and phoneme-level tokenization (supporting various G2P engines like pyopenjtalk and pypinyin). Additionally, it introduces text cleaning capabilities (\TextCleaner\) for languages such as Japanese (jaconv), Vietnamese, and Korean, and provides adapters for Hugging Face and OpenAI Whisper tokenizers.

funasr/tokenizer · high confidence

Launch of the FunASR product website and documentation hub

The www.funasr.com product site is now built from a new static Python/Jinja codebase in web-pages/product-site. Users benefit from a unified bilingual (zh/en) documentation system with a curated catalogue, a structured blog with categories and homepage selection, and a deployment registry that validates maturity, evidence, and hardware/runtime details. The site generates fingerprinted assets, search indexes, and sitemaps, while preserving legacy routes and anchors for historical content and API references.

web-pages/product-site · high confidence

New AudioLLMVicunaDataset for speech-to-text training

A new dataset class, AudioLLMVicunaDataset, has been added to the funasr library to support training large language models on audio data. This component enables users to load audio and text pairs, extract fbank features using a specified frontend, and process them through a tokenizer. It constructs input sequences and labels with specific padding and masking strategies (including audio masks and ignore indices) tailored for Vicuna-style instruction tuning, allowing for end-to-end speech transcription workflows.

_funasr/datasets/llm\_datasets\vicuna · high confidence

New C++ ONNX Runtime inference backend for FunASR

This change introduces a new C++ ONNX Runtime-based inference engine for FunASR, located in the \runtime/onnxruntime\ directory. It provides a standalone build system (CMake) and source code for core speech processing components, including audio handling, VAD (Voice Activity Detection), punctuation prediction (CT-Transformer), and bias language modeling. The implementation supports both CPU and GPU execution, integrates with FFmpeg for audio decoding, and includes specific build configurations for Windows (MSVC) and Linux, enabling users to compile and run the FunASR models using the ONNX Runtime library.

runtime/onnxruntime · high confidence

New C++ WebSocket runtime binaries for FunASR

The \runtime/websocket/bin\ directory now contains the build definitions and source code for the FunASR WebSocket runtime. This adds four new executables: \funasr-wss-server\ (offline mode), \funasr-wss-server-2pass\ (two-pass mode with sentence-level timestamps and hotword support), \funasr-wss-client\ (file-based testing), and \funasr-wss-client-2pass\ (real-time microphone input via PortAudio). The servers support SSL/TLS connections, configurable model paths, and hotword injection, while the clients provide the necessary interfaces for sending audio data to the ASR engine.

runtime/websocket/bin · high confidence

New C++ gRPC Paraformer server runtime

Added a new C++ gRPC-based runtime for the Paraformer ASR model in the \runtime/grpc\ directory. This introduces a standalone \paraformer-server\ executable that exposes speech recognition capabilities via a gRPC streaming interface, supporting online, offline, and two-pass decoding modes. The implementation includes the CMake build configuration, the gRPC service definition, and a server entry point that integrates with the existing FunASR C++ runtime and ONNX model inference, allowing users to deploy a high-performance ASR service endpoint.

runtime/grpc · high confidence

New C++ runtime API and audio processing headers

This change introduces the core C++ runtime interface for the FunASR ONNX runtime, making it available for integration. It adds \funasrruntime.h\, which defines the public C API for initializing and running ASR, VAD, and punctuation models, including support for offline, online, and two-pass modes, as well as WFST decoder integration. It also adds \audio.h\ for audio frame and stream management, \model.h\ for the abstract model interface, \offline-stream.h\ for the offline streaming client, and \com-define.h\ for common configuration constants. Additionally, it bundles the \tclap\ command-line argument parsing library headers to support local argument parsing without external dependencies.

runtime/onnxruntime/include · high confidence

New Colab quickstart notebook with multilingual documentation

Users can now run FunASR directly in a browser via a new Colab notebook (examples/colab/funasr\_quickstart.ipynb) that installs dependencies, automatically selects GPU or CPU, and transcribes audio using the paraformer-zh model with VAD and punctuation. The location also includes the English README and localized guides in Japanese, Korean, and Simplified Chinese, along with troubleshooting notes for common Colab issues like runtime resets and GPU availability.

examples/colab · high confidence

New Conformer-RWKV hybrid speech recognition model

Added a new speech recognition model combining a Conformer encoder with an RWKV-based Transformer decoder. The implementation includes the model definition, decoder logic supporting RWKV v4, v5, and v6 variants, and a configuration template for training. This allows users to leverage the efficiency of RWKV attention mechanisms within the existing Conformer architecture for improved inference performance.

_funasr/models/conformer\rwkv · high confidence

New FunASR MCP server for local speech transcription

This change introduces a new Model Context Protocol (MCP) server example that enables AI assistants to perform local speech-to-text transcription using the FunASR library. The \examples/mcp\_server\ directory now contains the \funasr\_mcp.py\ implementation, which exposes a \transcribe\_audio\ tool supporting Mandarin, Cantonese, English, Japanese, and Korean. Users can run this server directly via Python or as a Docker container (built from the provided \Dockerfile\), with configuration for device selection (CPU/CUDA/MPS) and model choice via environment variables. The entry includes registry metadata (\server.json\) for the official MCP Registry and Glama, along with smoke tests and unit tests to verify the stdio transport and tool contract.

_examples/mcp\server · high confidence

New FunASR dataset class with multi-context prompt support

Added a new FunASR dataset class and a MultiContextPrompt helper to enhance speech transcription inputs. The dataset class manages audio/text data loading and integrates with the FunASR registration system, while the prompt component dynamically constructs input sequences by combining historical transcriptions, one-pass results, and hotword lists (including negative sampling) in both English and Chinese.

_funasr/datasets/fun\_asr\datasets · high confidence

New German inverse text normalization pipeline

Added a complete German (de) inverse text normalization module, including finite-state transducers for classifying and verbalizing cardinals, ordinals, decimals, fractions, dates, times, measures, money, telephone numbers, and electronic addresses. The implementation mirrors the existing English pipeline structure, introducing new taggers and verbalizers in the de subdirectory to convert spoken German number-like text into structured tokens and back to written form.

_fun\_text\_processing/inverse\_text\normalization · high confidence

New LLM audio dataset classes for NAR and Qwen-Audio training

Added \AudioLLMNARDataset\ and \AudioLLMQwenAudioDataset\ classes in the \funasr/datasets/llm\_datasets\ and \funasr/datasets/llm\_datasets\_qwenaudio\ modules, along with a \TextPreprocessRemovePunctuation\ preprocessor. These new dataset implementations provide specialized data loading and formatting for training Large Language Models on audio tasks, supporting both Non-Autoregressive (NAR) and Qwen-Audio specific prompt templates and tokenization strategies.

_funasr/datasets/llm\datasets · high confidence

New LLM-based ASR inference and training demos

Added a suite of new example scripts and shell utilities in the \examples/industrial\_data\_pretraining/llm\_asr\ directory to demonstrate speech-to-text capabilities using the FunASR \AutoModel\. This includes \app.py\, a Gradio-based interactive web interface for real-time audio transcription with configurable system prompts, alongside batch inference scripts (\demo\_speech2text.py\, \demo\_speech2text\_multi.py\, \demo\_speech2text\_multi\_stream.py\) and shell wrappers (\demo\_infer.sh\, \infer\_speech2text.sh\) for evaluating models on datasets like LibriSpeech and AISHELL. Additionally, training and fine-tuning workflows are provided via \demo\_train\_or\_finetune.sh\ and \demo\_train\_or\_finetune2.sh\, which configure distributed training with DeepSpeed for LLM-ASR models.

_examples/industrial\_data\_pretraining/llm\asr · high confidence

New LLM-based ASR model with QFormer adaptor

Introduces a new \LLMASR\ model class in \funasr/models/llm\_asr\ that combines an audio encoder with a Large Language Model decoder for speech recognition. The module includes a new \adaptor.py\ file registering \Linear\, \Transformer\, and \QFormer\ adaptor classes to bridge audio features to the LLM, with the \QFormer\ adaptor specifically utilizing Hugging Face's \Blip2QFormerModel\. The main model file \model.py\ handles initialization of audio encoders (including Hugging Face and ModelScope hubs), LLMs (defaulting to Vicuna), and the adaptor, supporting configuration for freezing parameters and mixed-precision training.

_funasr/models/llm\asr · high confidence

New LibTorch and gRPC Python runtimes for ASR

This change introduces two new Python runtime options for FunASR. The \runtime/python/libtorch\ directory adds a \funasr\_torch\ package (v0.1.3) that provides a high-performance LibTorch backend for Paraformer, ContextualParaformer, SeacoParaformer, and SenseVoiceSmall models, complete with demo scripts and a setup configuration. Additionally, the \runtime/python/grpc\ directory adds a gRPC client and server implementation for streaming and full audio 2-pass decoding, including the protobuf definitions and a demo client script.

runtime/python/libtorch · high confidence

New ONNX Runtime Python bindings for FunASR models

This change introduces a new Python package at \runtime/python/onnxruntime/funasr\_onnx\ that provides ONNX Runtime-based inference wrappers for several FunASR models. Users can now import and use \Paraformer\, \ContextualParaformer\, \SeacoParaformer\, \Fsmn\_vad\, \Fsmn\_vad\_online\, \CT\_Transformer\, \CT\_Transformer\_VadRealtime\, and \SenseVoiceSmall\ directly via the \funasr\_onnx\ module. These bindings handle model loading (including automatic export from \funasr\ or download from \modelscope\), audio preprocessing, and inference, offering a standardized interface for speech recognition, voice activity detection, and punctuation prediction using ONNX models.

_runtime/python/onnxruntime/funasr\onnx · high confidence

New ONNX Runtime command-line binaries for FunASR inference

This change introduces a new build target (runtime/onnxruntime/bin) that compiles a suite of standalone executable binaries for running FunASR models via ONNX Runtime. The CMakeLists.txt defines and links executables for offline, online, 2-pass, VAD, and punctuation modes, including specific variants for real-time factor (RTF) benchmarking. These binaries provide command-line interfaces for users to perform speech recognition, voice activity detection, and punctuation restoration, supporting features like GPU inference, hotword injection, and JSONL output formatting.

runtime/onnxruntime/bin · high confidence

New OpenAI-compatible audio transcription example with deployment and client guides

The \examples/openai\_api\ directory now provides a complete, production-ready example for exposing FunASR as an OpenAI-compatible \/v1/audio/transcriptions\ endpoint. This includes a FastAPI-based server (\server.py\), Dockerfiles for both CPU and CUDA environments, and a Gradio-based browser demo for local testing. To help users integrate this service, the change adds extensive documentation covering client recipes for Python and JavaScript/TypeScript, workflow integration guides for tools like n8n and Dify, an OpenAPI specification, and security best practices for gateway deployment.

_examples/openai\api · high confidence

New OpenAI-style dataset support for SenseVoice

Added a new dataset module (\funasr/datasets/openai\_datasets\) that enables training with OpenAI-style JSONL conversation data. This includes an \OpenAIIndexDSJsonl\ class to parse and filter message histories (system, user, assistant) and an \OpenAIDataset\ class to handle tokenization, audio feature extraction (fbank), and batch construction, allowing users to fine-tune the model on instruction-following or chat-style datasets.

_funasr/datasets/openai\datasets · high confidence

New Paraformer LoRA finetuning workflow and updated documentation

This location now provides a complete, ready-to-use workflow for LoRA-based finetuning of the Paraformer model. Users can finetune using the new \lora\_finetune.sh\ script and \conf/paraformer\_lora.yaml\ configuration, run inference with \lora\_infer.py\/\lora\_infer.sh\, and evaluate results with \lora\_cer.sh\. The English and Chinese READMEs have been updated to document these new scripts, the LoRA configuration parameters (such as rank, alpha, and dropout), and the expected output directories.

_examples/industrial\_data\pretraining/paraformer · high confidence

New SenseVoiceDataset for audio-text data loading

A new SenseVoiceDataset class has been introduced to handle data loading for the SenseVoice model. This dataset manages audio preprocessing (including Fbank feature extraction and frontend handling), text tokenization, and batch collation. It supports configurable start/end-of-sequence tokens, retry logic for loading failures, and specific handling for Whisper frontends, providing the necessary data pipeline components for training or inference with SenseVoice.

_funasr/datasets/sense\_voice\datasets · high confidence

New SpecAugment and ProfileAugmentation modules for speech recognition training

This change introduces new data augmentation capabilities for speech recognition models. It adds a SpecAugment implementation (\SpecAug\, \SpecAugLFR\) that applies time warping, frequency masking, and time masking to spectrograms, with support for Low Frame Rate (LFR) processing. Additionally, it introduces a \ProfileAug\ module that performs speaker profile augmentation through split, merge, and disturb operations on speaker embeddings. These modules are registered in the \specaug\_classes\ table for use in training pipelines.

funasr/models/specaug · high confidence

New Whisper fine-tuning and inference examples

Added a new example directory for fine-tuning OpenAI Whisper models (including whisper-large-v3-turbo) using FunASR. This includes a README with data preparation and training instructions, shell scripts for fine-tuning and inference (from hub, local, or OpenAI sources), and Python demo scripts demonstrating model loading and generation with VAD integration.

_examples/industrial\_data\pretraining/whisper · high confidence

New desktop voice input example using FunASR

Added a new example in the \examples/voice\_input\ directory that provides a desktop voice input tool. This tool records audio via a configurable hotkey (default Ctrl+Shift+Space), sends it to a local FunASR server for speech recognition, and automatically pastes the transcribed text into the active application. It supports multiple platforms (macOS, Linux, Windows) and models (sensevoice, paraformer, fun-asr-nano).

_examples/voice\input · high confidence

New documentation for Cantonese, Chinese, and Fun-ASR-Nano integration

The product site now includes dedicated blog posts for Cantonese speech recognition using SenseVoice, a Mandarin speech recognition guide comparing Fun-ASR-Nano, SenseVoice, and Paraformer, a detailed usage guide for the Fun-ASR-Nano model, and a tutorial on integrating Fun-ASR-Nano with the native Hugging Face Transformers library.

web-pages/product-site/legacy · high confidence

New learning rate schedulers and unified scheduler registry

The funasr/schedulers module now provides a centralized registry for learning rate schedulers, exposing standard PyTorch schedulers (such as ReduceLROnPlateau, StepLR, and CosineAnnealingLR) alongside new custom implementations: NoamLR, WarmupLR, TriStageLR, and CustomLambdaLR. This change introduces a consistent API for configuring learning rate schedules during training, allowing users to select from a broader set of strategies including warmup, tri-stage (warmup, hold, decay), and custom lambda-based adjustments.

funasr/schedulers · high confidence

New metrics module with rapidfuzz-based CER/WER calculation

The funasr/metrics package has been introduced, providing new utilities for speech recognition evaluation. A key behavioral change is the migration of edit distance metrics to the rapidfuzz library, which is used in the new ErrorCalculator class to compute Character Error Rate (CER) and Word Error Rate (WER). This implementation includes specific guards to handle empty reference sequences, preventing errors during metric calculation. The module also adds standalone scripts for computing accuracy, Equal Error Rate (EER), and minimum Detection Cost Function (minDCF) for speaker recognition tasks.

funasr/metrics · high confidence

New modular dataloader entry point with map-style and iterable support

The \funasr/datasets\ module now includes a new \dataloader\_entry.py\ file that provides a structured way to initialize training and validation data loaders. This entry point registers \DataloaderMapStyle\ and \DataloaderIterable\ classes/functions via the \tables\ registry, allowing users to configure data loading through keyword arguments for frontend, tokenizer, and dataset configurations. The \DataloaderMapStyle\ implementation supports data splitting across epochs and handles batch sampling for both training and validation sets, while \DataloaderIterable\ provides a simpler interface for iterable-style datasets.

funasr/datasets · high confidence

New modular frontend architecture with speech enhancement and fused feature support

The \funasr/frontends\ package has been restructured into a modular system that introduces several new capabilities for audio processing. Users can now utilize \DefaultFrontend\ (also registered as \EspnetFrontend\) for a conventional pipeline including STFT, WPE noise suppression, MVDR beamforming, and Log-Mel Fbank feature extraction. A new \S3prlFrontend\ allows integration of pretrained speech representations from the S3PRL library. Additionally, \FusedFrontends\ enables combining multiple frontend feature streams (such as Default and S3PRL) via linear projection and alignment. The module also includes dedicated utilities for speech enhancement, specifically \DNN\_WPE\ for Wiener Power Prediction Estimation and \DNN\_Beamformer\ for multi-channel noise reduction, along with supporting utilities for complex tensor operations and feature transformation.

funasr/frontends · high confidence

New operational scripts for website contracts, growth metrics, and release automation

This change introduces several new Python scripts in the \scripts/\ directory to support ongoing maintenance and release workflows. \check\_funasr\_website\_static.py\ enforces content contracts on the funasr.com static site, verifying required text, links, and forbidden strings across English and Chinese pages. \collect\_growth\_metrics.py\ aggregates GitHub integration PR statuses, CI gate classifications, and PyPI download data to track ecosystem growth. \collect\_website\_traffic.py\ processes Nginx access logs to generate privacy-safe, anonymized visitor metrics for the site. \gen\_api\_docs.py\ automatically generates API reference documentation with source code previews from the FunASR codebase. \sync\_modelscope\_model\_cards.py\ patches and uploads ModelScope README files to ensure install guidance matches the latest FunASR version (v1.3.26+). \sync\_vllm\_guide.py\ maintains a generated mirror of the Chinese vLLM guide. Finally, \update\_release\_runtime\_downloads.py\ automates the attachment of prebuilt runtime assets to Python GitHub releases, validating asset integrity and publication order.

scripts · high confidence

New streaming ASR demo and export scripts for Paraformer

Added a new example directory for streaming speech recognition using the Paraformer model. This includes a Python demo script demonstrating both batch and chunked streaming inference with cache management, shell scripts for running inference and exporting the model to ONNX format, and a script for fine-tuning the model on custom data. The entry also adds a Chinese-language README documenting the usage of the streaming model, including configuration for chunk sizes and look-back windows.

_examples/industrial\_data\_pretraining/paraformer\streaming · high confidence

New text normalization and data processing tools for ASR pretraining

Added a suite of utility scripts in the \fun\_asr\_nano/tools\ directory to support industrial data pretraining. This includes \scp2jsonl.py\ for converting speech corpus files (SCP and transcripts) into JSONL format with audio duration and tokenization metadata, \whisper\_mix\_normalize.py\ for normalizing mixed Chinese-English-Japanese text using various normalizers, \cn\_tn.py\ for Chinese text normalization, \format5res.py\ for formatting recognition results, and \utils.py\ for audio loading and forced alignment helpers.

_examples/industrial\_data\_pretraining/fun\_asr\nano/tools · high confidence

New text normalization and timestamp anchoring utilities for FunASR Nano

The \funasr/models/fun\_asr\_nano/tools\ directory now includes a suite of new utility scripts to improve text processing and output precision. \cn\_tn.py\ provides comprehensive Chinese text normalization (handling numbers, currency, and punctuation), while \whisper\_mix\_normalize.py\ and \format5res.py\ add mixed-language (Chinese/English) normalization and result formatting capabilities. Additionally, \utils.py\ introduces \anchor\_punctuation\_timestamps\, which corrects punctuation timing in ASR outputs by anchoring punctuation tokens to the end of the preceding spoken segment, addressing timestamp inaccuracies during VAD merges. \scp2jsonl.py\ adds a new tool for converting SCP and transcript files into JSONL format for model training.

_funasr/models/fun\_asr\nano/tools · high confidence

New transformer scorer interfaces and implementations

The \funasr/models/transformer/scorers\ package has been introduced, providing the infrastructure for beam search scoring within the transformer model. This includes a \ScorerInterface\ hierarchy (\ScorerInterface\, \BatchScorerInterface\, \PartialScorerInterface\, \BatchPartialScorerInterface\) that defines how scorers interact with the decoder. Specific implementations include \CTCPrefixScorer\ for CTC-based prefix scoring (supporting both standard and batched/thresholded variants via \CTCPrefixScore\ and \CTCPrefixScoreTH\) and \LengthBonus\ for applying length penalties during generation. These components enable the transformer decoder to perform beam search with CTC integration and length normalization.

funasr/models/transformer/scorers · high confidence

New transformer utility modules for convolution and subsampling

Added a suite of new utility modules in the transformer utils package to support advanced encoder architectures. This includes Dynamic Convolution, Lightweight Convolution, and their 2D variants, which enable dynamic kernel generation for efficient sequence modeling. The update also introduces new subsampling layers (Conv2dSubsampling, VGG2L) and feed-forward blocks (MultiLayeredConv1d, FsmnFeedForward, Conv1dLinear) to replace standard position-wise feed-forward networks, alongside utilities for sequence padding, masking, and block repetition.

funasr/models/transformer/utils · high confidence

New utility module for FunASR

Added a new \funasr/utils\ package containing a collection of utility modules including \amp.py\ for PyTorch AMP compatibility, \export\_utils.py\ for model export, \fbank.py\ for feature extraction, \load\_utils.py\ for audio loading, \postprocess\_hotwords.py\ for text post-processing, and others.

funasr/utils · high confidence

Runtime documentation and deployment scaffolding added

The runtime area now includes comprehensive deployment guides (readme, quick start, release history) and server startup scripts (run\_server.sh, run\_server\_2pass.sh) that automatically configure thread parameters based on CPU count. An Android client project with Gradle wrapper is also added, providing a reference implementation for mobile WebSocket connections to the speech recognition service.

runtime · high confidence

SenseVoice example directory with continual fine-tuning guide and export demos

The \examples/industrial\_data\_pretraining/sense\_voice\ directory now provides a complete set of resources for the SenseVoice model. It includes a detailed English and Chinese guide on continual fine-tuning (adding new domains while retaining existing languages), along with scripts for ONNX and LibTorch export, speaker diarization, and standard inference. The fine-tuning guide specifically addresses retention constraints, replay manifests, and two-stage training strategies to prevent regression on existing languages.

_examples/industrial\_data\_pretraining/sense\voice · high confidence

Behavioural changes

FSMN KWS examples now support ModelScope hub integration and multi-task (MT) variants

The FSMN keyword spotting (KWS) examples have been updated to align with the ModelScope hub, enabling users to download pre-trained models directly via git or the AutoModel API instead of relying on local paths. The \demo.py\ and \infer.sh\ scripts now reference ModelScope model IDs (e.g., \iic/speech\_charctc\_kws\_phone-xiaoyun\), and the \finetune.sh\ scripts have been adapted to fetch and use these hub models for initialization. Additionally, a new multi-task (MT) variant (\fsmn\_kws\_mt\) has been introduced, featuring configurations and training/inference scripts that support dual token lists and multiple output heads, along with corresponding conversion utilities to handle the multi-task model structure.

_examples/industrial\_data\_pretraining/fsmn\_kws\mt · high confidence

Improved batch ASR example and real-time syntax fix

The examples directory now includes an improved batch ASR script (batch\_asr\_improved.py) that adds command-line configuration, recursive folder scanning, progress reporting, and per-file error handling for more robust offline transcription. Additionally, a syntax error in the real-time ASR calculation (Issue \#2720) has been corrected to prevent runtime failures, and the examples README files now redirect to the main tutorial documentation.

examples · high confidence

New modular audio dataset and sampling infrastructure

The audio data loading pipeline has been restructured into a new, modular system under \funasr/datasets/audio\_datasets\. This introduces a new \AudioDataset\ class that dynamically resolves preprocessors and tokenizers via a central registry, supporting both training and inference modes. It is paired with a suite of new distributed batch samplers (including \EspnetStyleBatchSampler\ and \CustomDistributedDynamicBatchSampler\) that enable length-based dynamic batching for more efficient GPU utilization. The change also adds utility scripts for converting between data formats (JSONL, SCP) and a new \AudioDatasetHotword\ class to support context-aware finetuning.

_funasr/datasets/audio\datasets · high confidence

New modular download subsystem for models and datasets

The \funasr/download\ package has been restructured into a set of dedicated modules to handle model and dataset retrieval. Users can now download models from ModelScope, Hugging Face, or OpenAI via a unified \download\_model\ entry point that resolves aliases and parses configuration files (config.yaml, configuration.json). A new \download\_dataset\_from\_hub\ module provides a helper to load datasets from ModelScope. Additionally, a \runtime\_sdk\_download\_tool\ CLI is available to export models to ONNX, TorchScript, or BladeDisc formats, and a new \file.py\ module introduces a storage abstraction (Local, HTTP) for reading and writing files.

funasr/download · high confidence

Redesigned documentation site with unified styling and new content sections

The FunASR documentation site has been redesigned with a new shared stylesheet (docs-hub.css) and a unified navigation structure across the English and Chinese hubs. The site now includes a comprehensive Training & Fine-tuning guide covering data preparation, multi-GPU training, DeepSpeed, and monitoring, alongside a new Model Selection Guide to help users choose the right model. Additionally, a dedicated page for the MOSS-Transcribe-Diarize model has been added, detailing its deployment and output contract, and the API reference has been updated with a new sidebar layout and search functionality.

gh-pages-output · high confidence

Redesigned product site with unified styling and new interactive features

The product site assets have been updated to support a comprehensive visual redesign and new interactive capabilities. A new unified CSS system (site.css) establishes a modern design language with updated typography, color palettes, and responsive layouts for the header, hero, and navigation. A dedicated experience.css provides specific styling for documentation pages, deployment guides, and the blog. Functionally, the site now includes a client-side search feature (experience.js) that allows users to filter documentation by model or runtime, and a deployment recommendation tool (site.js) that suggests optimal deployment configurations based on user-selected workload, hardware, and priority. Additionally, code blocks in documentation now feature one-click copy buttons, and the site incorporates a set of vendored Lucide icons for consistent UI elements.

web-pages/product-site/assets · high confidence

Redesigned product site with unified templates and deployment hub

The product site has been rebuilt with a new template structure, introducing a unified base layout, a dedicated deployment center (with index and detail pages for runtime contracts, commands, and limitations), a benchmarks page for reproducible measurements, a blog hub, and a redesigned homepage featuring a deployment selector and repository routing. The site now supports bilingual content (English and Chinese) with language switching, includes a 404 page, and adds structured data and SEO metadata to the base template.

web-pages/product-site/templates · high confidence

Windows build configuration for OpenFST

The OpenFST third-party library now includes a CMake build system and Windows-specific configuration. On Windows, the build is forced to static libraries, and the \HAVE\_BIN\ and \HAVE\_SCRIPT\ options are disabled to resolve MSVC compilation errors and prevent language model loading failures. The directory also now contains a Bazel build file, standard Autoconf/Automake infrastructure, and repository maintenance scripts for importing upstream releases.

_runtime/onnxruntime/third\party/openfst · high confidence

Test coverage

Added integration tests for major FunASR models and pipelines; Added test coverage for product site build, documentation, and editorial integrity; Browser test coverage for product site content and layout; Expanded test coverage for core inference, AMP compatibility, and documentation generation.

Dependencies

New dependency manifests for runtime integrations and web assets

This change introduces dependency specification files for several new or updated runtime integrations and web components. Python requirements are added for the FunASR Nano example (torch, funasr, websockets), gRPC and HTTP servers, and utility libraries. C\# projects are defined for AliFsmnVad, AliParaformerAsr, and WebSocket clients, targeting .NET 6 and 8 with packages like NAudio, ONNX Runtime, and Websocket.Client. An OpenClaw plugin integration is added with Node.js dependencies (ws, esbuild, vitest). Additionally, build and test manifests are included for the Android client (Gradle), iOS app (CocoaPods), and the product website (Vue.js, Playwright browser tests).

(dependencies) · high confidence

Updated bundled glog library to version 0.7.0

The ONNX Runtime's bundled glog logging library has been updated to version 0.7.0. This update includes a new CMake build system, updated source files for core logging and demangling functionality, and improved configuration headers to better support various build environments and platforms.

_runtime/onnxruntime/third\party/glog · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

This is the PUBLIC form of this artifact. Findings are listed in full, but the details of SECURITY findings — which rule fired, in which file, on which line, and how to fix it — are deliberately withheld, and any secret-scanner results are excluded entirely. Where detail is absent here it was REMOVED FOR PUBLICATION; it is not missing from the analysis. The complete artifact is available from the repository owner.

Score

  • CAI 45 → 45 (-0.5)
  • Rubric changed (rubric-2026.09.15 → rubric-2026.10.1) — scores are not directly comparable.

Lenses

  • Code Health 40 → 40 (+0.3)
  • Architecture 95 → 82 (-12.2)
  • Maturity 67 → 65 (-1.9)
  • Readiness 35 → 35 (-0.2)
  • Security 62 → 63 (+1.3)

Resolved (14)

  • Documentation: no installation or build instructions (README.md)
  • Documentation: no licence statement (README.md)
  • Documentation: no usage examples (README.md)
  • Documentation: written for insiders (integrations/openclaw/README.md)
  • Duplicated block (8 lines × 2) (runtime/csharp/AliParaformerAsr/AliParaformerAsr.Examples.MauiApp/RecognitionForFiles.xaml.cs)
  • High CVE: [GHSA redacted] (web-pages/package-lock.json)
  • High: security finding (details withheld)
  • Inconsistent naming for what appears to be the same or related concept within the same type. 'Token_nums' and 'Token_nums_length' use different suffixes ('nums' vs 'nums_length') which creates ambiguity about whether they represent the count itself or a length metric, especially when used together.
  • Medium CVE: [GHSA redacted] (web-pages/package-lock.json)
  • Off-boarding risk: anonymized user #1
  • Orphaned knowledge (fun_text_processing/text_normalization/normalize.py)
  • Typo in 'sanm_shfit' (should likely be 'sanm_shift'). While 'selfattention' is a compound word, 'shfit' is a clear misspelling of 'shift'. This is a spelling inconsistency within the codebase's vocabulary.
  • Typo in the type name 'OnnxRumtimeTypes' (should be 'OnnxRuntimeTypes'). The word 'Runtime' is misspelled as 'Rumtime'.
  • redundant comment (runtime/csharp/AliParaformerAsr/AliParaformerAsr.Examples.MauiApp/Platforms/Windows/App.xaml.cs)

New (1812)

  • (anonymous) (cognitive 16) (web-pages/product-site/assets/js/experience.js)
  • AliFsmnVad.GetSegmentsByStep (cognitive 35) (runtime/csharp/AliFsmnVad/AliFsmnVadSharp/AliFsmnVad.cs)
  • ApplyTimestampRules.apply (cognitive 23) (funasr/models/sense_voice/whisper_lib/decoding.py)
  • AudioDatasetHotword.collator (cognitive 37) (funasr/datasets/audio_datasets/datasets.py)
  • AudioDatasetHotword.collator (cyclomatic 19) (funasr/datasets/audio_datasets/datasets.py)
  • AutoModel.init (cognitive 19) (funasr/auto/auto_model.py)
  • AutoModel.init (cyclomatic 16) (funasr/auto/auto_model.py)
  • AutoModel.build_model (cognitive 50) (funasr/auto/auto_model.py)
  • AutoModel.build_model (cyclomatic 32) (funasr/auto/auto_model.py)
  • AutoModel.inference (cognitive 26) (funasr/auto/auto_model.py)
  • AutoModel.inference (cyclomatic 17) (funasr/auto/auto_model.py)
  • AutoModel.inference_with_vad (cognitive 161) (funasr/auto/auto_model.py)
  • AutoModel.inference_with_vad (cyclomatic 80) (funasr/auto/auto_model.py)
  • BaseTokenizer.init (cognitive 23) (funasr/tokenizer/abs_tokenizer.py)
  • BeamSearchDecoder.update (cognitive 24) (funasr/models/sense_voice/whisper_lib/decoding.py)
  • BeamSearchScama.forward (cognitive 21) (funasr/models/scama/beam_search.py)
  • BeamSearchScama.forward (cognitive 21) (funasr/models/uniasr/beam_search.py)
  • BeamSearchScamaStreaming.forward (cognitive 20) (funasr/models/scama/beam_search.py)
  • BeamSearchTransducer.align_length_sync_decoding (cognitive 29) (funasr/models/transducer/beam_search_transducer.py)
  • BeamSearchTransducer.default_beam_search (cognitive 22) (funasr/models/transducer/beam_search_transducer.py)
  • …and 1792 more

Changes since last survey

  • 46 commits — 27 feature/other, 19 fixes

By area

  • (repo) — 16 commits
  • .github/workflows — 7 commits
  • funasr/models — 4 commits
  • funasr/utils — 4 commits
  • docs/operations — 2 commits
  • examples/industrial_data_pretraining — 2 commits
  • funasr/auto — 2 commits
  • runtime/python — 2 commits
  • web-pages/product-site — 2 commits
  • docs/ascend_npu.md — 1 commit
  • docs/python_api.md — 1 commit
  • docs/troubleshooting.md — 1 commit
  • funasr/bin — 1 commit
  • funasr/train_utils — 1 commit

Notable commits

  • fix: Fix invalid escape sequences in text normalization and models (#3709)
  • fix: Merge main with timestamp and event mapping fixes
  • fix: ci: install ffmpeg for audio byte regression coverage
  • fix: fix(amp): resolve autocast/GradScaler device type for non-CUDA accelerators (#3735)
  • fix: fix(auto): warn when unavailable accelerators fall back to CPU
  • fix: fix(auto): warn when unavailable accelerators fall back to CPU (#3742)
  • fix: fix(hotwords): stop fuzzy postprocess matching from rewriting nearby text (#3744)
  • fix: fix(realtime): recover merged long-segment partials (#3755)
  • fix: fix(runtime): skip empty deduplicated event segments
  • fix: fix(runtime): skip empty deduplicated event segments (#3748)
  • fix: fix(sensevoice): keep words separate at standalone boundaries
  • fix: fix(sensevoice): normalize first word before joining subpieces
  • fix: fix(sensevoice): stop rendering the Cough event as the Sneeze emoji (#3747)
  • fix: fix(website): correct LiteLLM transcription route and guards
  • fix: fix: avoid deep-copying the checkpoint in load_pretrained_model (#3729)
  • fix: fix: count speaker overlap once per diarization segment
  • fix: fix: keep token separators in timestamp_sentence raw_text (#3746)
  • fix: fix: stop timestamp_sentence_en reusing a stale sentence start (#3741)
  • fix: fix: use y.device instead of hardcoded .cuda() in data2vec compute_var (#3710)
  • change: Merge pull request #3733 from modelscope/codex/speaker-overlap-accounting-20260928
  • …and 26 more

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

modelscope/FunASR was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 4 October 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit 66d7a4c264a5993a2a63ed00c1f402c296ee521a — the exact code this score is about.
  • Scored under rubric-2026.10.1 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-b94e107d0cec.