Skip to content
CAI
Software that uses CAICheck a score

soniqo/speech-swift

59.2

Adequate · 1 October 2026

144.8k

lines of production code

Swift

primary language

2

measurements over time

CAI band scale
CAI trend line
CAI lens gauges

What this system is

This system is a comprehensive Swift library for running local, on-device speech and audio AI models on Apple platforms. It provides a unified framework for automatic speech recognition, text-to-speech synthesis, voice cloning, and full-duplex voice conversations using various open-source models. The library includes utilities for audio processing, benchmarking, and streaming, along with demo applications and an OpenAI-compatible server interface.

Features

Add CSM (Conversational Speech Model) runtime

Introduces a new runtime for the Sesame CSM-1B model, enabling text-to-speech synthesis with voice cloning capabilities. The implementation includes the core CSMModel architecture (backbone and decoder towers), a Mimi codec loader for 32-codebook audio tokenization, and a CSMPipeline that handles tokenization, autoregressive frame generation, and waveform decoding. Users can now load pre-trained CSM models from Hugging Face via \CSMPipeline.fromPretrained\ and generate 24 kHz audio from text prompts using a reference audio clip for voice matching.

Sources/CSM · high confidence

Add Canary and Parakeet ASR models with CoreML support

Introduces two new offline automatic speech recognition engines: NVIDIA Canary (180M Flash) and Parakeet TDT (0.6B v3). Both models run entirely on-device using CoreML, with audio preprocessing (STFT, mel spectrogram) implemented in Swift using Accelerate/vDSP. Canary supports English, German, Spanish, and French with translation capabilities, while Parakeet supports 25 European languages natively. Both models include language steering, confidence scoring, and memory management APIs. Parakeet offers iOS-optimized variants with reduced memory footprint and supports both fixed-shape and enumerated-shape CoreML exports.

Sources/ParakeetASR · high confidence

Add Cohere Transcribe and Voxtral ASR models

Users can now perform speech recognition using two new MLX-based models: Cohere Transcribe and Voxtral. The Cohere Transcribe implementation (Sources/CohereTranscribeASR) includes a 16 kHz audio frontend with Slaney filterbanks and per-utterance mean-variance normalization, a Conformer encoder, and a Transformer decoder with KV-cache support, available in FP16, INT5, and INT8 quantizations. The Voxtral implementation (Sources/VoxtralASR) features a multi-modal architecture with an audio tower, a language model, and a multi-modal projector, also supporting FP16, INT5, and INT8 variants. Both modules provide public APIs for transcribing audio with optional language specification and decoding options.

Sources/CohereTranscribeASR, Sources/VoxtralASR · high confidence

Add CoreML-based Kokoro-82M TTS with multilingual phonemization

Introduces on-device text-to-speech using the Kokoro-82M model compiled to CoreML, enabling synthesis for English, Chinese, Japanese, Hindi, French, Spanish, Portuguese, and Italian. The implementation includes language-specific phonemizers (leveraging Apple's CFStringTokenizer for Asian languages and rule-based G2P for Latin scripts), pronunciation dictionaries, and a voice embedding system supporting 54 preset voices. It also features a post-processing step to remove trailing audio artifacts and a memory management interface for model unloading.

Sources/KokoroTTS · high confidence

Add DictateDemo: macOS menu bar streaming dictation

Introduces a new macOS demo application that provides continuous, streaming speech-to-text dictation directly from the menu bar. The app features a menu bar extra for starting and stopping recording, along with a floating HUD window that displays real-time transcription, sentence boundaries, and a word count. It leverages the Parakeet streaming ASR model and Silero VAD to process raw microphone audio, handling end-of-utterance detection and finalizing transcripts based on voice activity rather than just silence, ensuring smooth multi-sentence dictation.

Examples/DictateDemo/DictateDemo · high confidence

Add Fish Audio S2 Pro runtime

Introduces a new local text-to-speech runtime for the Fish Audio S2 Pro model (default ID: aufklarer/Fish-Audio-S2-Pro-MLX-fp16) within the FishAudioTTS module. This feature enables on-device speech generation using the MLX framework, supporting both direct text-to-speech and voice cloning via reference audio prompts. The implementation includes the full inference pipeline: a dual autoregressive language model for codebook generation, a neural codec for audio decoding, and a tokenizer, with configurable sampling parameters such as temperature, top-k, and repetition penalty.

Sources/FishAudioTTS · high confidence

Add LocalVQE acoustic echo cancellation and DeepFilterNet3 speech enhancement

Introduces two new audio processing capabilities to the SpeechEnhancement module. First, a LocalVQE-based acoustic echo canceller (AEC) is added, featuring a C++ adaptive filter frontend, a Core ML residual mask model, and Swift wrappers to process microphone and playback reference streams for echo removal. Second, DeepFilterNet3 speech enhancement is implemented with dual backend support: a Core ML path for on-device Neural Engine inference and an MLX path for GPU-accelerated inference via Metal, both sharing a common signal processing pipeline (STFT, ERB filterbank, deep filtering) implemented in Swift.

Sources/SpeechEnhancement · high confidence

Add MLX-based source separation models (HTDemucs and Open-Unmix)

Introduces two new music source separation engines powered by MLX. The Hybrid Transformer Demucs (Demucs v4) implementation provides a bag-of-models separator that loads fine-tuned sub-models for vocals, drums, bass, and other stems, supporting both fp16 and int8 precision bundles with overlapping window inference. The restored Open-Unmix module implements a BiLSTM-based architecture that separates stereo audio into the same four stems using vDSP-accelerated STFT processing and fused LSTM inference steps.

Sources/SourceSeparation · high confidence

Add Magpie TTS (9 languages, MLX INT8)

Introduces the Magpie-TTS Multilingual 357M model for Apple Silicon via MLX, supporting English, Spanish, German, French, Italian, Vietnamese, Chinese, Hindi, and Japanese. The implementation includes language-specific grapheme-to-phoneme (G2P) pipelines for Chinese, English, German, Spanish, and Japanese, as well as a byte encoder for French, Italian, and Vietnamese. Model weights are loaded in INT8 quantization from Hugging Face, with baked speaker identities for Sofia, Aria, Jason, Leo, and John Van Stan. The system supports both batch and streaming synthesis at 22.05 kHz.

Sources/MagpieTTS · high confidence

Add Qwen3-TTS 1.7B CoreML export and inference tooling

Introduces a new Python-based pipeline in \scripts/qwen3\_tts\_coreml\ to convert the Qwen3-TTS 1.7B model into six optimized CoreML components (TextProjector, CodeEmbedder, MultiCodeEmbedder, CodeDecoder, MultiCodeDecoder, and SpeechDecoder). The \convert\_coreml.py\ script handles the export with specific memory and numerical stability fixes, such as using explicit cache inputs for the CodeDecoder to avoid GPU backend state-mutation hazards and defaulting to FP32 precision for the decoder to ensure numerical accuracy. A reference runner (\run\_coreml.py\) and a validation suite (\validate\_coreml.py\, \test\_coreml\_variants.py\, \test\_state\_writes.py\) are included to execute inference and verify that the CoreML models match the original PyTorch behavior across different compute routes (CPU, GPU, ANE).

_scripts/qwen3\_tts\coreml · high confidence

Add SupertonicTTS-3 CoreML support

Introduces a new non-autoregressive flow-matching TTS model (SupertonicTTS-3) for Apple platforms via CoreML. This addition includes the model implementation, configuration, and a G2P-free tokenizer that handles text preprocessing (NFKD normalization, language wrapping) and chunking to fit the model's fixed text length constraints. The model supports 31 languages and multiple voices, synthesizing audio at 44.1 kHz.

Sources/SupertonicTTS · high confidence

Add VoiceChat 11B full-duplex runtime for MLX

Introduces a complete local VoiceChat 11B perception bundle for MLX, enabling real-time, full-duplex speech-to-speech inference. This includes a streaming FastConformer encoder with causal subsampling and context-aware attention, a 56-layer Nemotron-H language backbone adapted for fused audio/text embeddings, an EAR-TTS speech decoder, and a neural audio codec. The implementation handles model loading, quantization, and session management, allowing users to start stateful duplex conversations with native function calling support.

Sources/VoiceChat · high confidence

Add audio-to-avatar motion generation via NVIDIA Audio2Face-3D

Users can now generate 3D avatar motion coefficients from speech audio using the NVIDIA Audio2Face-3D model. This change introduces the Audio2Face3D library module, which includes an MLX runtime for on-device inference, a downloader for model weights, and configuration structures for identity-specific coefficient layouts (Mark, Claire, James). A new \avatar-motion\ CLI command is added to the \speech\ executable, allowing users to input audio and output JSONL frames of skin, tongue, jaw, and eye coefficients, with support for both local model directories and pre-trained Hugging Face bundles.

Sources/AudioCLILib · high confidence

Add local benchmark runner and dashboard for Qwen3-ASR

Added a local benchmarking toolchain for the Qwen3-ASR model, including a Python script to sweep batch sizes (1, 2, 4, 8) and measure performance metrics like RTF and throughput, a utility to monitor Apple Silicon GPU/CPU power consumption during runs, and a script to render an HTML dashboard from persisted JSON benchmark reports.

python · high confidence

Add native F5-TTS voice cloning runtime

Introduces a new F5-TTS voice cloning capability within the Sources/F5TTS module, allowing users to generate speech by providing reference audio and text. The implementation includes the core synthesis engine (F5TTSFlow) for mel-spectrogram generation, a Vocos-based vocoder (F5TTSVocos) for waveform decoding, and a Mandarin pinyin frontend (F5TTSPinyin) that enables support for Chinese text via a bundled lexicon. The model loads from Hugging Face bundles containing DiT and Vocos weights, supporting configuration via F5TTSConfig and offering synthesis options such as speed, CFG strength, and sampling steps.

Sources/F5TTS · high confidence

Add native IndexTTS2 voice cloning runtime

Introduces a complete native implementation of the IndexTTS2 voice cloning pipeline, including the BigVGAN vocoder, S2Mel flow, semantic GPT, and CAMPPlus speaker encoder. The runtime supports reference audio conditioning for voice cloning, explicit emotion control via presets or custom vectors, and bundle loading from Hugging Face with manifest validation.

Sources/IndexTTS2TTS · high confidence

Add native MLX-based GLiNER2.5-Decide classification and entity extraction

Introduces a new Swift implementation of the GLiNER2.5-Decide model using the MLX framework, enabling on-device single-label classification and entity span extraction. The library supports downloading and loading fp32, fp16, and int8 quantized model variants from Hugging Face, and includes a benchmarking tool to measure performance and memory usage for routing and extraction tasks.

Sources/GLiNER, Sources/GLiNERBenchmark · high confidence

Add native VoxCPM2 TTS backend

Introduces a new native text-to-speech backend based on the VoxCPM2 model, implemented in Swift using the MLX framework. This addition includes the core model architecture (MiniCPM4, AudioVAE, DiT decoder), configuration structures, and a long-rope rotary embedding implementation to support extended context lengths. The backend supports multiple model variants, including a default bf16 bundle and an int8 quantized version for reduced memory footprint, and handles model and tokenizer loading from Hugging Face.

Sources/VoxCPM2TTS · high confidence

Add on-device CoreML benchmark app for iPhone

A new iOS example app (\Examples/iOSBenchmark\) has been added to run on-device performance benchmarks for CoreML models. The app measures Real-Time Factor (RTF) for streaming ASR (Parakeet-EOU, Omnilingual) and TTS (Supertonic-3, Kokoro-82M), as well as tokens-per-second for the FunctionGemma LLM, while tracking peak memory usage. Results are displayed in the UI and saved to \results.json\, providing a standardized way to evaluate model efficiency on physical devices.

Examples/iOSBenchmark · high confidence

Chatterbox TTS adds Core ML runtime and MLX speaker encoder

Chatterbox TTS now supports a Core ML runtime for Apple Silicon, enabling on-device inference via new \ChatterboxFlashCoreMLModel\ and associated graph/config files. The library also introduces an MLX-native speaker encoder (\CAMPPlus\) to replace external dependencies for voice cloning, and adds memory management options (\ChatterboxMemoryOptions\) to control MLX cache limits and buffer clearing during generation.

Sources/ChatterboxTTS · high confidence

Introduce AudioServer with OpenAI-compatible HTTP and Realtime WebSocket endpoints

The new AudioServer component provides a local HTTP and WebSocket API that mirrors the OpenAI audio interface, exposing POST /v1/audio/transcriptions and /v1/audio/speech alongside a /v1/realtime WebSocket endpoint for streaming. It includes a model registry that lets clients introspect available ASR, TTS, and speech-to-speech variants via /v1/models and /v1/realtime/models, and adds server-side voice activity detection to manage real-time turn boundaries. This change introduces new runtime behavior for audio routing and session management rather than modifying existing services.

Sources/AudioServer · high confidence

Introduce CosyVoice3 TTS with voice cloning, multi-speaker dialogue, and memory management

This change adds the CosyVoice3 text-to-speech model to the CosyVoiceTTS module, implementing a three-stage pipeline (LLM, DiT flow matching, HiFi-GAN vocoder) that supports zero-shot voice cloning via a new CAM++ speaker encoder and multi-speaker dialogue synthesis with inline emotion tags. The model now loads bf16 or 8-bit quantized bundles, uses a fixed flow noise buffer for deterministic output, and exposes an unload() API to manage memory footprint. Dialogue input can now use \[Speaker\] and (emotion) tags to control speaker identity and style, with automatic crossfading between segments.

Sources/CosyVoiceTTS · high confidence

Introduce HibikiTranslate module for Zero-3B speech-to-speech translation

Adds the HibikiTranslate module, implementing the Hibiki Zero-3B model for translating French, Spanish, Portuguese, and German speech into English. This change introduces the core Swift implementation including configuration structures, a temporal transformer with Grouped-Query Attention and rope\_concat positional embeddings, a scheduled depformer for autoregressive codebook generation, and a driver that handles Mimi audio encoding/decoding. The implementation fixes critical generation behaviors such as pre-filling the token cache with correct initial tokens, aliasing EOS to PAD to prevent early truncation, and applying per-codebook delay un-shifting before decoding, ensuring the translation pipeline produces coherent output.

Sources/HibikiTranslate · high confidence

Introduce OmniVoiceTTS: MLX-based multilingual voice cloning

Added a new voice-cloning TTS module in Sources/OmniVoiceTTS that implements the OmniVoice model using MLX. This includes a Qwen3-based diffusion backbone, a Higgs-audio v2 codec for encoding and decoding audio, and a text front-end that supports over 600 languages and style instructions. The module handles downloading and loading model weights from Hugging Face, synthesizing speech from a reference clip and target text, and applying edge fades to the output.

Sources/OmniVoiceTTS · high confidence

Introduce Omnilingual ASR with CoreML and MLX backends

Added the Omnilingual ASR module, providing a language-agnostic speech recognition model that supports over 1,600 languages. The implementation includes a CoreML backend for on-device inference using the Meta CTC-300M model (with 5s and 10s window variants) and a new MLX backend supporting wav2vec2 variants (300M, 1B, 3B, 7B) with 4-bit and 8-bit quantization. Both backends conform to the \SpeechRecognitionModel\ protocol, handle audio resampling and layer normalization, and use a greedy CTC decoder with SentencePiece detokenization.

Sources/OmnilingualASR · high confidence

Introduce Parakeet EOU 120M streaming ASR with end-of-utterance detection

Adds a new streaming speech recognition module based on the Parakeet EOU 120M CoreML model, enabling real-time transcription with end-of-utterance (EOU) detection for continuous dictation. The implementation includes a streaming session that processes audio in chunks while maintaining encoder cache and LSTM decoder state, a mel spectrogram preprocessor optimized for streaming, and an RNNT greedy decoder that emits partial transcripts and final segments upon detecting EOU tokens. The module supports automatic language detection for a set of major languages and provides memory management utilities for the \~150 MB model footprint.

Sources/ParakeetStreamingASR · high confidence

Introduce PersonaPlex Swift/MLX inference engine with streaming and quantization

Adds the PersonaPlex speech-to-speech model implementation for Swift/MLX, enabling full-duplex, streaming audio generation on Apple Silicon. The new codebase includes the core model architecture (Temporal Transformer, Depformer, Mimi codec), streaming-optimized components like pre-allocated and ring KV caches, and support for 4-bit and 8-bit quantization. It also provides memory management APIs (unload, footprint tracking) and centralizes cache directory resolution to ensure voice prompts and tokenizers load correctly across different model variants.

Sources/PersonaPlex · high confidence

Introduce Qwen3-TTS CoreML inference with ANE-optimized chunked decoding

Adds a new \Qwen3TTSCoreML\ module that runs Qwen3-TTS (0.6B and 1.7B) on Apple Silicon using CoreML. The synthesizer routes text and codec embedding lookups to the CPU for FP32 precision, while the decoder and speech-decoder models run on the Neural Engine (or CPU for the 1.7B variant). To ensure the complex transformer graphs fit on the ANE, the CodeDecoder and MultiCodeDecoder are automatically split into smaller, stateless chunks with a separate head model when available. The module also includes a pure-Swift token sampler, NPY embedding file validation, and a \SpeechGenerationModel\ protocol extension for easy integration.

Sources/Qwen3TTSCoreML · high confidence

Introduce Qwen3-TTS inference engine with voice cloning and streaming support

Adds the Qwen3-TTS text-to-speech implementation, including the CodePredictor model, configuration structures, and sampling logic. This change enables high-fidelity voice cloning via In-Context Learning (ICL) and standard x-vector methods, supports streaming synthesis with incremental audio output, and provides batch processing for parallel text synthesis. It also introduces cooperative cancellation for generation tasks, memory management APIs to unload models, and a content-addressed cache for voice-cloning reference artifacts to optimize repeated use of the same speaker.

Sources/Qwen3TTS · high confidence

Introduce SpeechDemo app with Dictate, Speak, and Echo tabs

The SpeechDemo app now provides three distinct interaction modes: Dictate for real-time speech-to-text transcription using Parakeet TDT (and Qwen3-ASR on macOS), Speak for text-to-synthesis using Qwen3-TTS (macOS) or system speech (iOS), and Echo for a full voice-pipeline demo with VAD, ASR, TTS, and Smart Turn end-of-turn detection (macOS). The Echo tab includes an echo reference gate to suppress TTS playback leaking into the microphone, and users can toggle Smart Turn to confirm end-of-turn decisions via a probability classifier.

Examples/SpeechDemo/SpeechDemo · high confidence

Introduce VibeVoiceTTS module for Microsoft VibeVoice long-form text-to-speech

Adds the new VibeVoiceTTS module, providing an MLX-based implementation of the Microsoft VibeVoice long-form text-to-speech model (Realtime 0.5B and 1.5B). This release includes the full inference pipeline, featuring a DPM-Solver diffusion scheduler, streaming causal 1D convolutions for low-latency audio generation, and a KV-cache system for autoregressive language model steps. It also provides weight loading utilities that map and transpose model checkpoints, along with the necessary neural network layers (normalization, timestep embedding, acoustic tokenizers) to run the synthesis.

Sources/VibeVoiceTTS · high confidence

Introduce XGrammar bridge for JSON schema-constrained decoding

Added a new C-language bridge (CXGrammarBridge) that wraps the XGrammar library to enforce JSON-schema response formats. This component compiles JSON schemas into grammar rules and provides a bitmask-based matcher to restrict token selection, ensuring that generated outputs strictly adhere to the defined schema structure.

Sources/CXGrammarBridge · high confidence

Introduce local ASR benchmarking tool with multi-engine support

A new command-line benchmark runner (asr-bench) has been added to evaluate Automatic Speech Recognition engines on a shared dataset, reporting Word Error Rate (WER), Character Error Rate (CER), and Real-Time Factor (RTF). The tool supports a wide range of engines including Qwen3 (CoreML and MLX variants), Cohere Transcribe, Voxtral, Moss (CoreML and MLX), Nemotron (CoreML and MLX), Omnilingual, Parakeet, and Whisper (ASR and WhisperKit). It handles dataset loading from LibriSpeech-style directories or TSV manifests, normalizes text for fair comparison, and tracks memory usage (RSS and physical footprint) to provide a comprehensive performance profile for each engine.

Sources/AsrBenchmark · high confidence

Introduce local diarization benchmark runner

A new command-line tool, \diarization-bench\, is available to evaluate diarization engines against audio manifests with RTTM references. It supports multiple backends including CoreML and MLX variants for Sortformer, MOSS, Nemotron 3, and Pyannote, and reports detailed metrics such as Diarization Error Rate (DER), Jaccard Error Rate (JER), speaker count accuracy, and throughput (xRT). Users can configure specific engines, tuning parameters, and output reports in JSON format.

Sources/DiarizationBenchmark · high confidence

Introduce native Apple Silicon diarization with Community-1 Core ML and Silero VAD

The SpeechVAD module now includes a complete native implementation of the pyannote Community-1 speaker diarization pipeline, leveraging Core ML for segmentation and embedding extraction while performing PLDA transformation, VBx clustering, and timeline reconstruction in Swift. This addition brings high-performance, on-device diarization to Apple platforms, supported by a Core ML-optimized Silero VAD for voice activity detection and a DER scoring utility for evaluating diarization accuracy against ground truth.

Sources/SpeechVAD · high confidence

Introduce native Nemotron-3.5 streaming ASR with CoreML and MLX backends

Users can now perform low-latency, multilingual streaming speech recognition on Apple Silicon using the Nemotron-3.5 ASR 0.6B model. This change adds a new \NemotronStreamingASR\ module that supports two native inference backends: a CoreML runtime for general Apple devices and an MLX runtime for optimized performance on Apple Silicon. The implementation features a cache-aware FastConformer encoder and a prompt-conditioned RNN-T decoder, enabling incremental text output as audio streams in. It supports 76 languages via a language prompt mask, native punctuation and capitalization, and includes word-boosting capabilities to bias decoding toward specific phrases. The module handles audio resampling, streaming mel-spectrogram preprocessing, and provides both full-buffer transcription and real-time streaming APIs.

Sources/NemotronStreamingASR · high confidence

Introduce on-device FunctionGemma 270M tool-calling LLM

Adds a new on-device LLM component (FunctionGemma) that runs the 270M CoreML model to perform single-turn function calling. The implementation includes prompt formatting that mirrors the model's training grammar, a recursive-descent parser to extract structured tool calls from the model's output, and a pipeline bridge allowing VoicePipeline to use FunctionGemma as its LLM stage. It also supports loading models from Hugging Face with environment-variable-based endpoint mirroring.

Sources/FunctionGemma · high confidence

Introduce on-device wake-word detection using CoreML Zipformer

Adds a new SpeechWakeWord module that performs streaming keyword spotting on-device using a 3.49M-parameter KWS Zipformer model exported to CoreML. The module includes a Kaldi-compatible fbank feature extractor, a Viterbi-based BPE tokenizer to correctly decompose keyword phrases, and a modified beam-search decoder with an Aho-Corasick context graph to boost registered keywords. Users can load the model from Hugging Face, register custom wake-word phrases with per-phrase thresholds and boosts, and process audio streams to receive keyword detections with timestamps.

Sources/SpeechWakeWord · high confidence

Introduces device-aware memory tiering and a configurable voice pipeline

SpeechCore now includes a MemoryTier system that automatically detects available RAM on iOS and macOS to select an appropriate model configuration (full, standard, constrained, or minimal), preventing out-of-memory errors by falling back to Apple Speech/TTS or disabling on-device LLMs on lower-end devices. Additionally, a new VoicePipeline class provides a high-level Swift wrapper for the C++ pipeline engine, managing the full conversational loop (audio → VAD → STT → LLM → TTS → audio) with configurable modes (voice, transcribe-only, echo), end-of-turn detection, and event callbacks.

Sources/SpeechCore · high confidence

New AudioCommon module with unified audio I/O, streaming, and CoreML loading

The new AudioCommon module consolidates audio utilities into a single location, providing a reusable AudioIO manager for microphone capture and playback with optional acoustic echo cancellation, a pull-based AudioFileStream for bounded-memory file decoding, and a CoreMLLoader that instruments model loading to warn users when the Neural Engine silently falls back to CPU. It also introduces a thread-safe AudioRingBuffer using os\_unfair\_lock for real-time audio thread safety, a unified AudioModelError type, and a CoreMLComputeUnitsResolver that honors the SPEECH\_COREML\_COMPUTE\_UNITS environment variable to prevent CI hangs.

Sources/AudioCommon · high confidence

New CoreML ASR and ForcedAligner backends for Qwen3-ASR

This change introduces new CoreML-based inference backends for the Qwen3-ASR module, enabling full speech recognition and word-level timestamp alignment on Apple Neural Engine hardware. It adds \CoreMLASRModel\ (with separate compute unit configuration for the encoder and decoder), \CoreMLForcedAligner\ (supporting both FP16 and 8-bit quantized variants), and \CoreMLTextDecoder\ (utilizing MLState KV caches). These components allow the library to run the entire Qwen3-ASR pipeline on CoreML, reducing reliance on MLX GPU execution and lowering power consumption on macOS and iOS devices.

Sources/Qwen3ASR · high confidence

New CoreML backend for Magpie TTS

Added a new CoreML-based inference engine for the Magpie TTS model, accessible via the \--engine magpie-coreml\ CLI flag. This implementation uses compiled CoreML bundles (text encoder, decoder, and nano-codec) to run on Apple Silicon, leveraging the Neural Engine for significantly faster synthesis compared to the previous MLX backend. It supports the same set of speakers and languages (excluding Japanese, which falls back to the MLX backend) and includes optimized streaming capabilities with reduced first-packet latency.

Sources/MagpieTTSCoreML · high confidence

New Gemma 4 chat backend and Qwen3 quantization/formatting support

This location introduces a complete Gemma 4 chat backend (Gemma4Chat, Gemma4Model, Gemma4Tokenizer) alongside updated Qwen3 infrastructure. The Gemma 4 backend implements the model's specific architecture (sliding attention, KV-sharing, Proportional RoPE) and enforces a new chat template that suppresses the reasoning channel during decoding. It also adds constrained JSON-schema response formats via a Swift token matcher and on-device MLX sampling. For Qwen3, the module adds a unified ChatQuantization resolver to correctly read INT4/INT5/INT8 from various config formats, a ChatSampler with optimized top-K selection, and a ChatTemplate that handles Qwen3.5's thinking blocks.

Sources/Qwen3Chat · high confidence

New MLX inference primitives and audio processing utilities

This update introduces a suite of foundational components for the MLX inference engine. It adds GPU memory management tools, including a Metal budget API and memory pinning to prevent paging under pressure. Core model architecture support is expanded with an LSTM cell for EnCodec audio encoding, quantized MLP and embedding layers, and a shared scaled dot-product attention (SDPA) helper. Additionally, it includes a Slaney mel spectrogram front-end for audio analysis and robust utilities for loading and verifying quantized model weights.

Sources/MLXCommon · high confidence

New OpenAI Realtime API WebSocket client example

Added a new HTML-based client example (\Examples/websocket-client.html\) that demonstrates how to interact with the OpenAI Realtime API via WebSocket. This tool allows users to connect to a local server endpoint (defaulting to \ws://127.0.0.1:8080/v1/realtime\) to perform real-time speech recognition (ASR) by recording or uploading audio, and text-to-speech (TTS) synthesis. It supports selecting different engines (CosyVoice, Qwen3-TTS) and languages, and provides a visual log of session events, audio chunks, and transcripts.

Examples · high confidence

New SpeechUI module for streaming ASR transcripts

A new \SpeechUI\ module has been added to provide minimal SwiftUI components for streaming speech recognition applications. It includes a \TranscriptionView\ that displays a scrolling transcript, distinguishing between committed final lines and the in-progress partial line, with optional auto-scrolling. It also provides a \TranscriptionStore\ class that acts as an observable model to manage final and partial text state, allowing developers to easily wire up any streaming ASR backend to the UI.

Sources/SpeechUI · high confidence

New benchmarking support for diarization and VAD scoring

Added the BenchmarkSupport module, which provides utilities for loading benchmark manifests and reference data (including RTTM and simple segment formats), scoring Voice Activity Detection (VAD) results against references, and aggregating diarization error rates (DER) by speaker count. This enables users to run local benchmarks and view performance metrics grouped by the number of speakers in the reference recordings, rather than just a single corpus-wide score.

Sources/BenchmarkSupport · high confidence

New build, CI, and benchmark tooling scripts

This change introduces a suite of new scripts to the repository to improve build reliability, testing isolation, and benchmarking capabilities. The \build\_mlx\_metallib.sh\ script now builds the MLX Metal shader library with a content-hash cache to prevent stale artifacts, while \check\_demos.sh\ verifies that demo packages compile. CI stability is enhanced by \ci\_resolve.sh\, which heals corrupt SwiftPM dependency caches, and \test\_e2e\_isolated.sh\ / \test\_shard\_isolated.sh\, which run E2E tests in separate processes to prevent memory exhaustion. Additionally, \run\_benchmarks.sh\ provides a local runner for ASR, VAD, and diarization benchmarks, and several Python scripts (\export\_moss\_mlx.py\, \nemotron\_spm\_parity.py\, \probe\_magnet\_logits.py\, \profile\_gliner\_memory.py\, \prepare\_fleurs\_asr\_manifest.py\) support model exports, tokenizer parity checks, and memory profiling.

scripts · high confidence

New iOS Echo Demo app with on-device voice pipeline

The Examples/iOSEchoDemo location now contains a complete iOS application that demonstrates an on-device voice echo pipeline. Users can speak into the microphone and hear the response immediately. The demo uses Parakeet ASR and Kokoro TTS (CoreML) on physical devices, and Parakeet ASR with Apple's built-in TTS on the simulator. It includes voice activity detection (Silero VAD), end-of-turn detection (Smart Turn v3.2), adaptive echo prevention, and a diagnostics view showing CPU, memory, and VAD levels. The app automatically downloads models (\~500 MB) from HuggingFace on first launch.

Examples/iOSEchoDemo · high confidence

PersonaPlexDemo adds full-duplex voice interaction and offline caching

The PersonaPlexDemo app now supports a Full-Duplex mode that allows simultaneous listening and speaking, controlled via a new Mode picker in the SwiftUI interface. This mode uses a continuous audio capture path with optional echo suppression to prevent feedback when using speakers. The demo also introduces an offline cache policy that validates model file sizes and existence, allowing the app to load models from local storage instead of downloading them on every run. Additionally, the app includes a SentencePiece decoder to display the model's internal text tokens as transcripts and uses streaming audio playback for low-latency output.

Examples/PersonaPlexDemo · high confidence

Removals

Removal of legacy Qwen3-ASR CLI entry point

The standalone \main.swift\ entry point for the legacy \qwen3-asr-cli\ executable has been removed. This file previously handled direct command-line invocation for the 0.6B model, including argument parsing, model loading, and transcription logic. Its removal indicates that this specific CLI interface is no longer maintained or is being superseded by a unified command structure.

Sources/Qwen3ASRCLI · high confidence

Behavioural changes

The .codex directory now includes symlinks for Swift-specific project skills (benchmark, build, review-pr, test), which point to the corresponding skills in the .claude directory. This allows the Codex agent to access and utilize these existing Claude skills for Swift-related tasks.

.codex · high confidence

Native CoreML Whisper large-v3 turbo implementation

The WhisperASR module now uses a native CoreML runtime for the Whisper large-v3 turbo model instead of a previous wrapper. This change introduces a new model loading path that downloads and caches specific CoreML bundle files (MelSpectrogram, AudioEncoder, TextDecoder, and TextDecoderContextPrefill models along with tokenizer assets) from Hugging Face, validates the cache integrity, and initializes the CoreML models with optimized compute units (CPU+GPU for MelSpectrogram, CPU+NeuralEngine for the rest). The transcription pipeline now performs native greedy decoding with language detection, handles audio resampling to 16 kHz, and includes logic to prevent repeated-word loops during generation. A new native ByteLevelTokenizer is used for decoding model outputs, replacing any prior tokenization approach.

Sources/WhisperASR · high confidence

Rename CLI binary to speech-server and introduce HTTP API server

The command-line interface binary has been renamed from \audio-server\ to \speech-server\, with the old name retained as a deprecated alias that prints a warning upon use. This change accompanies the introduction of a new HTTP API server mode, which exposes endpoints for speech-to-text, text-to-speech, speech enhancement, and OpenAI-compatible transcription and speech synthesis, along with a WebSocket endpoint for the OpenAI Realtime API.

Sources/AudioServerCLI · high confidence

Fixes

Added script to regenerate KWS fbank reference fixtures

A new Python script has been added to the scripts/kws directory to regenerate the reference Fbank fixtures used by the SpeechWakeWordTests. This tool generates a deterministic 1-second PCM WAV file and the corresponding binary Fbank reference data (100 frames × 80 mel bins) to ensure parity between the Python export pipeline and the Swift implementation.

scripts/kws · high confidence

Test coverage

Added E2E tests for DictateDemo streaming pipeline; Added comprehensive test coverage for AudioServer endpoints and components; Added comprehensive test coverage for Canary and Parakeet ASR models; Added comprehensive test coverage for Chatterbox TTS components; Added comprehensive test coverage for F5-TTS voice cloning and Mandarin pinyin support; Added comprehensive test coverage for Kokoro TTS and phonemizers; Added comprehensive test coverage for Qwen3 ASR CoreML and MLX decoding paths; Added comprehensive test suite for AudioCLI commands and model runtimes; Added comprehensive test suite for CosyVoice TTS; Added comprehensive test suite for MADLADTranslation; Added comprehensive test suite for MAGNeT music generation; Added comprehensive test suite for Magpie TTS; Added comprehensive test suite for Nemotron Streaming ASR; Added comprehensive test suite for Qwen3Chat model configuration, sampling, and decoding; Added comprehensive test suite for the OmniVoice MLX-Swift TTS port; Added configuration and end-to-end tests for Hibiki Translate Zero-3B; Added configuration and end-to-end tests for Stable Audio 3 Music Gen; Added end-to-end and configuration tests for the Magpie CoreML TTS backend; Added end-to-end and runtime tests for Indic-Mio TTS; Added end-to-end and tokenizer tests for Supertonic TTS; Added end-to-end and unit tests for DeepFilterNet3 and LocalVQE speech enhancement; Added end-to-end and unit tests for Omnilingual ASR CoreML and MLX backends; Added end-to-end tests for VoiceChat real-time inference and tool calling; Added test coverage for SpeechVAD diarization, VAD, and scoring components; Added test suite for Audio2Face3D runtime and configuration; Added test suite for Qwen3-TTS CoreML module; Added test suite for Sidon Core ML speech restoration; Added test suite for SpeechWakeWord module; Added tests for CSM runtime and pipeline; Added tests for Cohere Transcribe and Voxtral ASR models; Added tests for Fish Audio S2 Pro runtime support; Added tests for FlashSR audio super-resolution module; Added tests for FunctionGemma parser and prompt formatting; Added tests for GLiNER2.5-Decide MLX integration; Added tests for HTDemucs and Open-Unmix source separation; Added tests for IndexTTS2 TTS bundle, tokenizer, and text processing; Added tests for Whisper model bundle cache validation; Added tests for diarization benchmark aggregation and reporting; Added tests for memory tier detection, pipeline configuration, and turn completion logic; Added unit and E2E tests for SpeechUI transcription store; Added unit tests for AudioPlayer state machine; Added unit tests for Parakeet streaming ASR configuration, vocabulary, and audio preprocessing; Added unit tests for PersonaPlex configuration, sampling, and cache directory logic; Added unit tests for VibeVoice TTS configuration and quantization; Added unit tests for VoxCPM2 TTS configuration and model args; Expanded test coverage for AudioCommon audio processing and download infrastructure; Expanded test coverage for Qwen3-TTS model loading, generation, and voice cloning.

Dependencies

Major platform upgrade and library expansion in speech-swift

The speech-swift package has been renamed from Qwen3ASR to Qwen3Speech and raised its minimum platform requirements to macOS 15 and iOS 18. This update introduces a significantly expanded set of libraries, including new ASR modules (CohereTranscribeASR, VoxtralASR, CanaryASR, WhisperASR, MossTranscribe, OmnilingualASR), TTS engines (Qwen3TTS, CosyVoiceTTS, ChatterboxTTS, OmniVoiceTTS, F5TTS, HiggsTTS, IndexTTS2TTS, VibeVoiceTTS, VoxCPM2TTS, MagpieTTS, SupertonicTTS, KokoroTTS), and utility modules (AudioCommon, SpeechVAD, SpeechLanguageID, SpeechEnhancement, SpeechRestoration, SourceSeparation, SpeechCore, SpeechUI, SpeechWakeWord, GLiNER). Additionally, new demo applications (DictateDemo, PersonaPlexDemo, SpeechDemo, iOSEchoDemo) and Python scripts for CoreML export have been added to support these capabilities.

(dependencies) · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

Score

  • CAI 61 → 59 (-1.7)
  • Rubric changed (rubric-2026.09.11 → rubric-2026.09.18) — scores are not directly comparable.

Lenses

  • Code Health 80 → 78 (-1.9)
  • Architecture 100 → 93 (-7.0)
  • Maturity 67 → 67 (+0.1)
  • Readiness 48 → 43 (-4.7)
  • Security 67 → 81 (+14.7)
  • Performance 75 (new)

Resolved (42)

  • Dependency hygiene PARTLY measured — SwiftPM pinning read, dependency currency NOT established
  • Documentation: no installation or build instructions (README.md)
  • Documentation: written for insiders (docs/models/qwen3-dense-chat.md)
  • Documentation: written for insiders (docs/models/qwen35-chat.md)
  • Documentation: written for insiders (docs/models/supertonic-tts.md)
  • Documentation: written for insiders (docs/models/voicechat.md)
  • Duplicated block (10–19 lines × 2) (Sources/Qwen3ASR/Qwen3ASR.swift)
  • Duplicated block (11 lines × 2) (Sources/Qwen3TTSCoreML/Qwen3TTSCoreML.swift)
  • Duplicated block (11 lines × 3) (Sources/SpeechVAD/ReDimNet2Speaker.swift)
  • Duplicated block (13 lines × 2) (Sources/MADLADTranslation/MADLADConfig.swift)
  • Duplicated block (20 lines × 2) (Sources/IndexTTS2TTS/IndexTTS2S2MelFlow.swift)
  • Duplicated block (2–12 lines × 5) (Sources/MLXCommon/WeightLoading.swift)
  • Duplicated block (5 lines × 6) (Sources/MLXCommon/WeightLoading.swift)
  • Duplicated block (7–19 lines × 2) (Sources/VibeVoiceTTS/Models/VibeVoice15BModel.swift)
  • Edited copy of a member (27 corresponding lines) (Sources/Qwen3Chat/Qwen3ChatConfig.swift)
  • Hotspot: Examples/PersonaPlexDemo/PersonaPlexDemo/AudioRecorder.swift (Examples/PersonaPlexDemo/PersonaPlexDemo/AudioRecorder.swift)
  • Hotspot: Examples/SpeechDemo/SpeechDemo/EchoViewModel.swift (Examples/SpeechDemo/SpeechDemo/EchoViewModel.swift)
  • Hotspot: Sources/AsrBenchmark/AsrBench.swift (Sources/AsrBenchmark/AsrBench.swift)
  • Hotspot: Sources/AudioCLILib/VoiceChatMCP.swift (Sources/AudioCLILib/VoiceChatMCP.swift)
  • Hotspot: Sources/AudioCommon/FullDuplexAudioIO.swift (Sources/AudioCommon/FullDuplexAudioIO.swift)
  • …and 22 more

New (258)

  • AudioProcessing.swift.computeERBFilterbank (cognitive 16) (Sources/SpeechEnhancement/AudioProcessing.swift)
  • AudioServer.swift.dispatchSynthesize (cyclomatic 16) (Sources/AudioServer/AudioServer.swift)
  • AudioServer.swift.handleRealtimeWS (cognitive 299) (Sources/AudioServer/AudioServer.swift)
  • AudioServer.swift.handleRealtimeWS (cyclomatic 101) (Sources/AudioServer/AudioServer.swift)
  • AudioServer.swift.parseRealtimeTurnDetection (cognitive 20) (Sources/AudioServer/AudioServer.swift)
  • AudioServer.swift.parseRealtimeTurnDetection (cyclomatic 17) (Sources/AudioServer/AudioServer.swift)
  • Benchmark.main (cognitive 30) (Sources/GLiNERBenchmark/GLiNERBench.swift)
  • BenchmarkOptions.parse (cognitive 19) (Sources/GLiNERBenchmark/Options.swift)
  • BenchmarkOptions.parse (cyclomatic 17) (Sources/GLiNERBenchmark/Options.swift)
  • ClassTooLong: JSONSchemaMatcher (Sources/Qwen3Chat/JSONSchema/JSONSchemaMatcher.swift)
  • Coverage not measured — Swift suite
  • DERScoring.swift.assignedOptimalMapping (cognitive 26) (Sources/SpeechVAD/DERScoring.swift)
  • DERScoring.swift.computeDER (cognitive 21) (Sources/SpeechVAD/DERScoring.swift)
  • DERScoring.swift.maximumWeightAssignment (cognitive 33) (Sources/SpeechVAD/DERScoring.swift)
  • DERScoring.swift.maximumWeightAssignment (cyclomatic 17) (Sources/SpeechVAD/DERScoring.swift)
  • DiarizationBench.swift.maximumJaccardAssignment (cognitive 33) (Sources/DiarizationBenchmark/DiarizationBench.swift)
  • DiarizationBench.swift.maximumJaccardAssignment (cyclomatic 17) (Sources/DiarizationBenchmark/DiarizationBench.swift)
  • Duplicate functionality across types. Both AudioIO and SystemAudioTap expose identical method signatures for starting microphone capture (startMicrophone/startTimestamped). While they may target different use cases (app mic vs system capture), the API surface is nearly identical, leading to confusion about which type to use for simple microphone access.
  • Duplicated block (10 lines × 2) (Sources/SpeechEnhancement/AudioProcessing.swift)
  • Duplicated block (10 lines × 4) (Sources/SpeechVAD/Nemotron3MLXModel.swift)
  • …and 238 more

Changes since last survey

  • 28 commits — 24 feature/other, 4 fixes

By area

  • (repo) — 11 commits
  • (root) — 3 commits
  • Sources/Qwen3Chat — 3 commits
  • scripts/qwen3_tts_coreml — 2 commits
  • .github/workflows — 1 commit
  • Sources/CXGrammarBridge — 1 commit
  • Sources/ParakeetASR — 1 commit
  • Sources/Qwen3TTSCoreML — 1 commit
  • Sources/SpeechVAD — 1 commit
  • Tests/Qwen3ASRTests — 1 commit
  • Tests/SpeechVADTests — 1 commit
  • docs/benchmarks — 1 commit
  • scripts/build_mlx_metallib.sh — 1 commit

Notable commits

  • fix: Merge pull request #484 from dreamweald/fix/qwen3-asr-cancellation
  • fix: Merge pull request #487 from soniqo/fix/metal15-pinned-runtime
  • fix: Merge pull request #493 from soniqo/fix/gemma4-no-reasoning-channel
  • fix: fix(qwen3-asr): make cooperative cancellation an explicit throwing API
  • change: Add GLiNER2.5-Decide: native MLX classification and entity spans (#494)
  • change: Add JSON-schema response formats with a Swift token matcher
  • change: Add Nemotron 3 diarization with Core ML and MLX backends
  • change: Add incremental Nemotron 3 diarization session
  • change: Add reproducible Qwen3-TTS 1.7B CoreML export
  • change: Carry dlpack's licence beside the vendored header
  • change: Cover minimal cache capacity in CoreML tracing
  • change: Document Nemotron 3 in all READMEs and update diarization results
  • change: Enforce JSON-schema response formats with XGrammar
  • change: Expose per-word timestamps from Parakeet TDT decoding
  • change: Merge pull request #488 from soniqo/feat/qwen17-coreml-export
  • change: Merge pull request #489 from soniqo/feat/qwen17-coreml-runtime
  • change: Merge pull request #490 from soniqo/feat/nemotron3-diarization
  • change: Merge pull request #491 from soniqo/feat/nemotron3-streaming-session
  • change: Merge pull request #492 from soniqo/feat/json-schema-response-format
  • change: Merge pull request #496 from netlemur/word-timestamps
  • …and 8 more

Architecture

  • Unchanged — 0 containers · 1 contexts · 0 edges

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

soniqo/speech-swift was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 1 October 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit ca382eec35c3675be9670081612e19ad9496f4c7 — the exact code this score is about.
  • Scored under rubric-2026.09.18 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-e569280dd5e2.