FluidInference/FluidAudio
62.2
Adequate · 30 September 2026
102.3k
lines of production code
Swift
primary language
2
measurements over time
What this system is
FluidAudio is a Swift SDK for macOS and iOS that provides on-device speech and audio processing capabilities using CoreML. It supports automatic speech recognition, speaker diarization, voice activity detection, and text-to-speech synthesis through a variety of integrated neural models. The system also includes utilities for audio enhancement, inverse text normalization, and decision scoring, all designed to run locally without external dependencies.
Features
Add Cohere Transcribe ASR with mixed-precision support and long-form transcription
Users can now transcribe audio using the Cohere model via a new Core ML pipeline that supports mixed-precision execution (INT8 encoder with FP16 decoder) for optimized performance. The implementation includes a sliding-window mechanism for long-form audio (up to 35 seconds per chunk with 5-second overlap) and supports 14 languages. A new CLI command (\cohere-transcribe\) allows users to specify model directories, language, and compute units, while a benchmark tool is provided to evaluate performance on datasets like FLEURS.
Sources/FluidAudio/ASR/Cohere, Sources/FluidAudioCLI/Commands/ASR/Cohere · high confidence
Add Nemotron 3 8-speaker streaming diarizer
Introduces a new streaming diarization engine backed by NVIDIA's Nemotron 3 model, supporting up to 8 speakers with 10 ms resolution output. The implementation includes \Nemotron3Diarizer\ for processing audio via fixed 80 ms frames, \Nemotron3Models\ for CoreML inference with optimized IOSurface-backed output arrays to prevent memory exhaustion on long runs, and \Nemotron3StateUpdater\ for host-side speaker cache and FIFO management. The feature provides multiple latency-optimized presets (offline, low, verylow, ultra, fast, efficient) and supports both monolithic and split-graph CoreML model variants.
Sources/FluidAudio/Diarizer/Nemotron3 · high confidence
Add NeuTTS-2E emotional English TTS backend
Introduces a new NeuTTS-2E synthesis backend that generates emotional English speech at 24 kHz using a Qwen3 language model and NeuCodec vocoder. Users can now synthesize audio with specific speakers (emily, paul, sophie, steven) and emotions (angry, disgusted, fearful, happy, neutral, sad, surprised) via the NeuTtsManager API. This feature requires macOS 15 or iOS 18 and automatically downloads the necessary CoreML models and tokenizer assets on first use.
Sources/FluidAudio/TTS/NeuTts · high confidence
Add Paraformer-large (zh) CoreML transcription with timestamps
Introduces a new non-autoregressive Paraformer-large (zh) ASR backend using CoreML models (SANM encoder, CIF predictor, parallel decoder) running on the Neural Engine. Users can now transcribe Chinese audio via \ParaformerManager\ with support for fp16 or int8 precision, and access per-token timestamps through the new \transcribeWithTimestamps\ method, which uses an energy-based boundary refinement for accurate start/end times.
Sources/FluidAudio/ASR/Paraformer · high confidence
Add SenseVoiceSmall CoreML transcription with multilingual support
Introduces a new SenseVoiceSmall automatic speech recognition backend using Apple's CoreML framework. This non-autoregressive model supports multilingual transcription (including Chinese, English, Japanese, Korean, and Cantonese) and detects language, emotion, and audio events. The implementation includes a three-stage pipeline (preprocessor, encoder, and greedy CTC decoder) running on CPU and the Neural Engine, with configurable encoder precision (fp16, int8, or fp32) and automatic model downloading and caching.
Sources/FluidAudio/ASR/SenseVoice · high confidence
Add Sortformer streaming and offline diarization
Introduces the Sortformer diarizer, providing both streaming and offline modes for speaker identification. The streaming implementation (\SortformerDiarizer\) processes audio in chunks with a fixed 4-speaker slot architecture, maintaining internal state via FIFO and speaker caches to produce real-time diarization timelines. The offline implementation (\OfflineSortformerDiarizer\) processes fixed 30.72-second windows and uses a permutation-based stitcher (\SortformerSpeakerStitcher\) to align speaker identities across overlapping windows. Both modes support CoreML execution with automatic compute-unit selection to avoid ANE compile hangs on low-RAM devices, and allow loading models from local paths or HuggingFace.
Sources/FluidAudio/Diarizer/Sortformer · high confidence
Add StyleTTS2 CoreML backend for zero-shot TTS
Introduces the StyleTTS2 LibriTTS (iteration\_3) backend, enabling zero-shot text-to-speech synthesis using CoreML models. The implementation includes a manager that orchestrates model loading, a phonemizer using a Misaki lexicon cache with BART G2P fallback, and a synthesizer that processes text into 24 kHz mono audio. It supports reference audio for voice cloning, configurable style blending, and automatic text chunking to handle inputs longer than the model's default token limits.
Sources/FluidAudio/TTS/StyleTTS2 · high confidence
Add configurable streaming chunk sizes for Parakeet EOU model
The Parakeet End-of-Utterance (EOU) streaming manager now supports three distinct chunk sizes (160ms, 320ms, and 1280ms) to allow users to trade off latency for throughput. The 160ms mode remains the default, while the new 320ms and 1280ms modes offer higher processing throughput at the cost of increased latency between outputs, with specific CoreML encoder configurations and mel-spectrogram parameters tailored for each mode.
Sources/FluidAudio/ASR/Parakeet/Streaming/EOU · high confidence
Add multilingual grapheme-to-phoneme support via CoreML
The G2P module now includes a multilingual grapheme-to-phoneme converter powered by a CoreML-based CharsiuG2P ByT5 model. This addition enables phonemization for nine languages (American English, British English, Spanish, French, Hindi, Italian, Japanese, Brazilian Portuguese, and Mandarin Chinese) alongside the existing English-only G2P implementation. The system automatically detects the target language from Kokoro voice identifiers and handles byte-level tokenization for robust multilingual input.
Sources/FluidAudio/TTS/G2P · high confidence
Add streaming VAD and speech segmentation capabilities
VadManager now supports real-time audio processing and precise speech boundary detection. Users can process audio in chunks via \processStreamingChunk\, which emits \speechStart\ and \speechEnd\ events using a Silero-style state machine with configurable hysteresis thresholds. Additionally, a new \segmentSpeech\ API allows segmenting entire audio buffers or precomputed VAD results into \VadSegment\ objects, with options to extract the actual audio samples for each segment. These features are controlled by the new \VadSegmentationConfig\, which exposes parameters for minimum/maximum speech and silence durations, padding, and threshold offsets.
Sources/FluidAudio/VAD · high confidence
Added FastCluster wrapper for centroid linkage clustering
A new C++ wrapper around the fastcluster library has been added to expose centroid linkage hierarchical clustering via a C API, enabling Swift integration. This component provides the efficient clustering algorithm required by the pyannote community-1 speaker diarization pipeline to cluster speaker embeddings.
Sources/FastClusterWrapper · high confidence
Added NeMo Sortformer AMI benchmark script and documentation
Introduced a new Python benchmark tool (\nemo\_ami\_benchmark.py\) and its documentation (\README.md\) in the \Scripts/nemo\_ami\_benchmark\ directory. This tool runs NVIDIA's original NeMo Sortformer model on the AMI SDM dataset to establish a baseline for comparing against the Swift/CoreML implementation. The script supports streaming inference with configurable chunking parameters, handles RTTM ground truth files (including auto-download from pyannote), and outputs detailed diarization metrics such as DER, miss rate, and real-time factor.
_Scripts/nemo\_ami\benchmark · high confidence
Added Parakeet Unified benchmarking tool
A new benchmarking utility (\UnifiedBenchmark.swift\) has been added to evaluate the Parakeet Unified 0.6B ASR backend. This tool measures Word Error Rate (WER) and Real-Time Factor (RTFx) for both batch and streaming modes using the LibriSpeech dataset. It supports configurable encoder precision (defaulting to int8), streaming attention context windows (defaulting to 70,13,13 frames), and allows users to specify custom model directories or input files for targeted testing. Results are logged and can be written to a Markdown file for comparison.
Sources/FluidAudioCLI/Commands/ASR/Parakeet/Unified · high confidence
Added Spanish and French language support for Kokoro ANE TTS
This change introduces grapheme-to-phoneme (G2P) frontends and number normalization for Spanish and French variants of the Kokoro ANE text-to-speech engine. It adds \SpanishG2P\ and \FrenchG2P\ modules that convert text into IPA using rule-based phonology and lexicon lookups to match the specific training data conventions of the Kokoro voices. A shared \RomanceNumberNormalizer\ ensures that digits, decimals, and percentages are correctly verbalized in both languages (e.g., converting '21' to 'veintiún' or 'vingt et un'), while \KokoroAneLexicon\ handles the loading of language-specific pronunciation caches.
Sources/FluidAudio/TTS/KokoroAne/G2P, Sources/FluidAudio/TTS/KokoroAne/G2P/French, Sources/FluidAudio/TTS/KokoroAne/G2P/Spanish · high confidence
Beta CAM++ speaker-embedding backend via CoreML
FluidAudio now includes a beta CAM++ speaker-embedding backend that uses CoreML to extract 192-dimensional, L2-normalized embeddings from 16 kHz mono audio. The library exposes a \CampPlusEmbedder\ actor that automatically downloads the required CoreML models (a CPU-based preprocessor and a GPU/ANE-based model) from Hugging Face and provides methods to generate embeddings or compute cosine similarity for speaker verification. A corresponding CLI command, \fluidaudio campplus-embed\, is available on macOS to generate embeddings for a single file or compare two files to determine if they belong to the same speaker.
Sources/FluidAudio/Speaker, Sources/FluidAudioCLI/Commands/Speaker · high confidence
Beta FSMN-VAD backend using CoreML models
A new Voice Activity Detection (VAD) backend based on an FSMN model is available in beta. It uses CoreML models (preprocessor on CPU, scorer on the Neural Engine) to detect speech segments in audio files, returning start and end times in milliseconds. This backend is exposed via a new CLI command \fluidaudio fsmn-vad-segment \<audio-file\>\ for command-line usage.
Sources/FluidAudio/VAD/Fsmn, Sources/FluidAudioCLI/Commands/VAD · high confidence
Beta support for Inflect v2 CoreML TTS backend
Added a new beta text-to-speech backend using the Inflect v2 CoreML models (Micro and Nano variants). This change introduces the necessary infrastructure to download, cache, and load the fixed-shape CoreML encoder and synthesizer bundles from HuggingFace, along with the shared English G2P frontend assets. Users can now synthesize speech using the ultra-tiny VITS-family models, with options to control compute units, noise scale, and speed, while benefiting from deterministic, seedable audio generation.
Sources/FluidAudio/TTS/Inflect · high confidence
Beta support for NVIDIA Canary-1B-v2 transcription with custom vocabulary boosting
This change introduces a beta implementation of the NVIDIA Canary-1B-v2 automatic speech recognition engine, allowing users to transcribe audio using a CoreML-based pipeline that supports int4 (Neural Engine), fp16, and int8 (CPU) precision modes. The update includes a \CanaryKeywordBooster\ that integrates a CTC-based keyword spotter to detect and inject custom vocabulary terms into the transcript via fuzzy matching or timestamp-guided insertion, improving accuracy for domain-specific terms. The \CanaryManager\ handles audio chunking for inputs longer than the 15-second fixed window, stitching results using a longest-common-substring merge strategy.
Sources/FluidAudio/ASR/Canary · high confidence
FluidAudio CLI initial release with comprehensive audio processing commands
The FluidAudio CLI is now available as a standalone tool for macOS, providing commands for audio transcription, diarization, voice activity detection, and text-to-speech. Users can run benchmarks for ASR, TTS, and diarization, perform streaming transcription with models like Nemotron and Parakeet, and utilize the new LuxTTS G2P dump utility for phoneme validation. The CLI also includes utilities for speaker embedding, dataset downloading, and multi-stream processing.
Sources/FluidAudioCLI · high confidence
Introduce LS-EEND diarization pipeline and DER evaluation metrics
Added the LS-EEND neural diarization implementation, including the \Diarizer\ protocol for streaming and offline processing, \DiarizerTimeline\ for managing speaker segments and post-processing configuration, and \DiarizationDER\ for computing frame-wise Diarization Error Rate with Hungarian speaker mapping and collar support. This introduces a new end-to-end diarization capability alongside standard evaluation metrics.
Sources/FluidAudio/Diarizer · high confidence
Introduce LS-EEND streaming diarizer for long-form audio
Added a new LS-EEND (Long-form Streaming End-to-End Neural Diarization) implementation for FluidAudio. This feature provides a CoreML-based diarizer that supports both streaming (chunked) and offline processing of long audio files. It includes a preprocessor for audio resampling and mel-spectrogram extraction, a model wrapper for loading LS-EEND models from HuggingFace, and a diarizer class that manages the inference loop, speaker enrollment, and timeline updates.
Sources/FluidAudio/Diarizer/LS-EEND · high confidence
Introduce LuxTTS zero-shot voice cloning backend
Adds a new LuxTTS (ZipVoice-Distill) backend for 48 kHz zero-shot voice cloning on iOS and macOS using CoreML. This feature includes a native Swift English text normalizer and a G2P engine that reproduces espeak-IPA phonemization from a bundled lexicon, a mel spectrogram extractor, and a flow-matching synthesizer that handles long-text synthesis via continuation-prompted spans with spurious-pause detection and re-drawing. The backend downloads CoreML models from HuggingFace and supports GPU (macOS) and Neural Engine (iOS) compute units.
Sources/FluidAudio/TTS/LuxTts · high confidence
Introduce Parakeet ASR configuration, types, and model loading infrastructure
This change adds the foundational types and configuration structures for the Parakeet ASR engine. It introduces \ASRConfig\ to manage transcription settings, including options for parallel chunk concurrency, streaming thresholds, and specific behaviors for long-form transcription like seam-gap repair and dual-decode arbitration. It also defines the result structures (\ASRResult\, \TokenTiming\, \WordTiming\) that expose transcript text, confidence, and granular timing data to users. Additionally, it provides the \ParakeetLanguageModels\ container and loading logic to handle downloading and initializing CoreML models from the ModelHub, supporting both int8 and fp32 encoder variants.
Sources/FluidAudio/ASR/Parakeet · high confidence
Introduce Parakeet Unified 0.6B ASR backend with streaming and offline batch modes
Adds a new Parakeet Unified 0.6B (FastConformer-RNNT) speech-to-text backend supporting both real-time streaming and offline batch transcription. The streaming manager (\StreamingUnifiedAsrManager\) uses a chunked-attention encoder with configurable latency tiers (e.g., 320ms to 2.08s) and exposes per-token timings for downstream word-level attribution. The offline batch manager (\UnifiedAsrManager\) processes long audio using overlapping 15-second windows merged via time-tolerant token matching, achieving lower word error rates than the streaming path. Both paths use a native Swift log-mel extractor (\UnifiedMelExtractor\) instead of a CoreML preprocessor, support int8 and fp16 encoder precision, and include vocabulary boosting capabilities to rescore transcripts against custom keyword lists. Model loading now uses \ModelHub\ with recovery logic to prevent bricking on interrupted downloads.
Sources/FluidAudio/ASR/Parakeet/Unified · high confidence
Introduce Supertonic-3 CoreML TTS pipeline with multilingual support
Added the Supertonic-3 text-to-speech pipeline to the FluidAudio library, enabling high-quality, multilingual voice synthesis using CoreML models. This new feature includes a text chunker that handles paragraph and sentence splitting with abbreviation awareness, a Unicode processor for text normalization and language tagging, and a synthesizer that orchestrates CoreML inference stages (text encoding, duration prediction, latent denoising, and vocoding). The implementation supports multiple languages, configurable voice styles, and speed adjustments, with specific optimizations for CJK scripts and efficient memory handling via MLMultiArray utilities.
Sources/FluidAudio/TTS/Supertonic3/Pipeline · high confidence
Introduce in-memory speaker clustering and management for diarization
The diarization module now includes a new \SpeakerManager\ struct and supporting types (\Speaker\, \SpeakerUtilities\) to handle in-memory speaker tracking and clustering. This change adds the capability to initialize known speakers, assign new speakers based on embedding similarity thresholds, and maintain consistent speaker identities across audio chunks. Users benefit from improved speaker consistency and configurable thresholds for assignment and embedding updates, replacing previous ad-hoc or external management approaches with a dedicated, thread-safe clustering component.
Sources/FluidAudio/Diarizer/Clustering · high confidence
Introduce native Swift StyleTTS2 synthesizer with CoreML backend
Adds a complete, native Swift implementation of the StyleTTS2 LibriTTS (iteration\_3) text-to-speech pipeline, replacing the previous Python-dependent inference path. This new synthesizer runs entirely on-device using CoreML models, featuring a custom mel-spectrogram extractor, a fused diffusion sampler with Karras noise scheduling, and CPU-side glue operations for duration prediction, alignment, and style blending. The pipeline includes a dedicated English phonemizer that leverages a local lexicon cache and a BART G2P CoreML model for out-of-word handling, along with a native text cleaner and tokenizer to ensure consistent tokenization. This change enables asynchronous, high-performance speech synthesis with configurable alpha/beta style blending and noise seeds, significantly reducing latency and dependency overhead for English TTS.
Sources/FluidAudio/TTS/StyleTTS2/Pipeline · high confidence
Introduce offline speaker embedding extraction and PLDA transformation
The offline diarization pipeline now includes dedicated components for extracting speaker embeddings and transforming them for clustering. A new \OfflineEmbeddingExtractor\ handles the inference of 256-dimensional embeddings from audio chunks using Core ML models, while \PLDATransform\ converts these embeddings into 128-dimensional PLDA-space features using a Core ML model and computes cosine similarity scores. Additionally, \WeightInterpolation\ provides utilities to resample segmentation masks, ensuring the offline pipeline's weight interpolation matches the reference Pyannote implementation.
Sources/FluidAudio/Diarizer/Offline/Extraction · high confidence
Introduces offline audio segmentation processor for diarization
Adds the \OfflineSegmentationProcessor\ struct, which implements the core logic for splitting audio into segments using a sliding window approach and a Core ML model. This component handles audio buffering, feature extraction, and model inference to produce segmentation outputs, serving as the foundational step for the offline diarization pipeline.
Sources/FluidAudio/Diarizer/Offline/Segmentation · high confidence
Introduces shared audio processing and infrastructure components
Adds a suite of foundational utilities to the FluidAudio shared module, including ANE-optimized memory management (ANEMemoryOptimizer, ANEMemoryUtils) for efficient speaker diarization, a unified AppLogger with configurable console mirroring, and a generic AssetDownloader. It also introduces core audio handling classes such as AudioConverter for format resampling, AudioMelSpectrogram for native Swift spectrogram computation, and AudioStream for real-time sliding-window buffering, alongside supporting types like AudioSampleSource and AudioSourceFactory.
Sources/FluidAudio/Shared · high confidence
Inverse Text Normalization (ITN) post-processor with native NeMo engine and context awareness
ASR output is now post-processed to convert spoken-form text into written form (e.g., "two hundred" to "232", "five dollars" to "$5.50") using a new \TextNormalizer\ class. This feature links directly to the bundled NeMo \text-processing-rs\ engine for high-performance normalization, replacing previous runtime discovery methods. It includes sentence-mode normalization with sliding window support and uses Apple's NaturalLanguage framework to prevent false positives on ambiguous words (like "period") by checking their grammatical context. The native engine is available by default but can be opted out via the \NemoTextProcessing\ trait for ASR-only consumers, in which case text passes through unchanged.
Sources/FluidAudio/ITN · high confidence
KokoroAne ANE TTS engine adds multi-language variants and detailed synthesis API
The KokoroAne module now supports Mandarin, Japanese, Spanish, and French variants in addition to English, each with specific default voices and optimized text frontends (including MeCab for Japanese and custom lexicons for Mandarin). The high-level synthesis API now auto-chunks long text to handle the model's frame limits, and the detailed synthesis method exposes normalized text and phonemes for debugging or custom processing. The engine also includes robust error handling for non-finite model outputs and OS-specific BNNS crash advisories.
Sources/FluidAudio/TTS/KokoroAne · high confidence
LocalVQE speech enhancement with streaming and demo app
Adds a Core ML port of the LocalVQE model for joint acoustic echo cancellation, noise suppression, and dereverberation on 16 kHz speech. The library introduces \LocalVqeManager\ for whole-clip processing and \LocalVqeStream\ for safe, stateful streaming with configurable variants (v1.3, v1.2) and chunk sizes (16 ms for live, 256 ms for batch). A new macOS SwiftUI demo (\LocalVQEDemo\) exercises these APIs in two modes: offline file enhancement with A/B listening and Parakeet TDT v3 transcription, and live capture that pairs a far-end playback with mic input via time-aligned streaming, complete with level meters and WAV export.
Examples, Sources/FluidAudio/Enhancement · high confidence
Nemotron streaming ASR: native Swift mel front-end, multilingual support, and vocabulary biasing
The Nemotron streaming ASR manager now uses a native Swift log-mel extractor instead of a CoreML preprocessor, fixing iPadOS cold-start failures that previously produced empty transcripts. It adds support for the Nemotron 0.6B multilingual model, including language detection, prompt-based language hints, and a tokenizer that strips language tags from output. A new decode-time vocabulary biasing engine boosts custom hotwords during greedy decoding without requiring a CTC head. The system also introduces configurable chunk sizes (2240ms, 1120ms, 560ms) to trade off latency and throughput, and includes a blank-span rescue mechanism to recover words that the streaming decoder silently dropped.
Sources/FluidAudio/ASR/Parakeet/Streaming/Nemotron · high confidence
New AMI, Japanese, and Multilingual benchmark dataset support
The CLI now includes dedicated parsers and downloaders for several new benchmark datasets. The AMI corpus is fully supported with robust annotation parsing that fails loudly if ground truth is missing (preventing silent scoring against placeholder data), and downloads now prefer a HuggingFace mirror with the Edinburgh server as a fallback. Japanese ASR benchmarking is added via JSUT-basic5000 and Common Voice (Japanese) datasets, with metadata and audio downloaded from HuggingFace. Additionally, multilingual ASR benchmarking is extended to include FLEURS, LibriSpeech, and Earnings22, with a unified configuration enum handling cache paths, HuggingFace repositories, and language mappings for these datasets.
Sources/FluidAudioCLI/DatasetParsers · high confidence
New ASR benchmarking and multi-stream transcription tools
The CLI now includes a suite of benchmarking utilities for evaluating ASR performance on standard datasets (LibriSpeech, FLEURS, Earnings22) and analyzing decoding characteristics (CTC decode strategies, emission delays). These tools provide detailed metrics including Word Error Rate (WER), Character Error Rate (CER), Real-Time Factor (RTFx), and streaming latency. Additionally, a new multi-stream command allows users to process multiple audio sources (e.g., microphone and system audio) in parallel using a shared ASR model session, enabling comparative testing or simultaneous transcription scenarios.
Sources/FluidAudioCLI/Commands/ASR/Parakeet/SlidingWindow · high confidence
New ASR commands for Canary, SenseVoice, Paraformer, and Japanese TDT models
Users can now transcribe audio and run benchmarks using four new speech-to-text engines: Canary-1B-v2 (multilingual, with optional custom-vocabulary boosting), SenseVoiceSmall (multilingual, non-autoregressive), Paraformer-large (Chinese, with timestamps), and Japanese TDT models (via JSUT and Common Voice benchmarks). These are available as new CLI commands (canary-transcribe, canary-earnings-benchmark, sensevoice-transcribe, sensevoice-benchmark, paraformer-transcribe, and japanese-asr-benchmark) that load the respective CoreML models, perform transcription or evaluation, and report metrics like WER/CER and real-time factors.
Sources/FluidAudioCLI/Commands/ASR · high confidence
New CLI data models and serialization support for diarization results
The FluidAudioCLI now includes structured data models for processing and benchmarking results, enabling users to inspect detailed output such as speaker segments, timing metrics, and performance assessments. This change introduces \ProcessingResult\, \BenchmarkResult\, and \BenchmarkSummary\ structs for structured data handling, along with \Codable\ extensions for \DiarizerConfig\ and \TimedSpeakerSegment\ to allow these internal types to be serialized for CLI output. Additionally, a \PerformanceAssessment\ enum is added to evaluate diarization quality based on DER thresholds, providing users with clear pass/fail/needs-work status indicators and corresponding exit codes for scripting integration.
Sources/FluidAudioCLI/Models · high confidence
New CLI utilities for ASR evaluation and diarization metrics
Added a suite of utility modules in Sources/FluidAudioCLI/Utils to support advanced speech evaluation and user-facing reporting. This includes DiarizationMetricsCalculator for computing diarization error rate (DER) and Jaccard error rate (JER) with deterministic speaker mapping, and WERCalculator for Word Error Rate and Character Error Rate (CER) scoring. The WERCalculator supports multiple normalization strategies: a HuggingFace-compatible normalizer for English, a conservative 'basic' normalizer for multilingual ASR, and a CJK-aware metric that uses character-level edit distance for languages without word spaces. TextNormalizer implements these rules, including British-to-American spelling conversion via a new english.json dictionary. Additional utilities include RTTMParser for loading ground-truth speaker segments, InlineDiff for generating ANSI-highlighted word-level diffs between reference and hypothesis transcripts, ResultsFormatter for printing benchmark tables and timing breakdowns, and TerminalUI for progress bars and box-drawing characters.
Sources/FluidAudioCLI/Utils · high confidence
New Chatterbox TTS backends: Multilingual and Nano
This change introduces two new local text-to-speech backends for Chatterbox. The Multilingual backend uses a T3 (Llama-520M) + S3Gen + HiFT pipeline to synthesize speech in 19 supported languages (e.g., en, de, es, fr, ja, zh) at 24 kHz, requiring macOS 15 / iOS 18. The Nano backend uses a smaller T3 (GPT2-small, 110M) + S3Gen + HiFT pipeline for English-only synthesis with paralinguistic tag support (e.g., \[laugh\], \[chuckle\]). The Nano backend offers two output capacities: standard (\~9.9s audio) and extended (\~29.9s audio, requiring an additional \~280 MB download). Both backends download CoreML models and assets from HuggingFace on first use, validate asset dimensions to prevent crashes, and enforce strict token budgets for text length and generation duration.
Sources/FluidAudio/TTS/Chatterbox · high confidence
New English text frontend for KokoroAne TTS
Introduces the \KokoroAneEnglishPhonemizer\ to handle English text normalization for the KokoroAne speech synthesis chain. This component implements a 10-stage word resolution pipeline that prioritizes a caller-supplied custom lexicon, followed by specific overrides for uppercase initialisms (e.g., spelling out 'AI' or 'FBI' as letter names rather than blending them), and then falls back to a case-sensitive and case-insensitive Misaki lexicon. It also handles smart-apostrophe contractions, hyphenated compound splits, and possessive stems (e.g., 'C-section's') to ensure accurate pronunciation, finally resorting to a BART G2P CoreML fallback for out-of-vocabulary words. Punctuation is preserved and attached to preceding words to match Kokoro's prosody requirements.
Sources/FluidAudio/TTS/KokoroAne/G2P/English · high confidence
New Mandarin (v1.1-zh) G2P pipeline with context-aware pronunciation
The Kokoro-ANE Mandarin text-to-speech pipeline now uses a comprehensive G2P system that significantly improves pronunciation accuracy for Chinese text. This update introduces context-aware polyphone disambiguation using a g2pW BERT model, which resolves correct readings for ambiguous characters based on sentence context. It also adds support for erhua (r-coloring) merging, tone sandhi rules (including 3+3 promotion and 不/一 changes), and proper segmentation of proper nouns using a jieba HMM decoder. Users can now verbalize numbers, dates, times, percentages, fractions, and currency expressions correctly, and can provide custom pronunciation overrides via a user-supplied lexicon to handle specific names or terms.
Sources/FluidAudio/TTS/KokoroAne/G2P/Mandarin · high confidence
New Nemotron and Parakeet EOU streaming commands and benchmarks
The CLI now includes dedicated commands for the Nemotron Speech Streaming 0.6B model and the Parakeet EOU streaming model. Users can transcribe audio files using \nemotron-transcribe\ and \nemotron-multilingual-transcribe\, with support for custom chunk sizes, language selection, and decode-time custom vocabulary biasing. Additionally, new benchmarking tools (\nemotron-benchmark\, \nemotron-multilingual-fleurs-benchmark\, \nemotron-multilingual-multi-stream-bench\, and \nemotron-vocab-benchmark\) allow evaluation of Word Error Rate (WER), throughput (RTFx), and vocabulary biasing effects on datasets like LibriSpeech and FLEURS. The \parakeet-eou\ command provides streaming transcription with configurable end-of-utterance (EOU) debounce and compute unit selection.
Sources/FluidAudioCLI/Commands/ASR/Parakeet/Streaming · high confidence
New SSML processing engine for FluidAudioTTS
The FluidAudioTTS module now includes a new SSML processing pipeline (SSMLProcessor, SSMLTagParser, SayAsInterpreter) that parses and handles standard SSML tags including \<phoneme\>, \<sub\>, and \<say-as\>. This enables users to use phonetic overrides, text substitutions, and semantic interpretations for numbers, dates, times, and other content types directly within their TTS inputs.
Sources/FluidAudio/TTS/SSML · high confidence
New benchmark and download commands for diarization, enhancement, and G2P
The CLI now includes dedicated commands to evaluate and download datasets for speaker diarization, acoustic enhancement, and grapheme-to-phoneme conversion. The \diarization-benchmark\ command evaluates streaming and offline diarization pipelines against datasets like AMI and VoxConverse, reporting DER, RTFx, and latency metrics. The \enhance-benchmark\ command measures near-end word recall and far-end leakage for the LocalVQE enhancement model using the AEC-Challenge synthetic set. The \g2p-benchmark\ command evaluates the multilingual CharsiuG2P model's phoneme error rate across supported languages. A new \download\ command allows users to fetch the required benchmark datasets (e.g., AMI, VAD, LibriSpeech, JSUT) and models on demand.
Sources/FluidAudioCLI/Commands · high confidence
New benchmarking, profiling, and dataset materialization scripts
The Scripts directory now includes a suite of new tools to support evaluation and development workflows. A new Swift utility, ane\_profile.swift, reports the CoreML compute-unit plan (ANE, GPU, CPU) for .mlmodelc bundles. Several new benchmark scripts have been added: run\_benchmarks.py provides a unified Python entry point for ASR, VAD, and Diarization benchmarks with baseline comparisons; shell scripts (parakeet\_subset\_benchmark.sh, fleurs\_parakeet\_sub\_benchmark.sh, diarizer\_subset\_benchmark.sh) automate regression testing for Parakeet TDT, FLEURS multilingual, and diarization models respectively. Additionally, Python scripts (materialize\_earnings22\_longform.py, materialize\_alimeeting\_card\_audio.py, materialize\_notsofar\_card\_audio.py) and a dataset manifest (localvqe-dataset.json) help prepare specific evaluation datasets for long-form ASR and NVIDIA Nemotron 3 card-style benchmarks. Finally, verify\_localvqe\_benchmark.py and its test suite ensure the integrity of local VQE benchmark reports.
Scripts · high confidence
New offline clustering algorithms and speaker count constraints
The offline diarizer now includes new clustering implementations: Agglomerative Hierarchical Clustering (AHC) with Fastcluster integration, K-Means with deterministic multi-initialization (n\_init) for robustness, and Variational Bayes (VBx) clustering. It also introduces speaker count constraints (minSpeakers, maxSpeakers, numSpeakers) to limit the detected number of speakers, and a constrained cluster assignment strategy to ensure distinct local speakers within segmentation chunks map to distinct clusters, improving parity with pyannote.
Sources/FluidAudio/Diarizer/Offline/Clustering · high confidence
New shared TTS infrastructure and improved English text normalization
This change introduces a suite of shared utilities for the FluidAudio TTS system. It adds an AudioPostProcessor to improve output quality by removing low-frequency rumble, applying de-essing to reduce harsh sibilant sounds, and smoothing high frequencies. English text normalization is significantly enhanced with a new EnglishTextNormalizer that handles raw numbers, ordinals, decimals, 12-hour times, decades, bare years, and roman-numeral list markers, and it can optionally use a byte-exact NeMo text normalization engine for richer coverage. Additional shared components include EnglishInitialisms for correctly spelling uppercase acronyms, a PhonemeChunker for splitting long phoneme strings to fit model limits, a LexiconAssetCache for loading phoneme data, and a TtsCacheDirectory that moves model storage to Application Support on iOS to prevent system purging. A TtsComputeUnitPreset provides a unified way to configure hardware acceleration (ANE, GPU, CPU) across different TTS backends.
Sources/FluidAudio/TTS/Shared · high confidence
On-device CUA-S1-FORMS decision scoring via Core ML
Added a new Core ML-based scoring component for the CUA-S1-FORMS model within the FluidAudio decision engine. This change introduces the \CuaS1FormsManager\ actor, which loads the model from a local file or cache and scores 2–32 UI action options against a text context on-device. The implementation includes input encoding (\CuaS1FormsInput\) that truncates text to specific byte limits (224 bytes for context, 96 bytes per option), model validation logic, and a stable softmax calculation for probability outputs, returning a \CuaS1FormsResult\ with the selected option index and associated confidence scores.
Sources/FluidAudio/Decision · high confidence
Punctuation-aware streaming ASR with optimized argmax
Streaming ASR now provides sentence-aware text segmentation via a new PunctuationCommitLayer that separates finalized text from speculative 'ghost' text, committing segments at punctuation marks (., !, ?) or after a configurable debounce timeout. This layer is supported by a new LogitsArgmax utility that accelerates the per-frame argmax operation for SenseVoice CTC and Paraformer decoders by using Accelerate's vDSP\_maxvi instead of slow MLMultiArray subscript reads, including efficient handling of fp16 logits on platforms where Float16 is unavailable.
Sources/FluidAudio/ASR/Shared · high confidence
StyleTTS2 CoreML model download and caching infrastructure
The StyleTTS2 text-to-speech backend now manages its own CoreML model assets, downloading the LibriTTS (iteration\_3) models from HuggingFace and storing them in the shared Application Support cache. This change introduces a model store that loads eight default-stage models and lazily fetches bucket variants (T=64, 128, 256) for the BERT and diffusion sampler stages as needed. It also ensures shared Kokoro G2P assets are available, allowing StyleTTS2 to function without requiring a separate Kokoro installation.
Sources/FluidAudio/TTS/StyleTTS2/Assets · high confidence
Supertonic-3 multilingual TTS with on-device CoreML support
Introduces the Supertonic-3 text-to-speech pipeline, enabling on-device synthesis in 31 languages using CoreML-converted models. The update adds a new \Supertonic3Manager\ API for initialization and synthesis, along with built-in voice presets (F1–F5, M1–M5) and configurable inference options such as ANE-optimized quantized VectorEstimators, speech speed, and inter-chunk silence duration.
Sources/FluidAudio/TTS/Supertonic3 · high confidence
Architecture
FluidAudio module restructure and Swift 6 concurrency support
The FluidAudio library has been restructured into a cohesive Swift module with a new entry point (FluidAudioSwift.swift) that provides backward-compatible type aliases (SpeakerDiarizationConfig, SpeakerDiarizationError) and a namespace struct. The codebase now includes a comprehensive ModelNames enum defining all supported model repositories (ASR, TTS, VAD, Diarization) and a ModelRegistry system allowing programmatic override of the base URL and repository paths for mirrored deployments. Additionally, a new MachTaskSelfWrapper module was added to safely expose the mach\_task\self\ C macro to Swift 6, enabling strict concurrency compliance.
Sources/FluidAudio · high confidence
Behavioural changes
Diarizer Core: ModelHub integration, progress callbacks, and per-chunk embeddings
The Diarizer Core module has been restructured to replace the legacy DownloadUtils with ModelHub for model loading, which now supports optional progress callbacks during download. The main DiarizerManager API has been updated to accept an optional progressHandler in performCompleteDiarization, allowing users to track processing progress. Additionally, the DiarizationResult now optionally exposes per-chunk speaker embeddings (ChunkEmbedding) when enabled, providing finer-grained data for downstream consumers like cluster-level contamination correction.
Sources/FluidAudio/Diarizer/Core · high confidence
Download stack refactored with improved reliability and progress reporting
The model download system has been restructured into a new stack (ModelHub, FileDownloader, HFTreeLister, ModelCache) that replaces the previous DownloadUtils implementation. This change introduces a stall watchdog to detect and retry frozen transfers quickly, honors server Retry-After headers for rate limiting, and preserves the local model cache on transient network errors to allow resumable downloads. It also adds concurrent subdirectory fetching for faster downloads of multi-file models, unified progress reporting with phase tracking (listing, downloading, compiling), and stricter validation of downloaded artifacts to prevent caching invalid or HTML error responses.
Sources/FluidAudio/Shared/Download · high confidence
Improved custom vocabulary boosting with fuzzy matching and reliable tokenization
The Parakeet ASR engine now features a more robust custom vocabulary system that ensures terms are correctly tokenized for CTC rescoring, preventing silent failures when terms are added programmatically. It introduces a BK-tree for efficient fuzzy string matching, allowing the system to detect and correct misspelled or phonetically similar vocabulary terms (e.g., correcting 'nvida' to 'nvidia') using configurable similarity thresholds. The system also supports compound word detection and context-aware biasing, enabling it to boost specific phrases even when they appear as part of longer spoken sentences, while applying safety guards to prevent false positives on short or common words.
Sources/FluidAudio/ASR/Parakeet/SlidingWindow/CustomVocabulary · high confidence
Improved custom vocabulary rescoring with adaptive thresholds and opt-in controls
The VocabularyRescorer now uses CTC log-probabilities to verify that vocabulary terms actually match the audio, replacing heuristic-based replacements with a fair CTC-vs-CTC comparison. To reduce short-keyword over-fires, the system introduces an opt-in adaptive context-biasing weight that tapers boosts for short terms and scales them up for longer ones, along with optional minimum similarity floors for spotter-anchored rescues. Users can also opt-out of the acoustic spotter-rescue pass entirely to revert to pre-0.14.5 behavior, and the text alignment logic for UTF-8 ranges has been made iterative to prevent stack overflow issues.
Sources/FluidAudio/ASR/Parakeet/SlidingWindow/CustomVocabulary/Rescorer · high confidence
Improved custom vocabulary word-spotting accuracy with standalone CTC scoring
The custom vocabulary word-spotting engine now uses a standalone CTC model (Parakeet CTC 110M/0.6B) for scoring, replacing the previous approach. This introduces a new BPE tokenizer that matches the NeMo CTC pipeline, a corrected dynamic-programming algorithm that properly handles blank tokens and repeated words, and chunked audio processing for longer inputs. As a result, keyword detection scores are more probabilistically meaningful and accurate, though raw score magnitudes have changed (they are now more negative per token due to blank emission costs), so existing score thresholds may need adjustment.
Sources/FluidAudio/ASR/Parakeet/SlidingWindow/CustomVocabulary/WordSpotting · high confidence
In-process Japanese text frontend for Kokoro ANE
The Japanese text-to-speech pipeline now uses an in-process G2P (grapheme-to-phoneme) frontend instead of external dependencies. This change introduces a MeCab-compatible Viterbi tokenizer backed by a trimmed unidic-lite dictionary and implements Misaki's Cutlet rules for converting tokens to IPA, ensuring the Japanese voice output matches the training data pipeline. The system handles text normalization, dictionary word regrouping, and phoneme generation entirely within the application process.
Sources/FluidAudio/TTS/KokoroAne/G2P/Japanese · high confidence
Introduce sliding-window ASR with vocabulary boosting and temporal token deduplication
The SlidingWindowAsrManager now supports real-time streaming transcription using overlapping windows, featuring vocabulary boosting that rescoring every window against custom terms, and temporal token deduplication that prevents prefix-collision loss by gating deduplication to temporally adjacent tokens. The manager also exposes volatile and confirmed transcript states, handles window seam corrections to avoid word loss at boundaries, and includes a session manager for shared model loading across multiple audio sources.
Sources/FluidAudio/ASR/Parakeet/SlidingWindow · high confidence
Introduces audio validation and optimized segmentation processing
The diarization segmentation module now includes an AudioValidation component that checks for minimum duration, silence, and embedding integrity before processing. The core SegmentationProcessor has been rewritten to use ANE-aligned memory and zero-copy operations via vDSP, significantly improving inference performance for speaker segment detection. Additionally, a new SlidingWindow utility has been added to manage time-based segment calculations.
Sources/FluidAudio/Diarizer/Segmentation · high confidence
KokoroAne pipeline stability and compute routing fixes
The KokoroAne synthesis pipeline now handles OS-specific hardware routing and data integrity issues to prevent crashes. On macOS/iOS 27, the noise and tail stages are routed to the CPU to avoid intermittent MPSGraph aborts, while on earlier OS versions (and M5 Macs), these stages remain on the GPU to avoid libBNNS segfaults and ensure performance. The synthesizer now throws an error instead of crashing when the PostAlbert model produces non-finite (NaN/inf) duration outputs, a known issue on iOS 27 betas. Additionally, memory allocation for MLMultiArrays includes slack bytes to prevent segfaults caused by BNNS overreads on OS 27, and the prosody stage uses the fp32 KokoroProsody\_v2 model to fix onset artifacts on long utterances.
Sources/FluidAudio/TTS/KokoroAne/Pipeline · high confidence
Offline diarizer pipeline restructured with caching, speaker constraints, and zero-vote fixes
The offline diarization engine has been restructured to split processing into a cacheable \prepare()\ step (segmentation and embeddings) and a \cluster()\ step, allowing clustering to be re-run without repeating model inference. New configuration options allow users to constrain the number of detected speakers (min, max, or exact count) and opt into an embedding skip strategy to improve performance on overlapping windows. The pipeline now includes a post-processing pass that re-embeds audio spans where no speaker votes were detected, preventing silent misassignments. Additionally, the system warns on macOS 14 about a known CoreML crash and honors the \computeUnits\ configuration for model loading.
Sources/FluidAudio/Diarizer/Offline/Core · high confidence
PocketTTS asset loading and download logic refactored for v2 voice packs
The PocketTTS module now uses a new constants loader and resource downloader to support v2 voice packs with pre-baked KV cache snapshots, enabling faster synthesis by skipping voice prefill for shipped voices. The download process is optimized to skip unused model variants (such as different precision or placement types) and irrelevant repository artifacts, reducing download size and time. Additionally, the loader handles best-effort retrieval of optional assets like \bos\_before\_voice.bin\ and encoder recovery weights, ensuring that synthesis with shipped voices remains functional even if these optional files are missing or fail to download.
Sources/FluidAudio/TTS/PocketTTS/Assets · high confidence
PocketTTS pipeline refactored for v2.1 with MLState support and streaming sessions
The PocketTTS pipeline has been updated to support the v2.1 model architecture, which replaces the per-token conditioning step with a one-shot conditioner and fuses the flow decoder into a single CoreML dispatch for improved performance. The system now introduces persistent TTS sessions that maintain voice and audio state across multiple utterances to ensure seamless audio continuity. Additionally, a new \.aneState\ placement option is available for macOS 15+ and iOS 18+, utilizing the CoreML MLState API to manage KV caches more efficiently. The pipeline also includes robust schema discovery for CoreML models to handle varying output names across different language packs and model variants.
Sources/FluidAudio/TTS/PocketTTS/Pipeline · high confidence
Repository rebranded to FluidAudio with Swift 6.2+ opt-out for NeMo text processing
The project has been renamed from SeamlessAudioSwift to FluidAudio, including the Swift package name and all documentation. For Swift 6.2+ toolchains, the bundled NeMo text-normalization engine (\~8 MB) is now an opt-in trait, allowing ASR-only applications to exclude it by specifying \traits: \[\]\ in their package dependencies.
(repo-wide) · high confidence
Supertonic-3 CoreML model management and on-demand voice downloads
The Supertonic-3 TTS backend now manages its CoreML assets (text encoder, duration predictor, vector estimator, and vocoder) and companion config files through a new \Supertonic3ModelStore\ actor, which handles lazy loading and supports both dynamic and ANE-bucketed vector estimator modes. Model assets are downloaded from HuggingFace via \ModelHub\ and stored in the shared Application Support directory, ensuring consistency across TTS backends. Additionally, built-in voice styles can now be downloaded on-demand from the \voice\_styles\ subdirectory when first used, rather than requiring all assets to be pre-loaded.
Sources/FluidAudio/TTS/Supertonic3/Assets · high confidence
Unified streaming ASR engine with multiple latency tiers and fused decoding
The streaming ASR subsystem now supports a unified catalog of true streaming models (Parakeet EOU, Nemotron, and Parakeet Unified) accessible via a single \StreamingModelVariant\ enum, allowing users to select specific latency/throughput tiers (e.g., 160ms to 2240ms chunks) and engine families. This change introduces a common \StreamingAsrManager\ protocol and shared utilities for cache management, audio buffering, and tokenization, while adding an opt-in fused decoder+joint model path for the Parakeet EOU engine to reduce CoreML dispatch overhead and improve real-time factor without changing word error rate.
Sources/FluidAudio/ASR/Parakeet/Streaming · high confidence
Fixes
Fix heap-buffer-overflow in EmbeddingExtractor
The EmbeddingExtractor now clamps the calculated number of masks per chunk to the actual number of available masks. This prevents a heap-buffer-overflow that could occur when processing audio segments where the formula for determining mask count exceeds the bounds of the mask array, ensuring stable memory access during speaker embedding extraction.
Sources/FluidAudio/Diarizer/Extraction · high confidence
Fixes Kokoro ANE model cache compatibility on iOS 27
The KokoroAne module now automatically detects and repairs legacy CoreML model bundles that lack the required \FlexibleShapeInformation\ metadata. On iOS 27, these older artifacts caused invalid synthesis results; the new migration logic backs up incompatible bundles, downloads updated versions from Hugging Face, and cleans up the backups, ensuring reliable TTS output without user intervention.
Sources/FluidAudio/TTS/KokoroAne/Assets · high confidence
Fixes for Parakeet TDT v3 streaming and long-form transcription reliability
This update introduces a new sliding-window TDT implementation for Parakeet that resolves several critical transcription artifacts. It adds a recovery ladder to handle windows that decode to blank despite containing speech, ensuring no content is lost. Streaming and long-form processing now use temporal deduplication to prevent prefix-collision loss and seam artifacts, while the final window is end-aligned to the last speech-bearing frame to stop trailing-word loss. Additionally, sentence-final punctuation is correctly resolved from the vocabulary, and the system now supports Parakeet v3 variants including Redux and Ultra.
Sources/FluidAudio/ASR/Parakeet/SlidingWindow/TDT · high confidence
Introduce Parakeet TDT v3 decoder with streaming context continuity and language-specific fixes
The sliding-window decoder now implements the Token-and-Duration Transducer (TDT) v3 algorithm, enabling faster decoding via duration prediction and supporting streaming transcription with proper context continuity across chunk boundaries. This update resolves trailing-word loss in the final window, fixes Cyrillic emission errors on short Latin utterances, and reduces English drift on French recordings via a scoped token blocklist. It also introduces an opt-in no-mel decode path for long-form audio and ensures sentence-final punctuation IDs are correctly resolved from the loaded vocabulary.
Sources/FluidAudio/ASR/Parakeet/SlidingWindow/TDT/Decoder · high confidence
Test coverage
Added CI test suite for FluidAudio diarizer and model components; Added LuxTTS test suite and fixtures; Added comprehensive test coverage for Diarizer components; Added comprehensive test coverage for KokoroAne TTS variants and internals; Added comprehensive test suite for Parakeet TDT sliding-window ASR; Added deterministic VAD test helper; Added regression and unit tests for ASR token deduplication, punctuation commit layer, and logits argmax; Added regression tests for EmbeddingExtractor buffer overflow; Added test coverage for CLI scoring determinism, AMI data handling, and LocalVQE enhancement; Added test coverage for FluidAudio core components; Added test coverage for Parakeet sliding-window custom vocabulary components; Added test coverage for VAD chunking, segmentation, and streaming logic; Added tests for CUA-S1-FORMS Core ML integration and output handling; Added tests for Canary 1B v2 configuration and chunk merging; Added tests for Chatterbox TTS alignment, assets, capacity, and tokenizers; Added tests for Cohere ASR configuration, long-form merging, and attention masks; Added tests for English TTS text normalization, initialism handling, and phoneme chunking; Added tests for Inflect v2 CoreML backend components; Added tests for Nemotron 3 diarizer streaming, caching, and tensor layout; Added tests for SlidingWindow ASR configuration, model loading, and streaming regression fixes; Added unit tests for Parakeet ASR configuration, constants, and decoding logic; Added unit tests for PocketTTS synthesis logic and infrastructure; Added unit tests for Sortformer diarizer components; Added unit tests for StyleTTS2 text processing and audio glue operations; Added unit tests for Supertonic-3 TTS internals; Added unit tests for TTS audio post-processing, text normalization, and language handling; Added unit tests for the offline diarizer module.
Dependencies
Initial release of FluidAudio Swift SDK and demo
This change introduces the FluidAudio Swift package (version 0.17.4) for macOS 14 and iOS 17, providing local audio AI capabilities including speaker diarization, voice-activity detection, and transcription via CoreML. The SDK includes a new \fluidaudiocli\ executable target and an optional \NemoTextProcessing\ binary dependency (v0.3.1) for byte-exact text normalization. A new \LocalVQEDemo\ example project is added to demonstrate usage, and the package is configured for Swift 6.0 with strict concurrency compliance.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Score
- CAI 62 → 62 (+0.5)
- Rubric changed (rubric-2026.09.11 → rubric-2026.09.18) — scores are not directly comparable.
Lenses
- Code Health 73 → 74 (+0.3)
- Architecture 99 → 93 (-6.3)
- Maturity 64 → 68 (+4.2)
- Readiness 54 → 55 (+0.9)
- Security 61 → 61 (+0.0)
- Domain Modelling 100 → 100 (+0.0)
- Performance 80 (new)
Resolved (41)
- AsrModels.loadVocabulary (cognitive 17) (Sources/FluidAudio/ASR/Parakeet/SlidingWindow/TDT/AsrModels.swift)
- Documentation: no installation or build instructions (README.md)
- Documentation: no usage examples (README.md)
- Duplicated block (10–11 lines × 2) (Sources/FluidAudioCLI/Commands/ASR/Parakeet/SlidingWindow/TranscribeCommand.swift)
- Duplicated block (11 lines × 2) (Sources/FluidAudioCLI/Commands/ASR/Parakeet/SlidingWindow/TranscribeCommand.swift)
- Duplicated block (13 lines × 2) (Sources/FluidAudio/Diarizer/Sortformer/SortformerStateUpdater.swift)
- Duplicated block (23–35 lines × 2) (Sources/FluidAudio/TTS/PocketTTS/Pipeline/PocketTtsSynthesizer.swift)
- Duplicated block (3–5 lines × 4) (Sources/FluidAudio/TTS/NeuTts/NeuTtsPrompt.swift)
- Duplicated block (4–9 lines × 3) (Sources/FluidAudioCLI/Utils/WERCalculator.swift)
- Duplicated block (5 lines × 2) (Sources/FluidAudio/ASR/Parakeet/Streaming/Nemotron/StreamingNemotronMultilingualAsrManager+Decode.swift)
- Duplicated block (5 lines × 3) (Sources/FluidAudio/TTS/KokoroAne/Pipeline/KokoroAneSynthesizer+Conversion.swift)
- Duplicated block (6 lines × 2) (Sources/FluidAudio/ModelRegistry.swift)
- Duplicated block (6–8 lines × 2) (Sources/FluidAudio/ASR/Parakeet/Streaming/Nemotron/StreamingNemotronMultilingualAsrManager+Decode.swift)
- Edited copy of a member (19 corresponding lines) (Sources/FluidAudio/TTS/PocketTTS/Pipeline/PocketTtsStateEngine.swift)
- Hotspot: Sources/FluidAudio/ASR/Parakeet/Streaming/EOU/StreamingEouAsrManager.swift (Sources/FluidAudio/ASR/Parakeet/Streaming/EOU/StreamingEouAsrManager.swift)
- Hotspot: Sources/FluidAudio/ASR/Parakeet/Unified/StreamingUnifiedAsrManager.swift (Sources/FluidAudio/ASR/Parakeet/Unified/StreamingUnifiedAsrManager.swift)
- Hotspot: Sources/FluidAudio/Diarizer/Offline/Clustering/VBxClustering.swift (Sources/FluidAudio/Diarizer/Offline/Clustering/VBxClustering.swift)
- Hotspot: Sources/FluidAudio/Diarizer/Sortformer/Offline/OfflineSortformerDiarizer.swift (Sources/FluidAudio/Diarizer/Sortformer/Offline/OfflineSortformerDiarizer.swift)
- Hotspot: Sources/FluidAudio/TTS/Chatterbox/Nano/ChatterboxNanoSynthesizer.swift (Sources/FluidAudio/TTS/Chatterbox/Nano/ChatterboxNanoSynthesizer.swift)
- Hotspot: Sources/FluidAudio/TTS/Chatterbox/Pipeline/ChatterboxSynthesizer.swift (Sources/FluidAudio/TTS/Chatterbox/Pipeline/ChatterboxSynthesizer.swift)
- …and 21 more
New (140)
- AsrModels.parseVocabulary (cognitive 17) (Sources/FluidAudio/ASR/Parakeet/SlidingWindow/TDT/AsrModels.swift)
- ClassTooLong: Nemotron3DiarizeCommand (Sources/FluidAudioCLI/Commands/Nemotron3DiarizeCommand.swift)
- Coverage not measured — Swift suite
- CtcDecoder.swift.ctcBeamSearch (cognitive 45) (Sources/FluidAudio/ASR/Parakeet/SlidingWindow/CTC/CtcDecoder.swift)
- CtcDecoder.swift.ctcBeamSearch (cyclomatic 20) (Sources/FluidAudio/ASR/Parakeet/SlidingWindow/CTC/CtcDecoder.swift)
- Dependency hygiene PARTLY measured — SwiftPM pinning read, dependency currency NOT established
- Duplicated block (10 lines × 2) (Sources/FluidAudio/ASR/Parakeet/AsrTypes.swift)
- Duplicated block (10 lines × 2) (Sources/FluidAudio/TTS/Inflect/InflectError.swift)
- Duplicated block (10 lines × 2) (Sources/FluidAudio/TTS/StyleTTS2/StyleTTS2Error.swift)
- Duplicated block (10 lines × 2) (Sources/FluidAudioCLI/Commands/EnhanceBenchmarkCommand.swift)
- Duplicated block (10–12 lines × 2) (Sources/FluidAudio/Diarizer/Nemotron3/Nemotron3Models.swift)
- Duplicated block (10–13 lines × 2) (Sources/FluidAudio/Diarizer/Nemotron3/Nemotron3StateUpdater.swift)
- Duplicated block (11–12 lines × 2) (Sources/FluidAudioCLI/Commands/Nemotron3DiarizeCommand.swift)
- Duplicated block (12 lines × 2) (Sources/FluidAudio/TTS/PocketTTS/PocketTTSError.swift)
- Duplicated block (12 lines × 2) (Sources/FluidAudioCLI/Commands/ASR/Parakeet/SlidingWindow/AsrBenchmark.swift)
- Duplicated block (13 lines × 2) (Sources/FluidAudioCLI/Commands/ASR/Parakeet/SlidingWindow/TranscribeCommand.swift)
- Duplicated block (14 lines × 2) (Sources/FluidAudio/ASR/Parakeet/SlidingWindow/CTC/CtcDecoder.swift)
- Duplicated block (14–15 lines × 2) (Sources/FluidAudioCLI/Commands/ASR/Parakeet/SlidingWindow/TranscribeCommand.swift)
- Duplicated block (16–17 lines × 2) (Scripts/nemo_ami_benchmark/nemo_ami_benchmark.py)
- Duplicated block (17 lines × 2) (Scripts/nemo_ami_benchmark/nemo_ami_benchmark.py)
- …and 120 more
Changes since last survey
- 33 commits — 21 feature/other, 12 fixes
By area
- Sources/FluidAudio — 16 commits
- (root) — 11 commits
- Documentation/ASR — 1 commit
- Documentation/Architecture.md — 1 commit
- Documentation/Showcase.md — 1 commit
- Examples/LocalVQEDemo — 1 commit
- Sources/FluidAudioCLI — 1 commit
- Tests/FluidAudioTests — 1 commit
Notable commits
- fix: fix(build): bump NemoTextProcessing to v0.3.1 for Mac Catalyst (#949)
- fix: fix(build): bump podspec to 0.17.1
- fix: fix(build): bump podspec to 0.17.2
- fix: fix(build): bump podspec to 0.17.3
- fix: fix(build): bump podspec to 0.17.4
- fix: fix(diarizer/nemotron3): type output backings from the model description, retry if rejected (#952)
- fix: fix(tts/english-tn): read roman-numeral list markers as numbers (#974)
- fix: fix(tts/kokoro-ane): quiet onset on long utterances — use fp32 KokoroProsody_v2 (#947) (#963)
- fix: fix(tts/kokoro-ane): restore long-text chunking in synthesizeDetailed(text:) (#940) (#965)
- fix: fix(tts/luxtts): remove spurious mid-phrase pauses and chunk long text (#937) (#942)
- fix: fix(tts/pocket): keep cache-safe long sentences whole (#938)
- fix: fix(vocab): make alignBaseWordsToUTF8Ranges iterative (#961) (#962)
- change: Add CUA-S1-FORMS Core ML scoring and benchmarks (#936)
- change: Add Orca One to the FluidAudio showcase (#953)
- change: Load local Orukeet-compatible ASR bundles without repository fallback (#928)
- change: Pin diarization artifacts to immutable revision (#927) (#939)
- change: Update README to include Trendshift badge
- change: Update README with Trendshift badge and alignment
- change: docs(asr): refresh Redux/Ultra numbers on the v3-family long-form path (#957)
- change: docs(readme): add Banter to showcase (#966)
- …and 13 more
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
FluidInference/FluidAudio was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 30 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit c388107348134698135cfd34f3f59dc823b6e7ce — the exact code this score is about.
- Scored under rubric-2026.09.18 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-cb25ca4feafa.