argmaxinc/argmax-oss-swift
57.7
Adequate · 30 September 2026
32.8k
lines of production code
Swift
primary language
2
measurements over time
What this system is
This system is the Argmax Open-Source SDK, a Swift-based framework that consolidates on-device speech-to-text, text-to-speech, and speaker diarization capabilities. It provides modular libraries for transcribing audio with WhisperKit, generating speech with TTSKit, and identifying speakers with SpeakerKit, all managed through a shared core for model handling and concurrency. The SDK supports local execution via a command-line interface and an HTTP server, with extensive example applications demonstrating real-time streaming and device-specific optimizations.
How it got here
2024 — SDK consolidation and core refactoring
15 changes.
The project rebranded as the Argmax Open-Source SDK, consolidating WhisperKit, TTSKit, and SpeakerKit into a unified Swift package with expanded platform support. The core transcription engine was refactored for modularity and concurrency, introducing incremental audio loading and protocol-based extensibility. Example applications were significantly enhanced with new features, while comprehensive testing infrastructure was established to validate performance and accuracy.
2025–2026 — Expansion into TTS and Diarization
14 changes.
This period focused on expanding the platform's capabilities by introducing TTSKit for on-device text-to-speech and SpeakerKit for Pyannote-based speaker diarization. The work also established ArgmaxCore as a shared foundation for model management and vendored Hugging Face dependencies, while enhancing the CLI with local server support and comprehensive test coverage for the new modules.
Features
Add WhisperKit Local Server client examples for Python and shell
New example clients are provided in Examples/ServeCLIClient to interact with the WhisperKit Local Server. The Python client (whisperkit\_client.py) uses the OpenAI SDK to transcribe and translate audio, while the Curl client provides lightweight shell scripts (transcribe.sh, translate.sh) for the same operations. Both include test suites (test\_transcribe.py, test\_translate.py, test.sh) that validate functionality against sample audio files, and the Python example includes a uv.lock file to pin dependencies.
Examples/ServeCLIClient, scripts · high confidence
Added Xcode scheme for ArgmaxOSSDynamic product
A new Xcode scheme file (argmax-oss-swift-Package.xcscheme) has been added to configure the build environment for the ArgmaxOSSDynamic product. This scheme includes build entries for ArgmaxOSSDynamic alongside other core components like WhisperKit, TTSKit, and SpeakerKit, and sets the argmax-cli as the default launch target for debugging.
.swiftpm · high confidence
Added fastlane automation and debug configuration for WhisperAX benchmarking
The WhisperAX example project now includes a new Debug.xcconfig file to allow developers to specify their Apple development team for local builds, alongside a comprehensive fastlane setup. This setup introduces lanes to list connected iOS devices, run benchmarks against specific or all available devices (including the local Mac), extract xcresult attachments, and upload results to a Hugging Face dataset, with support for both full and debug regression test configurations.
Examples/WhisperAX, fastlane · high confidence
Argmax CLI now includes a local server for WhisperKit transcription
Users can now run a local HTTP server via the new \serve\ subcommand to access WhisperKit transcription and translation capabilities over the OpenAI-compatible API. This server, built on Vapor and OpenAPI, exposes endpoints for audio transcription and translation, supports streaming responses with Server-Sent Events (SSE), and allows configuration of model selection, compute units, and network binding (host/port). The CLI also adds a dedicated \diarize\ command for speaker diarization using SpeakerKit, and enhances the \tts\ command with support for Qwen3-TTS models, including options for speech decoder modes (latency vs. throughput optimized) and multi-code decoder modes (stepped vs. fused).
Sources/ArgmaxCLI · high confidence
Argmax OSS SDK: Unified package with WhisperKit, TTSKit, and SpeakerKit
The repository has been rebranded as the Argmax Open-Source SDK (argmax-oss-swift), consolidating WhisperKit (speech-to-text), TTSKit (text-to-speech with Qwen-TTS), and SpeakerKit (speaker diarization with Pyannote) into a single Swift package. Users can now install via Swift Package Manager or Homebrew (brew install whisperkit-cli) and choose to import the umbrella ArgmaxOSS product or individual kits. The SDK now requires Xcode 16.0+ and supports Swift 6 concurrency, while WhisperKit gains incremental file loading for large audio files to reduce memory usage.
(repo-wide) · high confidence
Example app adds speaker diarization and regression testing infrastructure
The WhisperAX example app now integrates SpeakerKit to provide Pyannote-based speaker diarization support, replacing the previous MarkdownUI dependency. Additionally, the project includes a new test suite (WhisperKitTests) with functional, unit, and regression tests, along with specific audio resources and evaluation utilities to validate transcription accuracy.
Examples/WhisperAX/WhisperAX.xcodeproj · high confidence
Introduce ArgmaxCore as a shared foundation for model management and concurrency
This change introduces the new \ArgmaxCore\ module, providing a unified base for model lifecycle management, Hugging Face Hub integration, and Swift concurrency utilities. Users gain access to a \ModelManager\ that handles the download, prewarming, and loading of ML models with a clear state machine, alongside a \ModelDownloader\ for resolving model files from the Hub. The module also exposes a \Logging\ singleton for centralized, thread-safe logging across frameworks, \ConcurrencyUtilities\ including \UnfairLock\ and \Protected\ wrappers for safe state access, and \AutoTokenizerWrapper\ for loading tokenizers. Additionally, it includes platform-specific optimizations for CoreML, such as \FloatType\ handling for macOS and IOSurface-backed \MLMultiArray\ creation for float16 support.
Sources/ArgmaxCore · high confidence
Introduce Qwen3-TTS support in TTSKit
TTSKit now supports the Qwen3-TTS model family, including the 0.6B and 1.7B variants. This adds a complete generation pipeline (Qwen3GenerateTask) with dedicated CoreML-backed components for text projection, code/multi-code embedding, code/multi-code decoding, and speech decoding. The implementation supports both legacy single-function and optimized multifunction CoreML assets (stepped and fused graphs), automatically detecting model dimensions and selecting the appropriate CoreML function (latency vs. throughput) at load time. Configuration is handled via TTSKitConfig, allowing selection of model variants, local or HuggingFace model folders, and specific component quantization variants.
Sources/TTSKit/Qwen3TTS · high confidence
Introduce SpeakerKit for Pyannote-based speaker diarization
Adds the SpeakerKit module, providing a new capability to perform speaker diarization on audio using Pyannote models. The library exposes a \Diarizer\ protocol and a \PyannoteDiarizer\ implementation that handles model loading, downloading, and inference. Users can now segment audio into speaker turns, access per-speaker centroid embeddings for linking speakers across calls, and configure runtime options such as concurrency, model variants (W8A16/W32A32), and clustering thresholds via \PyannoteConfig\ and \PyannoteDiarizationOptions\.
Sources/SpeakerKit · high confidence
Introduce TTSKit for local speech synthesis
Adds TTSKit, a new Swift framework for on-device text-to-speech using the Qwen3-TTS model. The library provides a modular architecture with swappable CoreML components (text projector, code embedders, decoders) and a unified \SpeechModel\ protocol. It supports streaming audio generation with adaptive playback buffering to prevent underruns, concurrent chunk processing, and prompt caching for repeated voices or languages.
Sources/TTSKit · high confidence
Introduce TTSKit with Qwen3-TTS support and audio utilities
Adds the TTSKit module to handle text-to-speech generation using the Qwen3-TTS model. This includes an AudioOutput component for real-time streaming playback with adaptive pre-buffering and edge-fading to prevent clicks, a TextChunker for splitting long text into sentence-bounded segments, and token sampling utilities (GreedyTokenSampler) for codec generation. The kit also provides KVCache and PromptCache implementations to manage and persist model state for efficient autoregressive decoding, along with EmbedTypes and EmbedUtilities to handle CoreML tensor operations across platforms.
Sources/TTSKit/Utilities · high confidence
New TTSKitExample demo app for on-device text-to-speech
The TTSKitExample demo app is now available in the Examples/TTS directory, providing a complete SwiftUI interface for the TTSKit framework. Users can download and manage on-device Core ML models (0.6B and 1.7B variants), generate speech with real-time streaming and live waveform visualization, and configure advanced generation settings like temperature, chunking, and compute units. The app supports 9 voices across 10 languages, allows saving generation history as M4A files with embedded metadata, and runs on macOS 15.2+ and iOS 18.2+.
Examples/TTS · high confidence
Support for incremental audio file loading and VAD-based chunking
Users can now transcribe large audio files with significantly reduced memory usage by switching to incremental loading. The new \AudioLoadingMode.incremental\ option streams audio in bounded-memory chunks, using a back-pressure mechanism to pause loading when the buffer is full. This mode relies on the new \VADAudioChunker\, which splits audio into segments based on voice activity detection (silence boundaries) to ensure efficient processing. The \AudioProcessor\ now exposes \AudioInputOptions\ to configure this behavior, allowing users to balance memory consumption against processing latency for long recordings.
Sources/WhisperKit/Core/Audio · high confidence
Watch app gains full model selection, settings, and streaming transcription UI
The WhisperAXWatchApp has been significantly expanded from a simple single-button demo into a feature-rich interface. Users can now select from multiple available Whisper models, configure inference parameters (such as language, silence threshold, and VAD settings), and view real-time transcription with streaming text output. The UI includes a navigation split view with a model selector, status indicators, and a detailed settings panel, replacing the previous static 'Transcribe Example' button.
Examples/WhisperAX/WhisperAXWatchApp · high confidence
Removals
Removal of WhisperKitCLI main executable
The \transcribe.swift\ file containing the \WhisperKitCLI\ command-line tool has been deleted from the source tree. This removes the ability to run audio transcription directly via the CLI using the previously defined arguments for model paths, compute units, and decoding options.
Sources/WhisperKitCLI · high confidence
Behavioural changes
Core transcription pipeline refactored for modularity and protocol-based extensibility
The transcription engine in Sources/WhisperKit/Core has been restructured to replace concrete implementations with protocols (AudioProcessing, FeatureExtracting, AudioEncoding, TextDecoding, SegmentSeeking) and generic output types, allowing for custom model backends and easier testing. The new \TranscribeTask\ class orchestrates this pipeline, introducing extensibility hooks (\windowPreprocess\, \windowPostProcess\) for subclasses to modify audio windows or segment results. Configuration is now centralized in \WhisperKitConfig\ and \DecodingOptions\, with \ModelComputeOptions\ updated to handle simulator-specific defaults and platform-specific compute unit fallbacks. Legacy files like \AudioProcessor.swift\, \LogitsFilter.swift\, and \SegmentSeeker.swift\ have been removed or replaced by these protocol-based abstractions, and the \ArgmaxCore\ dependency is now imported to support the new tensor types.
Sources/WhisperKit/Core · high confidence
Enhanced transcription controls and speaker diarization support
The example app now includes settings for speaker diarization via SpeakerKit, allowing users to enable and configure speaker identification modes. Transcription behavior has been refined with new options for VAD chunking strategies, token confirmation requirements, and concurrent worker counts. Additionally, the default fallback count has been increased to 5, the compression check window extended to 60, and the required segments for confirmation raised to 4, while the eager decoder setting has been renamed to 'eager decoding'.
Examples/WhisperAX/WhisperAX/Views · high confidence
New logits filtering and improved word alignment in transcription
The transcription engine now includes a new LogitsFilter system (LogitsFilter.swift) that enforces timestamp rules and suppresses specific tokens during decoding, ensuring more accurate segment timing and text output. Additionally, the SegmentSeeker has been refactored to improve word alignment and timestamp handling, resulting in better segmentation of audio into text segments. The TokenSampler has been moved to the Text module and updated to support Swift 6 concurrency, with async sampling methods and Sendable compliance for safer concurrent execution.
Sources/WhisperKit/Core/Text · high confidence
Refactored utilities and enhanced subtitle export with word-level timing
The Utilities module has been reorganized into dedicated files for concurrency, logging, model loading, and text/audio processing, including a new \TranscriptionUtilities\ struct for merging results and adjusting segment timings. A key behavioral change occurs in the \ResultWriter\ (moved from Core to Utilities): SRT and VTT exports now support word-level timestamps when available, iterating through \wordTimings\ to produce granular subtitles instead of only segment-level timing. Additionally, \WriteJSON\, \WriteSRT\, and \WriteVTT\ classes are now open for subclassing, and a backward-compatible typealias \TranscriptionPropertyLock\ is provided for \PropertyLock\.
Sources/WhisperKit/Utilities · high confidence
Vendor Hugging Face Hub and Tokenizers logic from swift-transformers
The Hub and Tokenizers subsystems in ArgmaxCore are now implemented by vendoring code from the Hugging Face swift-transformers library (version 1.1.6). This change brings in core functionality for downloading model files from the Hugging Face Hub, managing repository snapshots, and executing various tokenizer algorithms (including BPE and WordPiece). The implementation has been adapted to be internal to ArgmaxCore, removing external dependencies like Jinja and the Xet protocol, and introducing a \BinaryDistinctString\ type to ensure exact byte-sequence preservation for tokenizer vocabularies.
Sources/ArgmaxCore/External · high confidence
Fixes
Added Xcode scheme for WhisperAX example app
The WhisperAX example project now includes a defined Xcode scheme (WhisperAX.xcscheme) that configures the build, test, and launch actions. This scheme enables parallel builds, includes both unit and UI tests in the test action, and configures environment variables (MODEL\_NAME, MODEL\_REPO) for the launch action, facilitating easier local development and regression testing of the example app.
Examples/WhisperAX/WhisperAX.xcodeproj/xcshareddata · high confidence
Test coverage
Added comprehensive test suite for audio processing, streaming transcription, and regression evaluation; Added test suite for TTSKit text-to-speech functionality; Added tests for vendored Hugging Face Swift Transformers components; Cleanup of test file formatting; Initial test suite for SpeakerKit diarization and clustering; Updated test target import to match current app bundle.
Dependencies
Package renamed to argmax-oss-swift with new libraries and platform support
The Swift package has been renamed from 'whisperkit' to 'argmax-oss-swift' and now includes new libraries: TTSKit (text-to-speech), SpeakerKit (speaker diarization), and ArgmaxOSS (a combined library). The package now supports watchOS 10 and visionOS 1 in addition to iOS and macOS. The dependency on swift-transformers has been removed, with functionality integrated directly into ArgmaxCore. The CLI tool has been renamed to 'whisperkit-cli' (with 'argmax-cli' also available). Server-side dependencies (Vapor, OpenAPI) are now conditionally included. The minimum Swift version is now 5.10.
(dependencies) · high confidence
Housekeeping
Whitespace cleanup in UI test files
Removed unnecessary blank lines from the top of the class definitions in the WhisperAX and WhisperAX Watch App UI test source files to improve code formatting consistency.
Examples/WhisperAX/WhisperAXUITests, Examples/WhisperAX/WhisperAXWatchAppUITests · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Score
- CAI 63 → 58 (-5.1)
- Rubric changed (rubric-2026.09.11 → rubric-2026.09.18) — scores are not directly comparable.
Lenses
- Code Health 83 → 85 (+1.6)
- Architecture 100 → 91 (-8.8)
- Maturity 70 → 70 (+0.0)
- Readiness 54 → 41 (-12.9)
- Security 60 → 72 (+12.3)
- Domain Modelling 100 → 100 (+0.0)
- Performance 68 (new)
Resolved (8)
- Documentation: no installation or build instructions (README.md)
- Duplicated block (6–10 lines × 2) (Examples/ServeCLIClient/Swift/Sources/WhisperKitSwiftClient/CLI.swift)
- Hotspot: Sources/ArgmaxCLI/TTSCLI.swift (Sources/ArgmaxCLI/TTSCLI.swift)
- Hotspot: Sources/TTSKit/TTSKit.swift (Sources/TTSKit/TTSKit.swift)
- Hotspot: Sources/TTSKit/Utilities/KVCache.swift (Sources/TTSKit/Utilities/KVCache.swift)
- Hotspot: Sources/WhisperKit/Core/TextDecoder.swift (Sources/WhisperKit/Core/TextDecoder.swift)
- Members sharing a duplicated core (5 members, 50+ identical tokens) (Sources/TTSKit/Qwen3TTS/Qwen3Config.swift)
- Off-boarding risk: anonymized user #1
New (46)
- Coverage not measured — Swift suite
- Dependency not covered by the committed resolution: swift-openapi-generator
- Dependency not covered by the committed resolution: swift-openapi-runtime
- Dependency not covered by the committed resolution: swift-openapi-vapor
- Dependency not covered by the committed resolution: vapor
- Documentation: no project overview (README.md)
- Duplicated block (10 lines × 2) (Examples/ServeCLIClient/Swift/Sources/WhisperKitSwiftClient/CLI.swift)
- Duplicated block (10–11 lines × 2) (Examples/ServeCLIClient/Python/whisperkit_client.py)
- Duplicated block (11–12 lines × 2) (Examples/TTS/TTSKitExample/TTSKitExample/DetailView.swift)
- Duplicated block (12 lines × 2) (Examples/WhisperAX/WhisperAX/Views/ContentView.swift)
- Duplicated block (13 lines × 2) (Examples/TTS/TTSKitExample/TTSKitExample/MultiCodeDecoderModeView.swift)
- Duplicated block (13 lines × 3) (Sources/TTSKit/Qwen3TTS/Qwen3Models.swift)
- Duplicated block (14 lines × 2) (Examples/ServeCLIClient/Python/whisperkit_client.py)
- Duplicated block (15 lines × 2) (Examples/ServeCLIClient/Python/whisperkit_client.py)
- Duplicated block (4–5 lines × 4) (Sources/SpeakerKit/Pyannote/SpeakerEmbedderModel.swift)
- Duplicated block (5 lines × 3) (Sources/WhisperKit/Core/AudioEncoder.swift)
- Duplicated block (6 lines × 2) (Examples/WhisperAX/WhisperAX/Views/ContentView.swift)
- Duplicated block (6–10 lines × 2) (Examples/ServeCLIClient/Swift/Sources/WhisperKitSwiftClient/CLI.swift)
- Duplicated block (7 lines × 2) (Examples/WhisperAX/WhisperAX/Views/ContentView.swift)
- Duplicated block (7 lines × 2) (Sources/TTSKit/Qwen3TTS/Qwen3CodeDecoder.swift)
- …and 26 more
Changes since last survey
- 5 commits — 4 feature/other, 1 fixes
By area
- Sources/TTSKit — 2 commits
- Sources/WhisperKit — 2 commits
- Sources/ArgmaxCore — 1 commit
Notable commits
- fix: fix: stop recording when streaming transcription exits on error (#533)
- change: Don't log the model variant before it is detected (#536)
- change: Reuse compatible iOS audio session in TTSKit playback (#464)
- change: TTSKit: mark the CoreML import in Qwen3SpeechDecoder as @preconcurrency (#524)
- change: Throw when a Hub snapshot download is cancelled (#534)
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
argmaxinc/argmax-oss-swift was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 30 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit f4e5d6be37ec820614fb0d72037e76c22d4c16f7 — the exact code this score is about.
- Scored under rubric-2026.09.18 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-cb25ca4feafa.