Skip to content
CAI
Software that uses CAICheck a score

argmaxinc/argmax-oss-swift

57.7

Adequate · 30 September 2026

32.8k

lines of production code

Swift

primary language

2

measurements over time

CAI band scale
CAI trend line
CAI lens gauges

What this system is

This system is the Argmax Open-Source SDK, a Swift-based framework that consolidates on-device speech-to-text, text-to-speech, and speaker diarization capabilities. It provides modular libraries for transcribing audio with WhisperKit, generating speech with TTSKit, and identifying speakers with SpeakerKit, all managed through a shared core for model handling and concurrency. The SDK supports local execution via a command-line interface and an HTTP server, with extensive example applications demonstrating real-time streaming and device-specific optimizations.

How it got here

2024 — SDK consolidation and core refactoring

15 changes.

The project rebranded as the Argmax Open-Source SDK, consolidating WhisperKit, TTSKit, and SpeakerKit into a unified Swift package with expanded platform support. The core transcription engine was refactored for modularity and concurrency, introducing incremental audio loading and protocol-based extensibility. Example applications were significantly enhanced with new features, while comprehensive testing infrastructure was established to validate performance and accuracy.

2025–2026 — Expansion into TTS and Diarization

14 changes.

This period focused on expanding the platform's capabilities by introducing TTSKit for on-device text-to-speech and SpeakerKit for Pyannote-based speaker diarization. The work also established ArgmaxCore as a shared foundation for model management and vendored Hugging Face dependencies, while enhancing the CLI with local server support and comprehensive test coverage for the new modules.

Features

Add WhisperKit Local Server client examples for Python and shell

New example clients are provided in Examples/ServeCLIClient to interact with the WhisperKit Local Server. The Python client (whisperkit\_client.py) uses the OpenAI SDK to transcribe and translate audio, while the Curl client provides lightweight shell scripts (transcribe.sh, translate.sh) for the same operations. Both include test suites (test\_transcribe.py, test\_translate.py, test.sh) that validate functionality against sample audio files, and the Python example includes a uv.lock file to pin dependencies.

Examples/ServeCLIClient, scripts · high confidence

Added Xcode scheme for ArgmaxOSSDynamic product

A new Xcode scheme file (argmax-oss-swift-Package.xcscheme) has been added to configure the build environment for the ArgmaxOSSDynamic product. This scheme includes build entries for ArgmaxOSSDynamic alongside other core components like WhisperKit, TTSKit, and SpeakerKit, and sets the argmax-cli as the default launch target for debugging.

.swiftpm · high confidence

Added fastlane automation and debug configuration for WhisperAX benchmarking

The WhisperAX example project now includes a new Debug.xcconfig file to allow developers to specify their Apple development team for local builds, alongside a comprehensive fastlane setup. This setup introduces lanes to list connected iOS devices, run benchmarks against specific or all available devices (including the local Mac), extract xcresult attachments, and upload results to a Hugging Face dataset, with support for both full and debug regression test configurations.

Examples/WhisperAX, fastlane · high confidence

Argmax CLI now includes a local server for WhisperKit transcription

Users can now run a local HTTP server via the new \serve\ subcommand to access WhisperKit transcription and translation capabilities over the OpenAI-compatible API. This server, built on Vapor and OpenAPI, exposes endpoints for audio transcription and translation, supports streaming responses with Server-Sent Events (SSE), and allows configuration of model selection, compute units, and network binding (host/port). The CLI also adds a dedicated \diarize\ command for speaker diarization using SpeakerKit, and enhances the \tts\ command with support for Qwen3-TTS models, including options for speech decoder modes (latency vs. throughput optimized) and multi-code decoder modes (stepped vs. fused).

Sources/ArgmaxCLI · high confidence

Argmax OSS SDK: Unified package with WhisperKit, TTSKit, and SpeakerKit

The repository has been rebranded as the Argmax Open-Source SDK (argmax-oss-swift), consolidating WhisperKit (speech-to-text), TTSKit (text-to-speech with Qwen-TTS), and SpeakerKit (speaker diarization with Pyannote) into a single Swift package. Users can now install via Swift Package Manager or Homebrew (brew install whisperkit-cli) and choose to import the umbrella ArgmaxOSS product or individual kits. The SDK now requires Xcode 16.0+ and supports Swift 6 concurrency, while WhisperKit gains incremental file loading for large audio files to reduce memory usage.

(repo-wide) · high confidence

Example app adds speaker diarization and regression testing infrastructure

The WhisperAX example app now integrates SpeakerKit to provide Pyannote-based speaker diarization support, replacing the previous MarkdownUI dependency. Additionally, the project includes a new test suite (WhisperKitTests) with functional, unit, and regression tests, along with specific audio resources and evaluation utilities to validate transcription accuracy.

Examples/WhisperAX/WhisperAX.xcodeproj · high confidence

Introduce ArgmaxCore as a shared foundation for model management and concurrency

This change introduces the new \ArgmaxCore\ module, providing a unified base for model lifecycle management, Hugging Face Hub integration, and Swift concurrency utilities. Users gain access to a \ModelManager\ that handles the download, prewarming, and loading of ML models with a clear state machine, alongside a \ModelDownloader\ for resolving model files from the Hub. The module also exposes a \Logging\ singleton for centralized, thread-safe logging across frameworks, \ConcurrencyUtilities\ including \UnfairLock\ and \Protected\ wrappers for safe state access, and \AutoTokenizerWrapper\ for loading tokenizers. Additionally, it includes platform-specific optimizations for CoreML, such as \FloatType\ handling for macOS and IOSurface-backed \MLMultiArray\ creation for float16 support.

Sources/ArgmaxCore · high confidence

Introduce Qwen3-TTS support in TTSKit

TTSKit now supports the Qwen3-TTS model family, including the 0.6B and 1.7B variants. This adds a complete generation pipeline (Qwen3GenerateTask) with dedicated CoreML-backed components for text projection, code/multi-code embedding, code/multi-code decoding, and speech decoding. The implementation supports both legacy single-function and optimized multifunction CoreML assets (stepped and fused graphs), automatically detecting model dimensions and selecting the appropriate CoreML function (latency vs. throughput) at load time. Configuration is handled via TTSKitConfig, allowing selection of model variants, local or HuggingFace model folders, and specific component quantization variants.

Sources/TTSKit/Qwen3TTS · high confidence

Introduce SpeakerKit for Pyannote-based speaker diarization

Adds the SpeakerKit module, providing a new capability to perform speaker diarization on audio using Pyannote models. The library exposes a \Diarizer\ protocol and a \PyannoteDiarizer\ implementation that handles model loading, downloading, and inference. Users can now segment audio into speaker turns, access per-speaker centroid embeddings for linking speakers across calls, and configure runtime options such as concurrency, model variants (W8A16/W32A32), and clustering thresholds via \PyannoteConfig\ and \PyannoteDiarizationOptions\.

Sources/SpeakerKit · high confidence

Introduce TTSKit for local speech synthesis

Adds TTSKit, a new Swift framework for on-device text-to-speech using the Qwen3-TTS model. The library provides a modular architecture with swappable CoreML components (text projector, code embedders, decoders) and a unified \SpeechModel\ protocol. It supports streaming audio generation with adaptive playback buffering to prevent underruns, concurrent chunk processing, and prompt caching for repeated voices or languages.

Sources/TTSKit · high confidence

Introduce TTSKit with Qwen3-TTS support and audio utilities

Adds the TTSKit module to handle text-to-speech generation using the Qwen3-TTS model. This includes an AudioOutput component for real-time streaming playback with adaptive pre-buffering and edge-fading to prevent clicks, a TextChunker for splitting long text into sentence-bounded segments, and token sampling utilities (GreedyTokenSampler) for codec generation. The kit also provides KVCache and PromptCache implementations to manage and persist model state for efficient autoregressive decoding, along with EmbedTypes and EmbedUtilities to handle CoreML tensor operations across platforms.

Sources/TTSKit/Utilities · high confidence

New TTSKitExample demo app for on-device text-to-speech

The TTSKitExample demo app is now available in the Examples/TTS directory, providing a complete SwiftUI interface for the TTSKit framework. Users can download and manage on-device Core ML models (0.6B and 1.7B variants), generate speech with real-time streaming and live waveform visualization, and configure advanced generation settings like temperature, chunking, and compute units. The app supports 9 voices across 10 languages, allows saving generation history as M4A files with embedded metadata, and runs on macOS 15.2+ and iOS 18.2+.

Examples/TTS · high confidence

Support for incremental audio file loading and VAD-based chunking

Users can now transcribe large audio files with significantly reduced memory usage by switching to incremental loading. The new \AudioLoadingMode.incremental\ option streams audio in bounded-memory chunks, using a back-pressure mechanism to pause loading when the buffer is full. This mode relies on the new \VADAudioChunker\, which splits audio into segments based on voice activity detection (silence boundaries) to ensure efficient processing. The \AudioProcessor\ now exposes \AudioInputOptions\ to configure this behavior, allowing users to balance memory consumption against processing latency for long recordings.

Sources/WhisperKit/Core/Audio · high confidence

Watch app gains full model selection, settings, and streaming transcription UI

The WhisperAXWatchApp has been significantly expanded from a simple single-button demo into a feature-rich interface. Users can now select from multiple available Whisper models, configure inference parameters (such as language, silence threshold, and VAD settings), and view real-time transcription with streaming text output. The UI includes a navigation split view with a model selector, status indicators, and a detailed settings panel, replacing the previous static 'Transcribe Example' button.

Examples/WhisperAX/WhisperAXWatchApp · high confidence

Removals

Removal of WhisperKitCLI main executable

The \transcribe.swift\ file containing the \WhisperKitCLI\ command-line tool has been deleted from the source tree. This removes the ability to run audio transcription directly via the CLI using the previously defined arguments for model paths, compute units, and decoding options.

Sources/WhisperKitCLI · high confidence

Behavioural changes

Core transcription pipeline refactored for modularity and protocol-based extensibility

The transcription engine in Sources/WhisperKit/Core has been restructured to replace concrete implementations with protocols (AudioProcessing, FeatureExtracting, AudioEncoding, TextDecoding, SegmentSeeking) and generic output types, allowing for custom model backends and easier testing. The new \TranscribeTask\ class orchestrates this pipeline, introducing extensibility hooks (\windowPreprocess\, \windowPostProcess\) for subclasses to modify audio windows or segment results. Configuration is now centralized in \WhisperKitConfig\ and \DecodingOptions\, with \ModelComputeOptions\ updated to handle simulator-specific defaults and platform-specific compute unit fallbacks. Legacy files like \AudioProcessor.swift\, \LogitsFilter.swift\, and \SegmentSeeker.swift\ have been removed or replaced by these protocol-based abstractions, and the \ArgmaxCore\ dependency is now imported to support the new tensor types.

Sources/WhisperKit/Core · high confidence

Enhanced transcription controls and speaker diarization support

The example app now includes settings for speaker diarization via SpeakerKit, allowing users to enable and configure speaker identification modes. Transcription behavior has been refined with new options for VAD chunking strategies, token confirmation requirements, and concurrent worker counts. Additionally, the default fallback count has been increased to 5, the compression check window extended to 60, and the required segments for confirmation raised to 4, while the eager decoder setting has been renamed to 'eager decoding'.

Examples/WhisperAX/WhisperAX/Views · high confidence

New logits filtering and improved word alignment in transcription

The transcription engine now includes a new LogitsFilter system (LogitsFilter.swift) that enforces timestamp rules and suppresses specific tokens during decoding, ensuring more accurate segment timing and text output. Additionally, the SegmentSeeker has been refactored to improve word alignment and timestamp handling, resulting in better segmentation of audio into text segments. The TokenSampler has been moved to the Text module and updated to support Swift 6 concurrency, with async sampling methods and Sendable compliance for safer concurrent execution.

Sources/WhisperKit/Core/Text · high confidence

Refactored utilities and enhanced subtitle export with word-level timing

The Utilities module has been reorganized into dedicated files for concurrency, logging, model loading, and text/audio processing, including a new \TranscriptionUtilities\ struct for merging results and adjusting segment timings. A key behavioral change occurs in the \ResultWriter\ (moved from Core to Utilities): SRT and VTT exports now support word-level timestamps when available, iterating through \wordTimings\ to produce granular subtitles instead of only segment-level timing. Additionally, \WriteJSON\, \WriteSRT\, and \WriteVTT\ classes are now open for subclassing, and a backward-compatible typealias \TranscriptionPropertyLock\ is provided for \PropertyLock\.

Sources/WhisperKit/Utilities · high confidence

Vendor Hugging Face Hub and Tokenizers logic from swift-transformers

The Hub and Tokenizers subsystems in ArgmaxCore are now implemented by vendoring code from the Hugging Face swift-transformers library (version 1.1.6). This change brings in core functionality for downloading model files from the Hugging Face Hub, managing repository snapshots, and executing various tokenizer algorithms (including BPE and WordPiece). The implementation has been adapted to be internal to ArgmaxCore, removing external dependencies like Jinja and the Xet protocol, and introducing a \BinaryDistinctString\ type to ensure exact byte-sequence preservation for tokenizer vocabularies.

Sources/ArgmaxCore/External · high confidence

Fixes

Added Xcode scheme for WhisperAX example app

The WhisperAX example project now includes a defined Xcode scheme (WhisperAX.xcscheme) that configures the build, test, and launch actions. This scheme enables parallel builds, includes both unit and UI tests in the test action, and configures environment variables (MODEL\_NAME, MODEL\_REPO) for the launch action, facilitating easier local development and regression testing of the example app.

Examples/WhisperAX/WhisperAX.xcodeproj/xcshareddata · high confidence

Test coverage

Added comprehensive test suite for audio processing, streaming transcription, and regression evaluation; Added test suite for TTSKit text-to-speech functionality; Added tests for vendored Hugging Face Swift Transformers components; Cleanup of test file formatting; Initial test suite for SpeakerKit diarization and clustering; Updated test target import to match current app bundle.

Dependencies

Package renamed to argmax-oss-swift with new libraries and platform support

The Swift package has been renamed from 'whisperkit' to 'argmax-oss-swift' and now includes new libraries: TTSKit (text-to-speech), SpeakerKit (speaker diarization), and ArgmaxOSS (a combined library). The package now supports watchOS 10 and visionOS 1 in addition to iOS and macOS. The dependency on swift-transformers has been removed, with functionality integrated directly into ArgmaxCore. The CLI tool has been renamed to 'whisperkit-cli' (with 'argmax-cli' also available). Server-side dependencies (Vapor, OpenAPI) are now conditionally included. The minimum Swift version is now 5.10.

(dependencies) · high confidence

Housekeeping

Whitespace cleanup in UI test files

Removed unnecessary blank lines from the top of the class definitions in the WhisperAX and WhisperAX Watch App UI test source files to improve code formatting consistency.

Examples/WhisperAX/WhisperAXUITests, Examples/WhisperAX/WhisperAXWatchAppUITests · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

Score

  • CAI 63 → 58 (-5.1)
  • Rubric changed (rubric-2026.09.11 → rubric-2026.09.18) — scores are not directly comparable.

Lenses

  • Code Health 83 → 85 (+1.6)
  • Architecture 100 → 91 (-8.8)
  • Maturity 70 → 70 (+0.0)
  • Readiness 54 → 41 (-12.9)
  • Security 60 → 72 (+12.3)
  • Domain Modelling 100 → 100 (+0.0)
  • Performance 68 (new)

Resolved (8)

  • Documentation: no installation or build instructions (README.md)
  • Duplicated block (6–10 lines × 2) (Examples/ServeCLIClient/Swift/Sources/WhisperKitSwiftClient/CLI.swift)
  • Hotspot: Sources/ArgmaxCLI/TTSCLI.swift (Sources/ArgmaxCLI/TTSCLI.swift)
  • Hotspot: Sources/TTSKit/TTSKit.swift (Sources/TTSKit/TTSKit.swift)
  • Hotspot: Sources/TTSKit/Utilities/KVCache.swift (Sources/TTSKit/Utilities/KVCache.swift)
  • Hotspot: Sources/WhisperKit/Core/TextDecoder.swift (Sources/WhisperKit/Core/TextDecoder.swift)
  • Members sharing a duplicated core (5 members, 50+ identical tokens) (Sources/TTSKit/Qwen3TTS/Qwen3Config.swift)
  • Off-boarding risk: anonymized user #1

New (46)

  • Coverage not measured — Swift suite
  • Dependency not covered by the committed resolution: swift-openapi-generator
  • Dependency not covered by the committed resolution: swift-openapi-runtime
  • Dependency not covered by the committed resolution: swift-openapi-vapor
  • Dependency not covered by the committed resolution: vapor
  • Documentation: no project overview (README.md)
  • Duplicated block (10 lines × 2) (Examples/ServeCLIClient/Swift/Sources/WhisperKitSwiftClient/CLI.swift)
  • Duplicated block (10–11 lines × 2) (Examples/ServeCLIClient/Python/whisperkit_client.py)
  • Duplicated block (11–12 lines × 2) (Examples/TTS/TTSKitExample/TTSKitExample/DetailView.swift)
  • Duplicated block (12 lines × 2) (Examples/WhisperAX/WhisperAX/Views/ContentView.swift)
  • Duplicated block (13 lines × 2) (Examples/TTS/TTSKitExample/TTSKitExample/MultiCodeDecoderModeView.swift)
  • Duplicated block (13 lines × 3) (Sources/TTSKit/Qwen3TTS/Qwen3Models.swift)
  • Duplicated block (14 lines × 2) (Examples/ServeCLIClient/Python/whisperkit_client.py)
  • Duplicated block (15 lines × 2) (Examples/ServeCLIClient/Python/whisperkit_client.py)
  • Duplicated block (4–5 lines × 4) (Sources/SpeakerKit/Pyannote/SpeakerEmbedderModel.swift)
  • Duplicated block (5 lines × 3) (Sources/WhisperKit/Core/AudioEncoder.swift)
  • Duplicated block (6 lines × 2) (Examples/WhisperAX/WhisperAX/Views/ContentView.swift)
  • Duplicated block (6–10 lines × 2) (Examples/ServeCLIClient/Swift/Sources/WhisperKitSwiftClient/CLI.swift)
  • Duplicated block (7 lines × 2) (Examples/WhisperAX/WhisperAX/Views/ContentView.swift)
  • Duplicated block (7 lines × 2) (Sources/TTSKit/Qwen3TTS/Qwen3CodeDecoder.swift)
  • …and 26 more

Changes since last survey

  • 5 commits — 4 feature/other, 1 fixes

By area

  • Sources/TTSKit — 2 commits
  • Sources/WhisperKit — 2 commits
  • Sources/ArgmaxCore — 1 commit

Notable commits

  • fix: fix: stop recording when streaming transcription exits on error (#533)
  • change: Don't log the model variant before it is detected (#536)
  • change: Reuse compatible iOS audio session in TTSKit playback (#464)
  • change: TTSKit: mark the CoreML import in Qwen3SpeechDecoder as @preconcurrency (#524)
  • change: Throw when a Hub snapshot download is cancelled (#534)

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

argmaxinc/argmax-oss-swift was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 30 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit f4e5d6be37ec820614fb0d72037e76c22d4c16f7 — the exact code this score is about.
  • Scored under rubric-2026.09.18 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-cb25ca4feafa.