apple/coreai-models
63.7
Adequate · 1 October 2026
65.4k
lines of production code
Swift
with Python
2
measurements over time
What this system is
This system is a Swift-based runtime and tooling suite for executing AI models on Apple devices using Core AI. It provides infrastructure for exporting, loading, and running a wide variety of models, including large language models, vision-language models, image segmenters, object detectors, speech recognizers, and video diffusion pipelines. The system supports both command-line execution and an OpenAI-compatible HTTP server, enabling features like streaming transcription, guided text generation, and real-time video processing with optimized memory and performance management.
Features
Add DiffusionGemma-26B-A4B Core AI export support
Users can now export the Google DiffusionGemma-26B-A4B model for on-device inference via Core AI. This new capability adds an export script and documentation for models/diffusion\_gemma, supporting the generation of a fixed-length token canvas through iterative denoising. The export produces two Core AI components—an autoregressive encoder and a bidirectional decoder—and supports 4-bit weight quantization for the macOS platform.
_models/diffusion\gemma · high confidence
Add Parakeet TDT live streaming transcription support
Introduces a new live transcription capability for the Parakeet TDT model architecture, enabling real-time, chunked audio processing with incremental text updates. This change adds the necessary Swift inference components, including a mel spectrogram front-end, a transducer decoder with streaming state management, and a streaming session API that handles windowed audio chunks, endpointing, and partial transcription delivery.
swift/Sources/CoreAISpeech · high confidence
Add Wan 2.1 text-to-video generation pipeline
Introduces a new Wan 2.1 text-to-video pipeline that orchestrates tokenization, text encoding, denoising, and 3D VAE decoding to generate video frames. The implementation includes a new \VideoPipeline\ protocol and \VideoConfiguration\ struct, supporting quality presets (fast, balanced, best) and tiled VAE decoding to manage memory usage for high-resolution outputs.
swift/Sources/CoreAIVideoDiffusionPipeline · high confidence
Add Wan 2.1 text-to-video generation tool
Introduces the \videodiffusion-runner\ command-line tool, enabling users to generate videos from text prompts using the Wan 2.1 diffusion model. The tool supports configurable parameters for video quality, resolution, frame count, and duration, and includes advanced features such as tiled VAE decoding to reduce memory usage, CFG cutoff optimization, and intermediate latent dumping for debugging.
swift/Sources/Tools/videodiffusion-runner · high confidence
Add export support for Parakeet, Sana Sprint, Wan 2.1, OLMo 2, and SAM3 models
Users can now export several new model families to Core AI. The Parakeet TDT model is exported as three separate graphs (encoder, decoder step, joint) to support its autoregressive transducer loop, with configurable streaming window geometry. Sana Sprint 0.6B and Wan 2.1 text-to-video pipelines are supported with dedicated torch wrappers and dummy-input factories for export. OLMo 2 is supported for iOS with a re-authored transformer implementation using iOS-specific primitives. SAM3 Lite is supported for iOS with re-authored DETR, FPN, image encoder, and mask decoder components adapted to the BC1S layout.
python · high confidence
Add video input and output support for vision-language models
Developers can now pass video files to vision-language models. The \VideoInput\ API extracts frames from local video files using configurable sampling strategies (uniform count or fixed frame rate) and delivers them lazily to keep memory usage low. Additionally, the \VideoWriter\ and \StreamingVideoWriter\ components allow encoding generated frames into MP4, GIF, APNG, or WebP formats, with streaming support for memory-efficient sequential writing.
swift/Sources/CoreAIShared/Video · high confidence
Introduce DiffusionBundle and FlowTransformerPipeline for FLUX.2 and Sana Sprint
The diffusion pipeline now uses a unified \DiffusionBundle\ to load model assets and configuration from \metadata.json\ (schema 0.2), replacing the previous descriptor-based loading. This change introduces \FlowTransformerPipeline\, which supports two new model families: FLUX.2 (using Qwen3 text encoding, in-graph RoPE, and flow-matching Euler sampling) and Sana Sprint (using Gemma2 text encoding, TrigFlow/SCM stochastic sampling, and instruction-prefixed prompts). The pipeline also adds latent preview capabilities via \LatentPreview\ and \LatentPreviewTuner\, allowing real-time RGB previews from latent tensors using fitted linear projection coefficients.
swift/Sources/CoreAIDiffusionPipeline/Pipelines · high confidence
Introduce OpenAI-compatible LLM inference server
The llm-server tool now provides an HTTP server that implements the OpenAI API interface, exposing endpoints for chat completions (/v1/chat/completions), text completions (/v1/completions), and model listing (/v1/models). The server supports both streaming and non-streaming generation, includes a /ready health check and a /v1/stats endpoint for monitoring performance metrics, and allows users to control inference behavior via command-line arguments for sampling parameters, KV cache strategies, and request queue depth. A --replay mode is also available to process JSONL request files directly for testing and benchmarking without starting the HTTP service.
swift/Sources/Tools/llm-server · high confidence
Introduces core data models and concurrency controls for the LLM server
This change adds the foundational Swift types and infrastructure required for the new OpenAI-compatible LLM server. It defines request and response structures for both legacy completions and modern chat endpoints, including support for tool calls, reasoning effort controls, and reproducible generation seeds. Additionally, it introduces an async request queue to manage concurrency limits and a replay format to allow deterministic testing and debugging of server interactions.
swift/Sources/CoreAILMCommon · high confidence
New input and state handler infrastructure for static-shape and pipelined engines
The CoreAI language model runtime now includes a comprehensive set of new input and state handlers to support static-shape and pipelined inference engines. This adds \StaticInputHandler\ and \StaticBucketKey\ protocols for zero-allocation, pre-allocated buffer management across different batch and context buckets. New specific handlers include \DualRoPEInputHandler\ for precomputed rotary embeddings (supporting large contexts beyond 65k tokens), \SlidingWindowInputHandler\ for ring-cache attention masks, \PerLayerEmbeddingsInputHandler\ for INT8 embedding gathers, and \PipelinedTokenInputHandler\ for asynchronous Metal buffer rotation in pipelined engines. The \InputLayout\ struct provides shared descriptor analysis to determine query, logits, and prefill policies. State handling is extended with \FixedMTLBufferState\ and \FixedNDArrayState\ for non-truncatable persistent states, and \GrowingNDArrayState\ for dynamically growing KV caches. These changes enable more efficient memory usage and support for advanced model features like sliding window attention and hybrid models.
swift/Sources/CoreAILanguageModels/Handlers · high confidence
New runtime utilities and expanded model structure support in CoreAIShared
The CoreAIShared/Runtime module introduces several new capabilities and structural changes. It adds a backpressured stream to cap unconsumed elements in async producers, a bilinear resampler for image resizing, a bitset for efficient mask operations, and a reader for NumPy .npy files. Model structure detection is expanded to support SAM3 multi-function and video segmenters, which now route to the NeuralEngine and GPU respectively. Asset resolution is unified into a filename-agnostic approach that discovers .aimodel/.aimodelc files, and a new CLI flag allows clearing the Core AI specialization cache. Additionally, NDArray helpers now handle non-contiguous layouts and specific scalar-type filling, while timing utilities and resource management protocols have been moved to this shared location.
swift/Sources/CoreAIShared/Runtime · high confidence
New sampling controls and reproducibility knobs
Users can now configure Min-P filtering, repetition penalty (with a configurable window), and a deterministic seed for reproducible generation. The sampling pipeline applies these in a defined order: repetition penalty, temperature, Min-P, Top-P, and Top-K. These options are available in both the CPU composite sampler and the GPU MPSGraph sampler, with the GPU engine supporting bitmask expansion for constrained generation.
swift/Sources/CoreAILanguageModels/Samplers · high confidence
New video-segmenter CLI tool for SAM 3 video segmentation
A new command-line tool, \video-segmenter\, is introduced to segment and track objects in videos using SAM 3 video models. Users can provide input videos and text prompts to generate segmented outputs, with options to control rendering (mask opacity, bounding boxes, labels), limit processing to a specific number of frames, and export detection data to JSON. The tool also supports diagnostic features such as model cache clearing, warmup passes, and detailed timing logs. Additionally, a parity mode is available to validate the Swift implementation against a reference directory, ensuring consistency with expected detection IDs, scores, and bounding boxes.
swift/Sources/Tools/video-segmenter · high confidence
SAM 3 Video Segmentation export and runtime
Added a new export script and documentation for SAM 3 Video Segmentation, enabling users to convert the model into a bundle format with support for 336, 672, and 1008 input resolutions. The entry point \models/sam3\_video/export.py\ wraps the shared segmentation export logic, specifically overriding dependencies to require \transformers\>=5.5.4\ and \huggingface-hub\>=1.5.0\ to satisfy the model's requirements, and includes a README detailing the pipeline functions, setup, and Swift runtime integration.
_models/sam3\video · high confidence
SAM3 Video Segmentation runtime implementation
Adds the core SAM3 video segmentation engine, including frame preprocessing, a fixed-slot memory bank for object tracking, detection decoding with non-maximum suppression, mask postprocessing, and video overlay rendering.
swift/Sources/CoreAIVideoSegmenter · high confidence
Speech recognizer tool gains parity testing and streaming capabilities
The speech-recognizer tool now includes a --parity-test mode that validates the Swift runtime's output against PyTorch reference traces for Parakeet and Whisper architectures, ensuring numerical accuracy of the mel spectrogram, encoder states, and tokens. It also introduces a --stream flag for incremental transcription via a push-based API, supporting features like real-time pacing, endpointing configuration, and deferred decoding for diagnostic comparison against offline runs.
swift/Sources/Tools/speech-recognizer · high confidence
Support for Vision-Language Models and Per-Layer Embeddings
The language model bundle now supports Vision-Language Models (VLMs) by introducing a \VisionConfig\ field and requiring it for VLM bundles, enabling multimodal inference. Additionally, the system now supports externalized Per-Layer Embeddings (PLE) via a new \SafeTensorsReader\ and \PerLayerEmbeddings\ loader, allowing large INT8 embedding tables to be mapped from sidecar files rather than embedded in the graph. The bundle structure was also renamed from \LanguageBundle\ to \LanguageModelBundle\ to reflect this broader scope, and new configuration options for state handling, prefill chunking, and model-specific runtime overrides (such as sliding windows and RoPE parameters) were added to \LanguageConfig\.
swift/Sources/CoreAILanguageModels/Bundle · high confidence
Support for Vision-Language Models via FoundationModels protocol
Users can now load and run Vision-Language Models (VLMs) using the standard FoundationModels protocol. A new CoreAIVisionLanguageModel adapter allows loading VLM bundles (kind=vlm) and executing multi-modal prompts that include both text and image attachments. The implementation handles image encoding, tokenization, and streaming generation through the CoreAISequentialVLMEngine, enabling capabilities like image understanding within the existing language model session interface.
swift/Sources/CoreAILanguageModels/VLM · high confidence
Behavioural changes
Batched inference and dynamic input sizing for object detection
The object detector now supports processing multiple images in a single batched forward pass via the new \detect(images:)\ API, improving throughput for multi-image scenarios. Additionally, \DetectionParameters\ now accepts \inputHeight\ and \inputWidth\ settings (defaulting to 800x800) to allow dynamic spatial dimensions for models that do not have fixed input shapes, and the \warmup\ method has been updated to accept these parameters to ensure the backend is prepared for the actual batch size and dimensions used in subsequent detection calls.
swift/Sources/CoreAIObjectDetector · high confidence
Benchmark tool adds prefill chunking, cache clearing, and timing metrics
The \llm-benchmark\ tool now supports configurable prefill chunking via \--chunk-size\ and \--chunk-threshold\ flags to optimize inference performance, and includes a new \--clear-coreai-cache\ flag to force re-specialization of Core AI models before loading. Additionally, the benchmark now reports detailed timing for model preparation and warmup phases, including cache hit status, in both console output and JSON reports, providing deeper insight into model load performance.
swift/Sources/Tools/benchmark · high confidence
Configurable image preprocessing strategies and enhanced tokenizer masks
Users can now select specific image preprocessing strategies (stretch, center-crop, or pad) when preparing images for vision-language models, with the system automatically handling layout transposition to the required CHW format. Additionally, the CLI text tokenizer now provides an attention mask alongside token IDs, allowing downstream components to distinguish between actual prompt content and padding tokens.
swift/Sources/CoreAIShared/Image · high confidence
Core AI diffusion components refactor: caching, input validation, and API expansion
The diffusion pipeline components have been updated to improve stability and performance. CoreAIDiffusionModelFunction now caches InferenceFunctions instead of reloading them per call, which prevents GPU memory leaks, and includes a fail-fast check for missing model assets. The public API has been expanded with a new ModelInput enum supporting cached NDArrays and a prepackInput helper for safer buffer handling, alongside strict validation to reject input buffers that do not match the model's expected dimensions. Additionally, CoreAIDenoiser and CoreAILatentCodec now use immutable views for tensor data, and documentation in TextEncoderOutput has been corrected to reflect T5 encoder behavior.
swift/Sources/CoreAIDiffusionPipeline/Components · high confidence
Core AI diffusion runner overhaul: new flags, model support, and latent tuning
The diffusion runner has been refactored to use the new DiffusionBundle API, dropping support for older Stable Diffusion models (1.5, 2.1, 3.5 Medium) and requiring the --model flag for all generation. New capabilities include a --clear-coreai-cache flag to force model re-specialization, a --reference-grid option for controlling token density in img2img, and a --guidance-mode flag to switch between distilled and manual classifier-free guidance. The tool now supports Flux.2 and Sana Sprint 0.6B models via the FlowTransformerPipeline. Additionally, a new latent preview tuning system allows users to collect per-step latents (--tune-preview) and fit RGB coefficients (--tune-fit) to optimize image generation, alongside a hidden --parity-test mode for validating component outputs against numpy fixtures.
swift/Sources/Tools/diffusion-runner · high confidence
Engine architecture refactored for idempotency, implicit prefix caching, and GPU-constrained generation
The inference engine implementation has been restructured to support independent, reusable generation sessions and more efficient processing. Session state (KV cache, token history) is now decoupled from the engine into a caller-owned \GenerationSessionState\, enabling the sequential engine to serve multiple independent conversations and allowing the pipelined engine to perform implicit prefix caching (automatically reusing cached prefixes across turns). A new \ConstrainedGenerationCapable\ protocol adds GPU-accelerated grammar-constrained sampling, and a \ChunkedPrefill\ loop optimizes prefill performance by splitting prompts into chunks, particularly when a dedicated prefill graph is available.
swift/Sources/CoreAILanguageModels/InferenceEngines · high confidence
Enhanced guided generation with rollback, jump-forward, and direct bitmask support
The guided generation engine now supports more robust control over constrained text generation. Users can rollback the grammar state by a specified number of tokens (up to a limit of 64) to correct errors or retry steps, and the system can identify the longest deterministic string from the current state to enable 'jump-forward' optimization, skipping unnecessary token-by-token processing. Additionally, the session can now fill a bitmask directly into a caller-provided buffer, allowing for more efficient integration with downstream components like GPU buffers for token sampling constraints.
swift/Sources/CoreAILanguageModels/GuidedGeneration · high confidence
Enhanced profiling granularity and refined performance metrics API
The profiling system now exposes five new fine-grained engine sub-spans—GatherEmbeddings, MaskBuild, RopeBuild, PLEGather, and LogitsCopy—allowing users to attribute per-token runner-side work more precisely in Instruments traces. Additionally, the PerformanceMetrics API has been refined: token counts are now exposed via a computed totalTokenCount property and recorded through recordPromptTokens and recordGeneratedTokens methods, replacing the previous setter-based approach for better encapsulation.
swift/Sources/CoreAILanguageModels/Profiling · high confidence
Guided generation now supports rollback, jump-forward, and completion detection via XGrammar
The C bridge for the XGrammar library in swift/Sources/lib/CXGrammar has been extended to expose advanced guided-generation capabilities. Users can now roll back the grammar state by a specified number of tokens, retrieve the longest deterministic string for immediate output (jump-forward), and check if the grammar's root rule is fully completed. These features are implemented in the new xgrammar\_c\_bridge.cpp and updated xgrammar\_c\_bridge.h, replacing the previous header location and enabling more robust control over structured text generation.
swift/Sources · high confidence
Image segmenter CLI gains cache clearing and load timing
The image-segmenter tool now includes a --clear-coreai-cache flag that forces re-specialization of the Core AI model by clearing its cached specialization before loading. Additionally, the CLI now reports the model load time and indicates whether a cache hit occurred, helping users monitor performance. The implementation also adds macOS-specific guards for the file-opening helper and refactors bundle loading to use ImageSegmentationBundle instead of ModelBundle.
swift/Sources/Tools/image-segmenter · high confidence
Lazy model loading and agentic reasoning support
Model initialization now defaults to lazy loading, deferring the inference engine load until the first generation request, which reduces startup latency and memory usage; users can opt into eager loading via a new \LoadMode\ parameter. The system now supports agentic chain-of-thought formats, allowing models to emit reasoning content as structured messages rather than inline tags. Additionally, users can configure prefill chunk sizes and thresholds to optimize performance for specific hardware or model architectures.
swift/Sources/CoreAILanguageModels/LanguageModel · high confidence
Logger output redirected to stderr
The CLILogger now writes diagnostic messages to standard error instead of standard output. This ensures that stdout remains clean for programmatic output, such as JSONL results in replay mode, preventing log noise from interfering with structured data consumption.
swift/Sources/CoreAIShared/Logger · high confidence
Model exports default to release mode with optional debug info
The export scripts for CLAP, CLIP, Depth Anything, EDSR, Efficient SAM, PVT, RoBERTa, T5, and Wav2Vec2 now default to RELEASE mode when converting models to Core AI, producing smaller assets by embedding minimal debug information. Users can opt into DEBUG mode via the new --include-debug-info flag to embed full debug information for troubleshooting conversion issues. Additionally, the coreai-core and coreai-torch dependencies have been updated to versions 1.0.0b3 and 0.4.3 respectively, and the Efficient SAM export now defaults to 2 points per query to support box prompts.
(repo-wide) · high confidence
Object detector CLI now supports batch image processing and dynamic model configuration
The object detector tool has been updated to accept multiple input images via repeated --image flags, processing them in a single run and outputting results as a JSON array of per-image detection objects (or rendering separate output images into a directory). Users can now override dynamic model input dimensions using --input-height and --input-width, and can clear the Core AI specialization cache before loading a model with the new --clear-coreai-cache flag to force re-specialization.
swift/Sources/Tools/object-detector · high confidence
Refactored decoding strategies to use eager setup and async sequences
The decoding strategies in CoreAILanguageModels have been refactored to improve performance and API consistency. The \DecodingStrategy\ protocol now returns a typed \AsyncSequence\ (e.g., \VanillaDecodedSequence\, \ConstrainedDecodedSequence\) instead of a generic \AsyncThrowingStream\, with all session creation and tokenization performed eagerly before the sequence is returned. A new \PipelinedConstrainedDecodingStrategy\ has been added to enable GPU-accelerated grammar-constrained decoding by applying bitmasks directly in the MPSGraph sampler, avoiding CPU logit transfers. Additionally, \StopSequences\ now supports additional EOS token IDs from tokenizer configs, and log probability calculations utilize a new \LogProbabilities\ struct for better numerical stability.
swift/Sources/CoreAILanguageModels/DecodingStrategies · high confidence
SAM3 export now supports iOS-optimized lite mode alongside full HF export
The SAM3 export tool now provides a default 'lite' export path targeting iOS, restructuring the model into a BC1S layout with palettized encoders (4-bit image, 6-bit text) and splitting it into three independently optimized functions (image\_encode, text\_encode, detect) at a 336x336 resolution. Users can also opt for a 'full' export via the new --full flag to retain the unmodified Hugging Face model structure. The CLI has been updated with new flags for controlling image size, palettization bits, and group size, and the default output path has changed to reflect the new lite bundle naming convention.
models/sam3 · high confidence
Scheduler refactoring: removal of legacy schedulers and expansion of DiscreteFlowScheduler
The scheduler module has been significantly restructured. The PNDMScheduler and DPMSolverMultistepScheduler implementations have been removed, and the SchedulerType enum now only exposes discrete flow matching. The DiscreteFlowScheduler has been expanded to support Flux, Wan, and Sana Sprint models, introducing a new initializer for explicit sigma schedules (e.g., TrigFlow), a stochastic sampling step for consistency-distilled models, and a public currentSigma property to expose the denoised x0 estimate for live previews. Additionally, the PredictionType enum has been moved to the top-level Scheduler.swift file.
swift/Sources/CoreAIDiffusionPipeline/Schedulers · high confidence
Streaming tool-call detection and grammar library cleanup
The CoreAILanguageModels module now includes a streaming parser (ToolCallParser) that detects tool-call blocks in the model's token stream, supporting JSON, ATEM, and Qwen3-Coder XML formats. A helper function (lastSafeIndex) ensures streaming markers are not split across deltas. Additionally, unused model shape configuration files and the bundled CXGrammar (XGrammar) framework headers have been removed from this location.
swift/Sources/CoreAILanguageModels · high confidence
Support for multi-function SAM3 lite models and smoother segmentation output
The image segmentation engine now automatically detects and supports two model asset shapes: the existing single-function models (like baseline SAM3 and EfficientSAM) and new multi-function models (produced by the SAM3 lite export) that use separate encode and detect graphs. This allows users to run SAM3 lite models without code changes. Additionally, the post-processing step now uses bilinear upsampling instead of nearest-neighbor for mask and probability grids, which eliminates staircase artifacts and produces smoother segmentation boundaries, particularly for the lower-resolution SAM3 lite outputs.
swift/Sources/CoreAIImageSegmenter · high confidence
Support for new model kinds and improved bundle validation
The model bundle system now recognizes \videoDiffusion\, \videoSegmenter\, and \speechRecognizer\ kinds, enabling the runtime to handle video diffusion (Wan), SAM 3 video segmentation, and speech recognition models. The \LanguageBundle\ type has been renamed to \LanguageModelBundle\ to reflect its scope. Additionally, bundle loading now validates that the input is a directory rather than a direct model asset file (\.aimodel\/\.aimodelc\), providing a clear error message if a user points the runner at a single asset file. The asset verification method was also renamed to \verifyAssetsExisting\ for clarity, and error messages were updated to provide more specific guidance on missing assets and bundle structure.
swift/Sources/CoreAIShared/Bundle · high confidence
Vectorized log-probability computation and enhanced generation API
Text generation now includes a new LogProbabilities struct that uses the Accelerate framework for vectorized log-softmax calculations, significantly improving performance for per-token log probability and top-K alternative extraction. The TextGenerator API has been updated to return token IDs alongside text and logits in generateWithLogits, and a new evaluateRawTokens method allows direct evaluation of pre-tokenized token sequences. Additionally, the decoding strategy is now accessed asynchronously via try await to support serialization of back-to-back turns.
swift/Sources/CoreAILanguageModels/TextGeneration · high confidence
Whisper export defaults to release mode with optional debug info and dynamic shapes
The Whisper model export now defaults to release mode, producing smaller assets by embedding minimal debug information; users can opt into full debug info via the new --include-debug-info flag. Additionally, the export process now traces the model with dynamic shapes (specifically for decoder input sequence length), which improves flexibility for variable-length inputs during inference.
models/whisper · high confidence
Yolo export defaults to release mode with optional debug info and updated Core AI dependencies
The Yolo model export process now defaults to release mode, resulting in smaller exported assets, while adding a --include-debug-info flag to opt-in to debug information embedding. This change also updates the coreai-core and coreai-torch dependencies to versions 1.0.0b3 and 0.4.3 respectively. Additionally, the documentation reflects new capabilities for batched detection and dynamic shape handling in both iOS/macOS applications and the Mac command-line tool, including support for explicit input dimensions and kernel warmup.
models/yolo · high confidence
llm-runner adds video input, MinP sampling, repetition penalties, and VLM image options
The llm-runner tool now supports vision-language models with new command-line flags for video input (--video, --video-frames, --video-sampling) and configurable image preprocessing (--image, --image-strategy, --image-info). It introduces MinP sampling (--min-p) alongside existing Top-P, and adds repetition penalty controls (--repetition-penalty, --repetition-penalty-window) with validation against constrained generation. Prefill chunking is refined with a separate --chunk-threshold option and a default of 4096 tokens on macOS with \>=64GB memory. A new --clear-coreai-cache flag forces re-specialization of the Core AI model cache, and the internal bundle type has been renamed from LanguageBundle to LanguageModelBundle.
swift/Sources/Tools/llm-runner · high confidence
Fixes
Fixes token count mismatch error in logits output handling
The \LogitsWriter\ now accepts generated token IDs directly instead of decoding generated text to derive them, which resolves a token count mismatch error that occurred when the encoded token count differed from the logits count. This change ensures that the validation logic correctly compares the number of provided token IDs against the number of logits vectors, preventing crashes or invalid states during logits saving or console printing.
swift/Sources/CoreAILanguageModels/Output · high confidence
Test coverage
Added comprehensive test coverage for CoreAISpeech audio processing and configuration; Added test coverage for CoreAIShared utilities; Added test coverage for CoreAIVideoSegmenter components; Added test coverage for diffusion pipeline configuration, token math, and input validation; Added tests for ConstrainedGenerationSession rollback, jump-forward, and bitmask filling; Added tests for ImageSegmentationBundle and updated segmentation engine tests; Added tests for ObjectDetector batch planning logic; Added tests for language model parsing, engine lifecycle, and configuration; Added tests for video input, reading, and writing in CoreAIShared; Added unit tests for CoreAILMCommon API types and server utilities; Removed unused XCTest plan file.
Dependencies
New video, speech, and LLM server capabilities with dependency updates
This release introduces new Swift libraries and tools for video diffusion (CoreAIVideoDiffusionPipeline), video segmentation (CoreAIVideoSegmenter), and speech recognition (CoreAISpeech), alongside a new llm-server executable for OpenAI-compatible LLM serving. The CXGrammar dependency is now a source-based target using XGrammar v0.2.2 instead of a binary framework. Python dependencies are updated, including transformers to v5.12.1, huggingface-hub to \>=1.5.0, and tokenizers capped at \<0.23, with new CLI scripts for VLM export and a flaky test marker added.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Score
- CAI 63 → 64 (+1.2)
- Rubric changed (rubric-2026.09.11 → rubric-2026.09.18) — scores are not directly comparable.
Lenses
- Code Health 85 → 85 (-0.6)
- Architecture 100 → 93 (-6.5)
- Maturity 58 → 58 (-0.0)
- Readiness 51 → 54 (+3.5)
- Security 79 → 84 (+4.6)
- Performance 83 (new)
Resolved (78)
- ClassTooLong: DiffusionRunner (swift/Sources/Tools/diffusion-runner/DiffusionRunnerMain.swift)
- ClassTooLong: Flux2Pipeline (swift/Sources/CoreAIDiffusionPipeline/Pipelines/Flux2Pipeline.swift)
- DiffusionRunner.runSD3ComponentParity (cognitive 20) (swift/Sources/Tools/diffusion-runner/DiffusionRunnerMain.swift)
- DiffusionRunner.runSD3ComponentParity (cyclomatic 19) (swift/Sources/Tools/diffusion-runner/DiffusionRunnerMain.swift)
- Documentation: no architecture or design documentation (models/muse_glimmer/README.md)
- Documentation: no installation or build instructions (README.md)
- Documentation: no licence statement (README.md)
- Documentation: no project overview (README.md)
- Documentation: no usage examples (README.md)
- Duplicated block (10 lines × 2) (swift/Sources/CoreAIDiffusionPipeline/Schedulers/DPMSolverMultistepScheduler.swift)
- Duplicated block (10 lines × 2) (swift/Sources/CoreAILanguageModels/InferenceEngines/CoreAISequentialEngine.swift)
- Duplicated block (10 lines × 2) (swift/Sources/Tools/diffusion-runner/DiffusionRunnerMain.swift)
- Duplicated block (10–11 lines × 2) (swift/Sources/CoreAIDiffusionPipeline/Pipelines/Flux2Pipeline.swift)
- Duplicated block (11 lines × 2) (swift/Sources/CoreAIDiffusionPipeline/Pipelines/Flux2Pipeline.swift)
- Duplicated block (11 lines × 2) (swift/Sources/CoreAILanguageModels/Bundle/LanguageConfig.swift)
- Duplicated block (11 lines × 3) (swift/Sources/Tools/diffusion-runner/DiffusionRunnerMain.swift)
- Duplicated block (12 lines × 3) (swift/Sources/Tools/diffusion-runner/DiffusionRunnerMain.swift)
- Duplicated block (12 lines × 5) (swift/Sources/CoreAIDiffusionPipeline/Pipelines/Flux2Pipeline.swift)
- Duplicated block (12 lines × 6) (swift/Sources/CoreAIDiffusionPipeline/Pipelines/PipelineConfiguration.swift)
- Duplicated block (13 lines × 2) (swift/Sources/CoreAIDiffusionPipeline/Pipelines/Flux2Pipeline.swift)
- …and 58 more
New (216)
- Associator.associate (cognitive 27) (swift/Sources/CoreAIVideoSegmenter/Tracking/Associator.swift)
- Associator.associate (cyclomatic 17) (swift/Sources/CoreAIVideoSegmenter/Tracking/Associator.swift)
- Change coupling: CoreAISequentialVLMEngine.swift ↔ LLMRunnerMain.swift (swift/Sources/CoreAILanguageModels/InferenceEngines/CoreAISequentialVLMEngine.swift)
- Change coupling: LLMRunnerMain.swift ↔ LLMServerMain.swift (swift/Sources/Tools/llm-runner/LLMRunnerMain.swift)
- ChatHandler.swift.runChatCompletion (cognitive 43) (swift/Sources/Tools/llm-server/ChatHandler.swift)
- ChatHandler.swift.runChatCompletion (cyclomatic 25) (swift/Sources/Tools/llm-server/ChatHandler.swift)
- ChatHandler.swift.runStreamingLoop (cognitive 40) (swift/Sources/Tools/llm-server/ChatHandler.swift)
- ChatHandler.swift.runStreamingLoop (cyclomatic 25) (swift/Sources/Tools/llm-server/ChatHandler.swift)
- ChatHandler.swift.tokenizeMessages (cognitive 20) (swift/Sources/Tools/llm-server/ChatHandler.swift)
- ClassTooLong: StaticShapeEngine (swift/Sources/CoreAILanguageModels/InferenceEngines/CoreAIStaticShapeEngine.swift)
- ConnectedComponents.areas (cognitive 35) (swift/Sources/CoreAIVideoSegmenter/Postprocessing/ConnectedComponents.swift)
- ConnectedComponents.areas (cyclomatic 18) (swift/Sources/CoreAIVideoSegmenter/Postprocessing/ConnectedComponents.swift)
- ConstrainedGenerationSession.swift._applyBitmask (cognitive 21) (swift/Sources/CoreAILanguageModels/GuidedGeneration/ConstrainedGenerationSession.swift)
- Coverage not measured — Swift suite
- Critical CVE: [GHSA redacted] (uv.lock)
- Duplicated block (10 lines × 2) (models/depth-anything/export.py)
- Duplicated block (10 lines × 2) (models/t5/export.py)
- Duplicated block (10 lines × 2) (python/src/coreai_models/models/macos/muse_glimmer_drafter_dflash.py)
- Duplicated block (10 lines × 2) (python/src/coreai_models/primitives/macos/cache.py)
- Duplicated block (10 lines × 2) (swift/Sources/CoreAIDiffusionPipeline/Components/CoreAIDiffusionModelFunction.swift)
- …and 196 more
Changes since last survey
- 31 commits — 29 feature/other, 2 fixes
By area
- swift/Sources — 22 commits
- python/src — 6 commits
- models/phi — 1 commit
- python/tests — 1 commit
- swift/Tests — 1 commit
Notable commits
- fix: Fix BFloat16 KV-state crash in reset and cache growth (#268)
- fix: Fix SSMState.update_states slice_update end index missing trailing dim (#275) (#283)
- change: Add --replay mode to the LLM server (#266)
- change: Add DiffusionGemma-26B-A4B model Core AI export (#261)
- change: Add reasoning_effort control to the LLM server (#263)
- change: Add reproducible generation: seed and system_fingerprint (#265)
- change: Broaden tool-calling to Phi and Qwen3-Coder dialects (#271)
- change: Consolidate VLM sequential engine: route construction through EngineFactory (C1+C2) (#262)
- change: Consolidate VLM sequential engine: shared iterator helpers, tidy (B3+B4) (#251)
- change: Consolidate model-bundle and diffusion asset resolution (#304)
- change: CoreAISequentialEngine: idempotent generation via external session state (#258)
- change: Default prefill chunk to 4096 only at >=64GB memory (#274)
- change: Exclude RMSNormGated from 4bit weight quantization (#277) (#281)
- change: Gemma 4 E2B / E4B iOS runner upgrades (#302)
- change: Honor ContextOptions.reasoningLevel in the language model executor (#267)
- change: Image and Video Segmentation Bundle Cleanup (#308)
- change: Migrate Diffusion Quantization to Pre-Export (#286)
- change: New Model: Sana Sprint 0.6B (#289)
- change: Preview denoised x0 estimate for FLUX.2 latent preview (#254)
- change: Recover exported turn-end stop tokens from tokenizer.json (Gemma, Phi) (#264)
- …and 11 more
Architecture
- Unchanged — 0 containers · 1 contexts · 0 edges
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
apple/coreai-models was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 1 October 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 7ec0a710e0de883d167b6d5b7e439f0442d7f5ad — the exact code this score is about.
- Scored under rubric-2026.09.18 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-e569280dd5e2.