Skip to content
CAI
Software that uses CAICheck a score

kellyvv/PhoneClaw

38.9

Weak · 1 October 2026

66.2k

lines of production code

Swift

with C, JavaScript

2

measurements over time

CAI band scale
CAI trend line
CAI lens gauges

What this system is

PhoneClaw is an open-source, on-device AI agent for iOS and macOS that runs local large language models to ensure user privacy. It provides a multimodal chat interface and a real-time 'Live' voice mode, supporting text, image, and audio inputs through a unified inference engine backed by LiteRT and MLX. The system features a modular skill architecture with native tool handlers for device services like calendar and health, alongside a remote inference gateway for Mac-based offloading.

Features

Add FastCluster wrapper and Parakeet ASR components to FluidAudio package

This change introduces the C++ FastCluster wrapper (including headers, internal algorithms, and module map) to enable centroid linkage hierarchical clustering for the speaker diarization pipeline, and adds the Swift source files for the Parakeet Automatic Speech Recognition subsystem, including configuration types, thread-safe audio buffering, CoreML model loading for multiple languages, ARPA n-gram language model parsing, and CTC decoding (greedy and beam search with LM rescoring).

Packages/FluidAudio · high confidence

Add Gemma 4 and MiniCPM-V 4.6 models to the catalog

The app now includes predefined descriptors for Gemma 4 (E2B and E4B variants) and MiniCPM-V 4.6, making them available for download and use. Gemma 4 models are configured for LiteRT with specific context budgets and capabilities like live mode (E2B) or structured planning (E4B). MiniCPM-V 4.6 is added as a multimodal model using the GGUF format with a companion multimodal projector file, supporting vision and live modes but excluding audio and thinking capabilities. These models are now listed in the available models catalog for users to select.

LLM/Models · high confidence

Add iOS 27 Foundation Models inference backend

Users on iOS 27 and macOS 27 can now use Apple's native Foundation Models for text-based inference. This change introduces a new backend that registers the 'Apple Foundation Models' (Gemma 4 family) as an available system model, handling model loading, session management, and text generation via the \FoundationModelsInferenceService\. The implementation is text-only, explicitly ignoring image and audio inputs in multimodal or live generation paths, and supports features like live mode, structured planning, and persistent sessions.

LLM/Backends/FoundationModels · high confidence

Added Three.js core library and license to web assets

The application now includes the Three.js 3D rendering library as a bundled web asset. This change adds the \three.core.min.js\ build file and the corresponding MIT license (\three/LICENSE\) to the \App/OrbWebAssets/three\ directory, enabling 3D graphics capabilities within the web client.

App/OrbWebAssets · high confidence

Gemma 4 audio support and memory optimizations

The Gemma 4 model now supports audio input, including a new feature extractor, audio encoder, and processor integration, allowing the model to process audio alongside text and images. To improve stability on memory-constrained devices, the model now uses 4-bit quantized KV caches for full attention layers (reducing peak KV memory by \~36%) and implements chunked prefill for long text prompts to avoid out-of-memory errors. Additionally, the vision model's masked scatter operation has been made more robust to handle mismatched source and position counts.

LLM/MLX/Gemma4 · high confidence

Initial release of MLX Swift package with CI and build infrastructure

This change introduces the MLX Swift package, providing a Swift API for the MLX array framework. It includes the core Swift source code for MLX and related libraries (MLXNN, MLXOptimizers, etc.), a CMake build system for macOS and Linux (including CUDA support), and a complete GitHub Actions CI pipeline for linting, testing, and releasing XCFrameworks. The package also adds standard repository configuration files such as issue templates, pull request templates, and contributor guidelines.

Packages/mlx-swift, python · high confidence

Initial release of PhoneClaw on-device AI agent

Introduces the PhoneClaw application, a local AI assistant that runs entirely on-device to protect user privacy. The app features a chat interface with streaming responses, multimodal support for image analysis, and a file-driven skill system that allows users to enable or disable specific capabilities. It includes a configuration UI for adjusting model parameters (such as temperature and token limits) and managing system prompts, backed by an MLX-based inference engine for efficient local processing.

PhoneClaw, PhoneClawMac, PhoneClawMac/Sources/PhoneClawMac · high confidence

Introduce LAN-based remote inference via Mac gateway

Adds a new remote inference backend that allows the iOS app to offload LLM generation to a Mac on the same local network. This includes a LAN discovery service using Bonjour to find Mac gateways, a binding system that persists Mac identity and secrets securely in the Keychain for automatic reconnection, and a Mac gateway implementation that acts as an OpenAI-compatible reverse proxy for local models like Ollama. The remote inference service integrates with the existing agent engine, translating local prompts into chat completions and streaming SSE responses back to the device.

LLM/Backends/Remote · high confidence

Introduce LiteRT backend with KV cache reuse and multimodal support

Adds a new LiteRT inference backend that enables persistent sessions for KV cache reuse, significantly reducing time-to-first-token for chat and Live modes. The backend supports both CPU and GPU execution paths, handles multimodal inputs (text, image, audio) via a dedicated Conversation API, and includes a diagnostic test suite to verify audio streaming behavior against non-streaming outputs.

LLM/Backends/LiteRT · high confidence

Introduce LiveLand real-time voice mode with Dynamic Island and widget support

This change adds the LiveLand feature, a real-time voice interaction mode that integrates with iOS system surfaces. It introduces a Live Activity to display the conversation state (listening, processing, skill execution, results) in the Dynamic Island and on the lock screen, complete with a new activity bridge and attributes. The feature includes a dedicated audio input engine for microphone capture, a Voice Activity Detection (VAD) service using Silero, and a turn controller to manage speech boundaries. Users can launch LiveLand via new App Intents, home screen widgets, or URL schemes, and the system provides haptic feedback and background task continuation to maintain the session.

LiveLand · high confidence

Introduce Mac Gateway app with LAN discovery and icon generation

Adds the PhoneClaw Gateway macOS application, providing a local LLM inference gateway that broadcasts via Bonjour on the local network. The app includes a SwiftUI dashboard for managing providers and paired devices, along with a menu bar extra. It requires Local Network permission to function and ships with a new build script and Python asset tools to generate the macOS app icon from a source render.

MacGateway · high confidence

Introduce MiniCPM-V backend with multimodal and live video support

Added a new MiniCPMVBackend implementation that integrates the OpenBMB mtmd-ios C API, enabling support for the MiniCPM-V model family. This backend manages a three-file model bundle (LLM, multimodal projector, and optional CoreML vision tower) and introduces live video inference capabilities via generateLive, which utilizes KV cache reuse for efficient streaming. The implementation includes a transaction gate to serialize live requests and prevent concurrency crashes, along with vision warmup and idempotent stop generation to improve stability and response quality.

LLM/Backends/MiniCPMV · high confidence

Introduce SkillKit draft generator for PhoneClaw skills

A new Python-based tool (SkillKit) has been added to generate PhoneClaw skill skeletons from a JSON manifest. It validates the Skill Contract V2 schema and renders localized SKILL files (Chinese, English, Japanese) along with a smoke-test scenario YAML, helping developers create consistent skill structures without manual boilerplate.

Tools/SkillKit · high confidence

Introduce dedicated LIVE voice model download and installation infrastructure

This change adds a new, independent installation subsystem for LIVE voice assets (ASR, TTS, and VAD) in the \LLM/MLX/Installation\ directory. It introduces \LiveModelStore\ and \LiveModelDownloader\ to manage the on-device download of these specific models from ModelScope and HuggingFace, utilizing \LiveDownloadPlanner\ to handle file listing via baked manifests or tree APIs. The system includes \LiveModelInstallFinalizer\ for atomic staging and validation, ensuring that required files are present before finalizing the installation. This infrastructure is distinct from the existing LLM model downloader, providing a dedicated path for voice capabilities without modifying the core MLX local LLM service.

LLM/MLX/Installation · high confidence

Introduces LLM Core architecture with safe generation lifecycle and unified inference protocol

This change establishes the LLM/Core module, introducing a new \InferenceService\ protocol that unifies text, multimodal, and Live mode generation behind a single interface, and a \ModelRuntimeCoordinator\ that manages model loading, backend switching, and session state. It adds \GenerationTransaction\ to enforce a safe cancellation flow—ensuring the underlying inference stream fully terminates before KV cache is reset to prevent state corruption—and \LiteRTBootstrap\ to preload GPU accelerators at process startup. The update also includes \InstallState\ for per-model installation tracking, \PromptTokenEstimator\ for context budgeting, and \WebFreshness\ for search result recency scoring.

LLM/Core · high confidence

Introduces MiniCPM-V multimodal inference engine with C-interop bridge

Adds the CMTMDBridge layer to safely expose the OpenBMB mtmd-ios C++ API to Swift without requiring Swift/C++ interop, preventing build conflicts with other modules. This enables the MTMDEngine to initialize and run MiniCPM-V models for image and video frame processing, supporting features like KV cache cleaning, image slicing configuration, and streaming token generation with proper UTF-8 buffer handling.

LocalPackages/PhoneClawEngine/Sources/MTMDEngine · high confidence

Introduces VAD service for real-time voice activity detection in Live mode

Adds a new VAD service that wraps the Silero VAD model (via FluidAudio) to detect speech activity in real-time. The service processes 16kHz audio chunks serially to ensure deterministic callback ordering, exposing events for speech start/end and probability updates. It supports loading a pre-downloaded CoreML model from the unified Live model download path, with a fallback to auto-download, and integrates with the shared audio engine via a handler to enable barge-in logic in the Live mode engine.

Live/VAD · high confidence

Introduces collapsible 'Thinking' cards and a porcelain light theme

The chat interface now displays the model's internal reasoning in a dedicated, collapsible 'Thinking' card, separated from the final response by parsing new \\[\[PHONECLAW\_THINK\]\]\ markers in the data layer. This is accompanied by a complete visual overhaul to a 'porcelain' light theme (champagne background, amber copper accents) and a new responsive scaling system that adjusts spacing and component sizes for large-screen devices without simply enlarging UI elements.

UI · high confidence

Launch of the PhoneClaw GitHub Pages documentation site

The repository now includes a complete, static documentation site hosted on GitHub Pages, providing a public-facing web presence for the PhoneClaw project. This site features a cohesive visual identity with a custom CSS theme (using the Space Grotesk font and a dark/light color palette) and includes dedicated pages for the framework overview, on-device model benchmarks (detailing memory footprints for Gemma 4 and MiniCPM-V), the FAQ, LiveLand interaction details, and Mac Gateway remote inference. It also supports internationalization with landing pages in Simplified Chinese and Japanese, and is optimized for search engines with comprehensive SEO metadata, Open Graph tags, and structured data (JSON-LD).

site · high confidence

Localized on-device TTS with language-specific backends

The Live TTS service now uses distinct neural models based on the device language: Chinese uses the 'keqing' model via sherpa-onnx, English uses the 'Piper libritts\_r-medium' model, and Japanese uses the 'Piper Plus' backend. The system shares a single audio engine with VAD to enable acoustic echo cancellation, and while non-Japanese locales fall back to the system TTS if the neural model is unavailable, Japanese strictly requires the on-device neural backend.

Live/TTS · high confidence

Localized system prompts and per-turn persona handling for Live Voice mode

The Live Voice mode now supports localized system prompts and user interactions for Chinese, English, and Japanese. A new \LiveLocale\ configuration provides distinct persona names (e.g., 'PhoneClaw' for English, '手机龙虾' for Chinese) and tailored system instructions for each language. The per-turn persona reminder, previously hardcoded in Chinese, is now language-aware: it is applied in Chinese to prevent identity drift, but removed in English to avoid the model treating the reminder as a stage direction that forces unnatural sentence openers. Additionally, camera state handling is localized, with specific markers for 'camera off' and vision task hints adapted to each locale's natural language patterns.

LLM/LiveVoice · high confidence

Native audio input support and prompt localization foundation

The LLM engine now accepts audio inputs alongside images and text, enabling voice-based interactions with local models. To support this and improve multilingual accuracy, a new PromptLocale system centralizes all model-facing instructions (system prompts, thinking mode cues, image/audio context markers, and time anchors) for Chinese, English, and Japanese, ensuring the AI responds in the user's language and handles multimodal context correctly.

LLM · high confidence

New LiteRT model store and secure zip extraction for MiniCPM-V 4.6

The app now uses a dedicated \LiteRTModelStore\ to manage \.litertlm\ model installations, ensuring thread-safe state updates to prevent crashes during UI rendering. It includes a one-time cleanup of obsolete CoreML/ANE companion files for MiniCPM-V 4.6 to reclaim disk space and a new \ZipExtractor\ to safely unpack model archives without external dependencies, protecting against path traversal vulnerabilities.

LLM/Installation · high confidence

New and updated local frameworks for LiteRT LM, Metal acceleration, and llama integration

The engine now ships with several new or updated XCFrameworks to support advanced inference capabilities. A new GemmaModelConstraintProvider framework is added to enable constrained decoding for Gemma models. The LiteRT LM engine is upgraded to the main HEAD with a new CLiteRTLM framework exposing a C API for session and conversation management, alongside LiteRtMetalAccelerator and LiteRtTopKMetalSampler frameworks to leverage Apple's Metal GPU acceleration. Additionally, the llama.xcframework is vendored to support MiniCPM-V integration, providing the underlying GGML backend and CPU/BLAS execution layers.

LocalPackages/PhoneClawEngine/Frameworks · high confidence

New chat engine data models and session persistence layer

The Agent/Engine module introduces a new data model for chat interactions, including ChatMessage (supporting image and audio attachments), ChatImageAttachment (with platform-specific image processing), and ChatAudioAttachment (with waveform generation). A new ChatSessionStore class handles session persistence, loading, and saving with debounced writes. The engine also adds a new GuardedSkillRouter for intent-based skill routing, an IOS27FoundationSkillRouter for Foundation model-based routing, and an ImageFollowUp system for managing multimodal conversation context. Additionally, a ConversationMemoryPolicy manages context window budgeting with legacy and hotfix strategies, and an OutputCleaner handles response sanitization and safety truncation.

Agent/Engine · high confidence

New iOS 27 Core AI research experiment and probe app

Added a new isolated experiment package (Experiments/iOS27CoreAI) and a standalone iOS probe app (Experiments/iOS27CoreAIProbeApp) to research iOS 27 Core AI and Foundation Models capabilities without affecting the main PhoneClaw runtime. The experiment package introduces a planning model abstraction with a heuristic baseline service for skill routing, a FoundationModels-based service for structured routing, and a CoreAI probe to benchmark model loading and function discovery. The probe app provides a UI for smoke-testing framework availability, system language model responses, and routing logic on iOS 27 devices.

Experiments · high confidence

New live camera service with periodic frame capture

A new LiveCameraService has been introduced in the Live/Camera module to manage AVFoundation camera sessions. It provides asynchronous start/stop capabilities, handles camera permissions, and implements a periodic frame-capture mechanism (every 3 seconds) that produces CIImage snapshots. The service ensures thread safety for session operations and snapshot access, and includes logic to handle cancellation during asynchronous permission requests.

Live/Camera · high confidence

New native tool handlers for Calendar, Contacts, Health, Reminders, Clipboard, and Web search

The application now includes dedicated handlers for six core system capabilities, enabling direct interaction with device data and services. Users can create and query calendar events, manage contacts (create, update, search, delete), and read health metrics (steps, heart rate, sleep, etc.) via HealthKit. Reminders can be created with natural language time parsing, clipboard contents can be read and written, and web searches are performed with a structured retrieval pipeline that includes result ranking, freshness handling, and a circuit-breaker for search providers. These handlers are registered into the tool registry and support bilingual (Chinese/English/Japanese) descriptions and parameter parsing.

Tools/Handlers · high confidence

New permission management and tool execution infrastructure

The Tools module now includes a centralized AppPermissions system that defines and checks access states for microphone, camera, calendar, reminders, contacts, and HealthKit, ensuring skills only run when authorized. A new ToolRegistry manages tool definitions with unified validation logic for required parameters and aliases, while Helpers provide robust JSON payload formatting and flexible date parsing using Apple's NSDataDetector to handle natural language time expressions. Additionally, SystemStores centralizes access to EventKit and Contacts databases to prevent redundant instance creation across the tool handlers.

Tools · high confidence

New shared audio engine and real-time frequency analyser for Live mode

Live mode now uses a unified AVAudioEngine to handle both microphone input and TTS output, enabling hardware-accelerated acoustic echo cancellation (AEC) via iOS voice processing. The new LiveAudioIO component manages audio routing, sample-rate conversion to 16kHz for VAD, and idle detection, while the new OrbAudioAnalyser provides real-time 16-bin frequency data using vDSP FFT for the visual orb. This change ensures that the TTS output is properly cancelled from the microphone input, improving audio clarity during live interactions.

Live/Audio · high confidence

New structured logging, language service, and diagnostics foundation

This update introduces three new core components in the Shared module. First, a typed localization system (L10n) and LanguageService replace scattered string literals and locale checks, enabling immediate UI language switching (Chinese, English, Japanese) with auto-detection from system settings. Second, a structured logging framework (PCLog) standardizes runtime diagnostics into categorized, machine-parseable logs (Model, Turn, Perf, Warn) while suppressing noisy LiteRT runtime output. Third, a diagnostics export feature allows users to generate a JSON bundle of device and runtime metadata for bug reporting, excluding user content.

Shared · high confidence

Open-sourced under Apache 2.0 with new community guidelines and repository structure

The project has switched its license from MIT to Apache 2.0, adding the standard \LICENSE\ file. To support this open-source transition, the repository now includes a \CODE\_OF\_CONDUCT.md\ (based on the Contributor Covenant) and a \CONTRIBUTING.md\ guide detailing build requirements and issue reporting. The repository structure has been formalized with a new \PROJECT\_STRUCTURE.md\ outlining module boundaries (e.g., \Agent/\, \Skills/\, \MacGateway/\), and developer tooling has been updated with a \.gitattributes\ file to correctly handle language statistics for vendored packages.

(repo-wide) · high confidence

PhoneClawEngine local package with GPU accelerator preloading and LiteRT integration

The PhoneClawEngine is now a local Swift package containing the core LiteRT-LM inference engine. This update introduces a public \LiteRTRuntime.preloadGpuAccelerator()\ API that synchronously loads the Metal GPU accelerator framework at app startup to ensure the GPU backend is registered before any model loading occurs. The engine class now exposes a \Status\ enum for tracking model state and supports text, vision, and audio inference via a unified Conversation API, with specific optimizations to skip loading vision encoder weights for text-only sessions.

LocalPackages/PhoneClawEngine/Sources/PhoneClawEngine · high confidence

Resumable background downloads for LLM assets

The LLM installation system now supports resumable downloads that persist progress across app suspensions and background transfers. Users benefit from reliable downloads that survive interruptions, automatic recovery of partial files, and background execution via URLSession to conserve battery and maintain progress even when the app is in the background. The system also validates file integrity, handles source switching on failure, and cleans up orphaned workspace data.

LLM/Installation/Download · high confidence

Shared Audio: Promotes Live and Chat ASR components to shared module

The Shared/Audio directory now contains the core implementation files for Automatic Speech Recognition (ASR) and output sanitization, including ASRBackend.swift, ASRService.swift, OutputSanitizer.swift, and SherpaOnnx.swift. This change centralizes the logic for selecting ASR backends (Sherpa Onnx for Chinese/English, Sherpa Offline for Japanese, and WhisperKit as a fallback), manages model initialization and lifecycle, and provides text cleaning for both Chat UI and Live Voice modes. These files are now available for use by the CLI harness and other parts of the application, ensuring consistent audio processing behavior across platforms.

Shared/Audio · high confidence

Unified inference backend dispatcher for multi-model support

The application now uses a BackendDispatcher to route inference requests to different backend implementations based on the loaded model. This allows seamless switching between LiteRT (for Gemma 4), MiniCPM-V 4.6 (for GGUF models with ANE), remote Mac inference via OpenAI-compatible gateways, and Apple Foundation Models, while maintaining a stable interface for the rest of the application.

LLM/Backends · high confidence

Behavioural changes

Agent engine restructured with new audio capture and configurable model settings

The Agent module has been refactored to introduce a new \AudioCaptureService\ that records microphone input as M4A files and decodes them for model processing, replacing previous PCM-based approaches. The \AgentEngine\ core has been reorganized into a modular structure with explicit dependencies on \ModelRuntimeCoordinator\ and \ChatSessionStore\, and now exposes observable state for UI binding. Model configuration is now handled via a dedicated \ModelConfig\ class, which defaults \maxTokens\ to 2048 and supports backend selection (CPU/GPU) and speculative decoding toggles. Additionally, a new \HotfixFeatureFlags\ system allows runtime control of prompt pipelines, tool result canonicalization, and skill routing via environment variables or user defaults.

Agent · high confidence

App initialization now ensures GPU acceleration and background download continuity

The app entry point now explicitly initializes the LiteRT runtime bootstrap before any other operations, ensuring the GPU Metal accelerator is registered correctly for model inference. Additionally, the app registers an AppDelegate to handle background URLSession transfer completion events, enabling reliable background downloads.

App · high confidence

Archive upstream LiteRT-LM patches for future engine rebuilds

Added a new \patches\ directory to the PhoneClawEngine local package to archive the local modifications applied to the upstream LiteRT-LM library. This includes the patched \sampler\_factory.cc\ (which implements a Mach-O symbol walker to resolve hidden C ABI symbols in the iOS Metal sampler dylib), the modified \c/BUILD\ file (which adds a build target for the engine dylib), the combined \litert-lm.patch\, and the \package-xcframework.sh\ script. This archive serves as a bus-factor safeguard, ensuring that the specific patches required to rebuild the \LiteRTLM.xcframework\ and its plugins (such as \LiteRtTopKMetalSampler\ and \LiteRtMetalAccelerator\) are preserved for future engine upgrades or rebuilds.

LocalPackages/PhoneClawEngine/patches · high confidence

Cross-turn KV cache reuse for faster text-only responses

The MLX inference service now reuses the Key/Value (KV) cache across conversation turns for text-only interactions, significantly reducing time-to-first-token (TTFT) by avoiding redundant prefilling of common prompt prefixes. This optimization is enabled by default for the Gemma 4 E4B model but remains disabled for the Gemma 4 E2B model due to observed generation issues. The feature is controlled by a \kvReuseEnabled\ flag and includes benchmarking tools to track cache hit rates and token consumption.

LLM/MLX · high confidence

Dynamic memory-aware generation and iOS tokenizer compatibility

The MLX core now uses a two-layer memory safety architecture: a linear budget formula for initial token estimates and a real-time headroom floor check during generation to prevent out-of-memory crashes. This replaces the previous total-sequence budget logic, which incorrectly reduced output limits based on prompt length. Additionally, the tokenizer loader now supports both SwiftPM and iOS Xcode builds by conditionally calling the correct API for the underlying tokenizer library, ensuring compatibility across build environments.

LLM/MLX/Core · high confidence

InferenceKit package scaffolding and Gemma-focused model registry

The Packages/InferenceKit directory has been established with standard repository infrastructure, including GitHub issue and pull request templates, a CI workflow for linting and macOS testing, Swift formatting and pre-commit configurations, and documentation. Within the MLXLLM library, the model registry has been narrowed to support only the Gemma family (Gemma, Gemma2, Gemma3, and Gemma3n), with the LLMModelFactory and LLMRegistry now exposing only these specific model configurations and types.

Packages/InferenceKit · high confidence

Introduces structured Live mode turn processing with semantic event parsing

This change introduces a new, structured pipeline for handling Live mode interactions, replacing the previous ad-hoc token streaming with a state-machine-based parser. The new \LiveOutputParser\ and \LiveTurnProcessor\ components parse the LLM's token stream into semantic events (\.complete\, \.interrupted\, \.thinking\, \.speechToken\, \.skillCall\, \.done\) rather than raw text. This allows the engine to react to specific turn states—such as ignoring text until a '✓' marker appears or handling tool calls wrapped in \\<tool\_call\>...\</tool\_call\>\ tags—enabling more robust voice interactions, multi-modal support, and future skill invocation capabilities.

Live/Turn · high confidence

Japanese localization for permission descriptions

Added Japanese translations for system permission usage descriptions (InfoPlist.strings) in the ja.lproj directory. This ensures that permission prompts for calendars, reminders, contacts, camera, microphone, and health data display correctly in Japanese when the device language is set to Japanese, covering features such as Live mode camera access and local health data processing.

ja.lproj · high confidence

Live Orb rendering engine restructured with native background and JS-driven 3D scene

The Live Orb UI has been restructured to separate the background rendering from the 3D orb visualization. A new \OrbBackgroundView\ now handles the deep purple radial gradient and film-grain noise overlay using SwiftUI and Canvas. The 3D orb itself is rendered via a \WKWebView\ (\OrbSceneView\) running a Three.js-based JavaScript renderer, which receives audio data from native analyzers at a capped 30Hz rate to ensure smooth performance. Additionally, \OrbShaderSource.swift\ provides a native Metal geometry modifier for SceneKit, porting the sphere deformation and normal-reconstruction logic from the previous JavaScript implementation to allow for native-side vertex manipulation.

Live/UI/LiveOrb · high confidence

Localized permission descriptions for Simplified Chinese users

Added a new InfoPlist.strings file for the zh-Hans localization to provide user-facing explanations for various system permissions. This ensures that when users on Simplified Chinese devices grant access for features like calendar, reminders, contacts, camera, microphone, and health data, they see clear, localized reasons for why the app needs these capabilities, such as creating calendar events, recording audio, or generating local health summaries.

zh-Hans.lproj · high confidence

Pre-commit hook enforces bilingual skill sync and LiteRT-based model harness

A new pre-commit hook has been added to enforce two quality gates: it always runs a bilingual synchronization check on skill files (ensuring English and Chinese translations remain in sync) and conditionally runs a CLI harness using the LiteRT runtime when changes are detected in agent, skill, prompt, router, or CLI source directories. The harness executes tests for both the E2B and E4B models, requiring both to pass (indicated by 'ALL GREEN') before allowing the commit, with an option to skip the harness via the PHONECLAW\_SKIP\_HARNESS environment variable.

.githooks · high confidence

Redesigned Live Mode interface with state-aware Orb and camera support

The Live Mode UI has been restructured into a dedicated \LiveModeUI.swift\ file, introducing a full-screen experience centered around the Orb visualization. The Orb now dynamically changes color and opacity based on the engine's state (e.g., dim gray when idle/preparing, warm amber when interactive), and a loading mask provides a smooth fade-in effect during initialization. The interface supports a camera preview mode that overlays the camera feed while dimming the Orb, and status text is now context-aware, displaying specific labels like 'Listening' or 'Processing' only when relevant. The layout also includes a top bar, a status capsule below the Orb, and a caption area for real-time text output.

Live/UI · high confidence

Refactored Live mode core engine and turn lifecycle management

The Live/Core module has been restructured into dedicated files, introducing a new LiveModeEngine that manages the full voice interaction pipeline (VAD, ASR, LLM, TTS) and a new VoiceTurnController that handles the user-turn state machine with a 100ms grace window for VAD jitter. The engine now includes a LiveTurnCompletionParser to handle streaming text markers (○, ◐, ✓) for incomplete turns, implements Pipecat-style interruption semantics to handle barge-ins during bot speech, and exposes state and metrics via the @Observable macro for UI binding.

Live/Core · high confidence

Skills library restructured with bilingual support and new capabilities

The Skills module has been refactored to support bilingual content (English and Japanese) via localized SKILL.md files, with Japanese users seeing Japanese or English content and Chinese users seeing Chinese. New built-in skills have been added: Health (with a health-steps-today tool), free web search, and multi-model capabilities including calendar, reminders, and contacts. The skill system now includes enhanced metadata such as skill type (device, content, network), activation modes, history policies, side-effect policies, and UI chip labels/prompts for better user interaction. User-editable skills are now stored in Application Support/skills/\<id\>/SKILL.md as overrides, with a new registration-based architecture replacing the previous file-scanning approach.

Skills · high confidence

Updated onnxruntime and piper-plus frameworks for iOS

The onnxruntime.xcframework has been updated to a newer version (API version 17), introducing new configuration options for the CoreML execution provider such as enabling the Apple Neural Engine (ANE) and restricting execution to static input shapes. Additionally, the piper\_plus.xcframework has been updated to version 1 of its API, adding support for ZH-EN code-switching dispatch to improve pronunciation of English loanwords in Chinese text, alongside standard synthesis and query capabilities.

Frameworks · high confidence

Xcode project restructured with new Live mode, download, and inference components

The PhoneClaw Xcode project has been significantly reorganized to support new capabilities. The build scheme is now explicitly defined, and the project file incorporates a large set of new source files including LiveModeEngine, LiveLand components (VoiceRuntime, VADService, TurnController), and a new download infrastructure (ResumableAssetDownloader, ModelInstaller). Inference capabilities are expanded with LiteRTBackend, WhisperKit, and piper\_plus.xcframework integration, while legacy CocoaPods references are removed in favor of local frameworks like PhoneClawEngine and MarkdownUI. Additionally, the project bundles Japanese language resources (OpenJTalk dictionary) and new UI assets (OrbWebAssets, HDR/EXR env maps) for the welcome screen.

PhoneClaw.xcodeproj · high confidence

Test coverage

Added Live Mode debugging and metrics tools; Added skill translation sync check and download subsystem tests; Added unit and contract tests for PhoneClawCore runtime and routing logic.

Dependencies

Added Swift Package Manager reference for LiteRTLM-Swift

A new symbolic link has been added to the Packages directory pointing to the LiteRTLM-Swift source location, establishing the package structure for Swift Package Manager integration.

Packages · high confidence

Dependency updates and removal of legacy MediaPipe pods

This change updates the project's dependency graph by removing the MediaPipe GenAI CocoaPods (MediaPipeTasksGenAI and MediaPipeTasksGenAIC) and pinning a new set of Swift Package Manager dependencies in the resolved lockfile, including swift-crypto 4.5.0, swift-markdown-ui 2.4.1, swift-transformers 1.1.9, and swift-jinja 2.3.5. It also introduces several new local Swift packages (IOS27CoreAIExperiment, PhoneClawEngine, FluidAudio, InferenceKit, and AudioTest) and removes the legacy PhoneClawMac package, consolidating the build system around the new local and remote SPM targets.

(dependencies) · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

Score

  • CAI 65 → 39 (-25.9)
  • Rubric changed (rubric-2026.09.11 → rubric-2026.09.18) — scores are not directly comparable.

Lenses

  • Code Health 77 → 78 (+1.2)
  • Architecture 98 → 98 (+0.3)
  • Maturity 72 → 72 (-0.1)
  • Readiness 57 → 56 (-0.9)
  • Security 63 → 73 (+9.2)
  • Event-Driven 10 (new)
  • Performance 62 (new)

Resolved (14)

  • Dependency hygiene PARTLY measured — Maven/Gradle declarations read, no dependency graph resolved
  • Documentation: no architecture or design documentation (Docs/iphone_agent_permissions_phase3.md)
  • Documentation: no installation or build instructions (README.md)
  • Documentation: no usage examples (README.md)
  • Edited copy of a member (13 corresponding lines) (MacGateway/Sources/PhoneClawGateway/GatewayProviders.swift)
  • Edited copy of a member (18 corresponding lines) (MacGateway/Sources/PhoneClawGateway/GatewayProviders.swift)
  • Hotspot: LLM/Backends/Remote/RemoteInferenceService.swift (LLM/Backends/Remote/RemoteInferenceService.swift)
  • Hotspot: LiveLand/Runtime/LiveLandVoiceRuntime.swift (LiveLand/Runtime/LiveLandVoiceRuntime.swift)
  • Hotspot: MacGateway/Sources/PhoneClawGateway/GatewayApp.swift (MacGateway/Sources/PhoneClawGateway/GatewayApp.swift)
  • Hotspot: MacGateway/Sources/PhoneClawGateway/GatewayProviders.swift (MacGateway/Sources/PhoneClawGateway/GatewayProviders.swift)
  • Hotspot: Shared/Audio/ASRService.swift (Shared/Audio/ASRService.swift)
  • Hotspot: Shared/LanguageService.swift (Shared/LanguageService.swift)
  • Hotspot: Tools/ToolRegistry.swift (Tools/ToolRegistry.swift)
  • Low cohesion: Module (LCOM4 5) (Packages/mlx-swift/Source/MLXNN/Module.swift)

New (87)

  • ChatModels.swift.buildDisplayItems (cognitive 61) (UI/ChatModels.swift)
  • ChatModels.swift.buildDisplayItems (cyclomatic 30) (UI/ChatModels.swift)
  • ContentView.body (cognitive 19) (UI/ContentView.swift)
  • ContentView.body (cyclomatic 17) (UI/ContentView.swift)
  • Duplicate intent across distinct service implementations. Both FoundationPlanningModelService and HeuristicPlanningModelService expose an identical availability() method returning the same type. While they are distinct implementations, the lack of a shared protocol or base class for this specific capability suggests a missed abstraction opportunity, though not a strict naming inconsistency.
  • Duplicated block (10 lines × 2) (UI/ResponseUI.swift)
  • Duplicated block (10 lines × 2) (UI/SkillsManagerView.swift)
  • Duplicated block (11 lines × 3) (Tools/AppPermissions.swift)
  • Duplicated block (11 lines × 3) (UI/AudioUI.swift)
  • Duplicated block (11 lines × 3) (UI/AudioUI.swift)
  • Duplicated block (11–12 lines × 2) (UI/ContentView.swift)
  • Duplicated block (12 lines × 2) (UI/AudioUI.swift)
  • Duplicated block (12–22 lines × 6) (MacGateway/Sources/PhoneClawGateway/GatewayApp.swift)
  • Duplicated block (13 lines × 2) (MacGateway/Sources/PhoneClawGateway/GatewayApp.swift)
  • Duplicated block (13 lines × 2) (UI/AudioUI.swift)
  • Duplicated block (13 lines × 2) (UI/ConfigurationsView.swift)
  • Duplicated block (13 lines × 2) (UI/ConfigurationsView.swift)
  • Duplicated block (14 lines × 2) (LLM/Backends/Remote/GatewayProviders.swift)
  • Duplicated block (14 lines × 2) (MacGateway/Sources/PhoneClawGateway/GatewayApp.swift)
  • Duplicated block (15 lines × 2) (LiveLand/Widget/PhoneClawLiveActivityWidget.swift)
  • …and 67 more

Architecture

  • Unchanged — 0 containers · 1 contexts · 0 edges

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

kellyvv/PhoneClaw was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 1 October 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit 31bf61d3219f597f62b7672f94475148bca5de19 — the exact code this score is about.
  • Scored under rubric-2026.09.18 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-e569280dd5e2.