Skip to content
CAI
Software that uses CAICheck a score

jundot/omlx

42.3

Weak · 19 September 2026

458.3k

lines of production code

Python

with Swift

1

measurement over time

CAI band scale
CAI lens gauges

What this system is

oMLX is a local LLM inference server built on the MLX framework, optimized for Apple Silicon with support for Vision-Language Models, OCR, and tiered KV caching. It provides a native macOS application for managing server lifecycle, downloading models, and running performance benchmarks, alongside a web-based admin interface for configuration and monitoring. The system includes specialized tooling for model conversion, custom kernel compilation, and comprehensive benchmarking of specific hardware paths like the Neural Engine.

Features

Introduce native Swift macOS app with menubar-first lifecycle

The macOS application is now a native Swift app that runs primarily in the menu bar, keeping the main window hidden until the user opens it via the status item or the first-run Welcome wizard. Cmd-Q and Dock Quit now hide all windows and remove the Dock icon while preserving the background server and menu bar status item; full termination requires using the menu bar's Quit command. The app also restores the CLI launch shim and manages server lifecycle, update confirmations, and signal handling through a new AppDelegate.

apps/omlx-mac/Sources/App · high confidence

Introduce native Swift networking layer for macOS app

The macOS application now uses a dedicated Swift networking stack (\OMLXClient\ and \AdminAPI\ endpoints) to communicate with the backend. This replaces previous HTTP handling with a typed, async client that manages cookie sessions, API key authentication, and JSON serialization for all \/admin/api/\*\ routes, including model management, settings, usage stats, and benchmarking endpoints.

apps/omlx-mac/Sources/Net · high confidence

Introduce native macOS settings interface with server lifecycle controls

The macOS application now features a native SwiftUI settings shell (AppView) that replaces previous UI approaches, providing a sidebar-based navigation for Server, Models, Benchmarks, and General settings. This change introduces direct server lifecycle management (start, stop, restart) and storage configuration (base path, model directories, port) within the app, allowing users to manage the oMLX server state and configuration without leaving the native interface. It also includes an update confirmation sheet for reviewing release notes before downloading, and integrates system-style appearance settings with an Enhanced Readability toggle.

apps/omlx-mac/Sources/AppView · high confidence

Introduces native macOS configuration and CLI management components

The macOS app now includes a dedicated configuration layer in the \Config\ source directory. This adds \APIKeyGenerator\ for creating secure, crypto-grade API keys for the Security screen and Welcome wizard, and \AppConfig\ as the single source of truth for managing \settings.json\, including logic to normalize wildcard bind addresses to loopback hosts and resolve the base path via environment variables or a bootstrap file. Additionally, \ShellEnvWriter\ is introduced to manage the CLI launch shim, ensuring the \omlx\ command is available in the user's PATH and handling shell rc file exports only after explicit user consent.

apps/omlx-mac/Sources/Config · high confidence

Native Swift macOS app launch and menubar controls

The oMLX macOS application has been rewritten in native Swift using SwiftUI, replacing the previous implementation. This change introduces a new menubar interface that displays live system stats and allows users to manage the local server lifecycle directly from the menu bar. The app now includes dedicated screens for server status, logs, models, downloads, and performance, along with a new context benchmark tool to measure usable context windows. Additionally, the update mechanism has been ported from Sparkle to a custom Swift-based GitHub Releases updater, and the app now supports local usage history tracking.

apps/omlx-mac/oMLX.xcodeproj · high confidence

Native Swift macOS updater replaces Sparkle with GitHub Releases integration

The macOS application's update mechanism has been rewritten in native Swift, removing the external Sparkle framework. The new updater checks for updates directly against GitHub Releases, supporting Stable, Release Candidate, and Dev channels with PEP 440 version parsing. It downloads macOS-specific DMG assets, mounts them, stages the new application bundle, and uses a crash-safe launchd worker to atomically swap the app in place upon user confirmation.

apps/omlx-mac/Sources/Updater · high confidence

Native Swift menubar with live metrics and system stats

The macOS menubar has been rewritten in Swift to provide native parity with the previous Python implementation. Users now see live inference activity (prompt and generation throughput) via dedicated status items that can be toggled on or off in Appearance settings, along with a combined system resource monitor for CPU, GPU, and memory usage. The main menu includes updated controls for server lifecycle, web dashboard access, and model management, while a new visibility watcher automatically detects and alerts users if the menubar icon is hidden by system tools like Bartender or macOS Tahoe's ControlCenter.

apps/omlx-mac/Sources/Menubar · high confidence

Native Swift server process management for macOS

The macOS application now manages the oMLX server lifecycle using a native Swift implementation instead of the previous Python-based approach. This change introduces a new \AppControlServer\ for local IPC, a \PortConflictResolver\ to detect and handle port usage, and a \PythonRuntime\ module to correctly locate and configure the bundled Python interpreter (including environment variables like PYTHONHOME and PYTHONPATH). The \ServerProcess\ class now handles the child process lifecycle, including health checks, automatic restarts with backoff, and signal handling to ensure clean termination of the server when the app exits.

apps/omlx-mac/Sources/Server · high confidence

Native macOS app introduces comprehensive network data models

The native Swift macOS application now includes a complete set of Data Transfer Objects (DTOs) to communicate with the server API. These models enable the app to fetch and display benchmark results and system metrics, manage global server settings and usage history, handle Hugging Face and ModelScope model downloads and uploads, and configure per-model settings including experimental features like DFlash and ANE prefill. The changes also support profile management, security controls via sub-keys, and real-time server statistics.

apps/omlx-mac/Sources/Net/DTO · high confidence

Native macOS app resource bundle and localization

The macOS application now ships with a complete native resource bundle, including a new app icon with dark/light support, dedicated menu bar icons, and an Info.plist that defines the 'oMLX' display name and registers the 'omlxapp' URL scheme for deep-linking. Localization is expanded to include Russian, Simplified Chinese, Traditional Chinese, Japanese, and Korean, with a comprehensive Localizable.xcstrings catalog covering settings, menus, and UI strings.

apps/omlx-mac/Resources · high confidence

New Homebrew formula for oMLX with macOS 27 and custom kernel support

Users can now install the oMLX LLM inference server via Homebrew on Apple Silicon (macOS Sequoia and later). The formula includes support for building native custom kernels (Bonsai, GLM-5.2, MiniMax M3, Qwen3.5/3.6/4) and structured output via xgrammar. It specifically addresses macOS 27 beta compatibility by building certain packages from source to avoid dylib stripping issues, patches xgrammar for correct library loading, and bundles the spaCy English model for Kokoro TTS functionality.

Formula · high confidence

New benchmark suite for MoE offload, Qwen4 sparse attention, and Qwen3.5 ANE performance

Added a suite of new benchmark scripts to evaluate specific performance characteristics of the model runtime. The \moe\_offload\_prefill\_bench.py\ and \deepseek\_v41\_offload\_bench.py\ scripts measure expert offload latency, memory footprint, and fetch throughput during prefill and decode phases. The \bench\_qwen4\_qsa\_sparse\_gqa.py\ script benchmarks the native Qwen4 sparse GQA attention kernel against a portable gathered reference. Additionally, several scripts (\qwen35\_ane\_down\_fused\_poc.py\, \qwen35\_ane\_down\_output\_split\_poc.py\, \qwen35\_ane\_gdn\_split\_poc.py\, \qwen35\_ane\_prefill\_bench.py\) probe and benchmark the Qwen3.5 ANE prefill paths, including fused SwiGLU/down projections, output-row splitting, and GDN projections across single/dual ANE and CPU/GPU configurations. The \bench\_m5\_sorted\_gather\_chunk.py\ script benchmarks MoE prefill cost vs chunk width on M5 sorted gather paths, and \tp\_identity\_probe.py\ verifies tensor parallelism correctness by comparing token IDs between single-node and distributed runs.

benchmarks · high confidence

New benchmarking, download, and integration management screens in the macOS app

The macOS app now includes dedicated screens for running accuracy and context benchmarks, managing model downloads from Hugging Face and ModelScope, and configuring third-party integrations like Claude Code and MCP. The Accuracy and Context benchmark view models allow users to select models, configure parameters (such as target tokens and prefill priority), and monitor job status via polling. The Downloads screen supports switching between Hugging Face and ModelScope sources, configuring mirrors, and viewing recommended models. The Integrations screen provides a centralized interface to configure Claude Code (including model selection and context scaling), other AI tools (Codex, Copilot, etc.), and MCP configurations, with live command generation for launching these tools.

apps/omlx-mac/Sources/AppView/ViewModels · high confidence

New benchmarking, i18n normalization, and cluster testing scripts

Added several new maintenance and testing scripts to the scripts directory: bench.py provides a standalone Python interface for running single-request and continuous-batching performance benchmarks against local models; build\_benchmark\_corpora.py generates stable UTF-8 text corpora from source code and public literary works for throughput testing; cluster\_context\_gate.py acts as a black-box harness to validate context window limits against an oMLX endpoint; normalize\_i18n.py synchronizes translation files by adding missing keys and removing extra ones based on the English source; and replay\_anthropic\_payload.py allows developers to replay captured Anthropic API payloads against an oMLX server for debugging.

scripts · high confidence

New first-run welcome wizard for macOS

A new three-step onboarding flow (intro, setup, success) is presented to users on their first launch of the macOS app. This transient window allows users to configure the base directory, model directory, and server port, validates these settings, and persists the configuration before starting the server. Returning users will not see this window, as it is skipped when existing settings are detected.

apps/omlx-mac/Sources/Welcome · high confidence

New macOS app screens for About, Benchmarks, Downloads, and Integrations

The macOS app now includes dedicated screens for viewing app information (About), running accuracy and context benchmarks, managing model downloads from Hugging Face and ModelScope, and configuring integrations like Claude Code. These screens provide native SwiftUI interfaces for tasks previously handled in the web admin, allowing users to measure model performance, download models, and set up local AI tool integrations directly within the macOS application.

apps/omlx-mac/Sources/AppView/Screens · high confidence

New macOS build and validation scripts for the native Swift app

Added \build.sh\, \check\_custom\_kernel\_python\_abi.py\, and \dev\_server.py\ to streamline the creation and verification of the native Swift oMLX.app. The build script automates the Swift compilation, venvstacks Python environment staging, and code-signing, while the new ABI checker ensures that bundled native custom kernels match the donor app's Python extension ABI to prevent runtime crashes. A lightweight dev server stub is also included to help developers verify the Swift app's server process spawning during local iteration.

apps/omlx-mac/Scripts · high confidence

New model conversion tools for FP16 cloning and ternary weight repacking

Added two new command-line utilities in the tools directory to support advanced model optimization workflows. The new \clone\_mlx\_model\_fp16.py\ script creates an FP16 clone of MLX quantized models by converting floating-point checkpoint tensors while preserving packed integer weights, specifically enabling the optional Qwen3.5/3.8 ANE+CPU prefill path. Additionally, \repack\_ternary\_t5.py\ provides a lossless method to repack MLX 2-bit ternary checkpoint weights into a base-3 (t5) format, reducing weight size by approximately 20% (from 2.0 to \~1.585 bits per weight) without changing dequantized values.

tools · high confidence

oMLX documentation and build configuration updates

The project introduces localized README files for French, Japanese, Korean, and Chinese, alongside a comprehensive SECURITY.md policy outlining vulnerability reporting and trust assumptions. The main README is updated to reflect new capabilities including Vision-Language Models (VLM), OCR support, tiered KV caching, and multi-model serving, while also documenting the new macOS app lifecycle commands (\omlx start/stop/restart\) and Homebrew service integration. Build configuration is adjusted in \setup.py\ to support native custom kernels for specific model families (GLM-5.2, MiniMax M3, Qwen3.5) via CMake extensions, and \pytest.ini\ is updated to include a marker for TurboQuant tests.

(repo-wide) · high confidence

Removals

Removal of the native macOS menubar application

The native macOS menubar application (oMLX App) has been removed from the codebase. This eliminates the PyObjC-based interface that previously allowed users to manage the oMLX server directly from the menu bar, including features such as start/stop controls, health monitoring, automatic updates, and configuration via Preferences and Welcome windows.

_packaging/omlx\app · high confidence

Behavioural changes

Introduce native SwiftUI theme components for the macOS app

The macOS app now uses a new set of native SwiftUI components for its interface, replacing previous custom implementations. This includes themed buttons (primary, destructive, plain), a click-to-copy code chip, structured list groups with labeled rows, native-style popup pickers, progress bars, and standardized section headers and footers. These components ensure the app's controls share bezel metrics, fonts, and disabled/pressed treatments with native macOS system controls, providing a consistent and native-feeling user experience.

apps/omlx-mac/Sources/Theme/Components · high confidence

Packaging pipeline retires PyObjC menubar app in favor of Swift bundle

The packaging directory no longer builds the macOS application bundle or DMG; it exclusively generates the venvstacks Python layers (runtime and framework) that the Swift macOS app embeds. The old \omlx-app\ layer using PyObjC has been removed, and the framework layer has been renamed from \mlx-framework\ to \mlx-base\. The build script now supports \--venvstacks-only\ and \--print-fingerprint\ to produce the \\_export/\ tree, while the Swift app build is handled separately by \apps/omlx-mac/Scripts/build.sh\. Additionally, a new \--macos-target\ flag allows swapping MLX wheels for specific macOS platform versions.

packaging · high confidence

Visible scrollbar styling for pill rows in admin panel

The admin panel now ensures that pill rows remain scrollable with a visible scrollbar on macOS, where scrollbars are typically hidden by default. A new CSS utility class (.pills-scrollbar) applies a thin, semi-transparent scrollbar style to improve the overflow cue for users.

omlx/admin · high confidence

macOS settings UI adopts native macOS design tokens and liquid-glass visuals

The macOS settings interface now uses a new theming system that resolves colors via dynamic AppKit system colors, ensuring the app automatically tracks system light/dark mode, accent colors, and accessibility contrast settings. This includes a new 'Enhanced Readability' toggle that increases text contrast by promoting secondary and tertiary text to primary weight and intensifying danger/warning colors. Visually, the UI now features a 'liquid glass' effect for sidebar and toolbar surfaces (using native macOS 26 APIs with fallbacks for earlier versions) and gradient-filled 'squircle' icons for sidebar items, aligning the app's appearance with current macOS conventions.

apps/omlx-mac/Sources/Theme · high confidence

Fixes

Remove TurboQuant and native MTP restrictions from the web UI

The web admin dashboard no longer hides or disables the TurboQuant and native MTP (Lightning MTP) feature toggles. These settings are now available for all models, allowing users to enable these optimizations directly from the UI without encountering previous interface restrictions.

omlx · high confidence

Test coverage

Added integration tests for TTS streaming, cache consistency, and real-model validation; Added test coverage for macOS app core components; Expanded test coverage for core engine and cache components.

Dependencies

Upgrade MLX stack and align dependencies for v0.32.2 compatibility

The core MLX dependency is pinned to version 0.32.2, requiring a corresponding update to nanobind (2.15.0) to ensure ABI compatibility for custom kernel extensions. The mlx-lm library is upgraded to a specific commit (872ae88) that supports full cache state and per-stream stop matchers, while mlx-vlm is updated to commit 3fb24e9. To support these updates, the transformers requirement is tightened to versions 5.14.0 through 5.17, and mistral-common is pinned to a minimum of 1.10 to prevent processor resolution issues. Additional dependencies such as huggingface-hub (\>=1.19.0), regex, and tiktoken are added or updated to support new model features, administrative tools, and tokenizer backends.

(dependencies) · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

Baseline

  • First survey — no prior run to compare against. CAI 42.

Lenses

  • Code Health 31
  • Architecture 97
  • Maturity 63
  • Readiness 43
  • Security 71
  • Domain Modelling 100
  • Accessibility 44

Changes since last survey

  • 300 commits — 84 feature/other, 216 fixes

By area

  • omlx/patches — 74 commits
  • omlx/admin — 38 commits
  • omlx/engine — 24 commits
  • omlx/api — 13 commits
  • omlx/cluster — 13 commits
  • apps/omlx-mac — 12 commits
  • docs/TESTING.md — 12 commits
  • omlx/cache — 12 commits
  • omlx/scheduler.py — 12 commits
  • (repo) — 8 commits
  • omlx/custom_kernels — 8 commits
  • omlx/_version.py — 7 commits
  • omlx/utils — 6 commits
  • (root) — 5 commits
  • omlx/engine_pool.py — 5 commits
  • omlx/eval — 4 commits
  • omlx/memory_monitor.py — 4 commits
  • omlx/oq.py — 4 commits
  • Formula/omlx.rb — 3 commits
  • omlx/models — 3 commits

Notable commits

  • fix: Fix/glm5 promoted boundary (depends on PR #3290) (#3710)
  • fix: Merge pull request #1 from mvdbos/fix/qwen4-qsa-path-log-dedup
  • fix: Merge pull request #3080 from monroewilliams/fix/ane-tuner-disable-specprefill
  • fix: fix(admin): add dark mode colors for remaining model setting pills
  • fix: fix(admin): improve dashboard accessibility (#2669)
  • fix: fix(admin): let the Lightning MTP toggle through for Qwen3.8 Flash Next (#3200)
  • fix: fix(admin): offer the API toggle on the New Profile form (#3658)
  • fix: fix(admin): persist mtp_num_draft_tokens via settings PUT (#3280)
  • fix: fix(admin): retry failed Xet downloads over HTTP (#3158) (#3186)
  • fix: fix(admin): size U32-packed MLX quants from Hub blob bytes (#3419)
  • fix: fix(admin): validate Lightning MTP draft depth (#3280)
  • fix: fix(admin): validate settings before saving profiles and templates (#3674)
  • fix: fix(ane): clear the per-module state caches when releasing ANE banks (#3340)
  • fix: fix(ane): explain a below-floor GDN fraction instead of compiling nothing
  • fix: fix(ane): keep compile cache cleanup inside the cache root
  • fix: fix(ane): keep compile cache lock files stable
  • fix: fix(ane): preserve GDN quality with recurrent-safe prefill offload (#3133)
  • fix: fix(ane): transfer dispatch tickets after thread spawn
  • fix: fix(api): parse attribute-style <function name> tool calls (MiniCPM5) (#3530)
  • fix: fix(api): preserve Anthropic plain-text document content (#3617)
  • …and 280 more

Architecture

  • 0 containers · 1 bounded contexts · 0 dependency edges (baseline)

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

jundot/omlx was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 19 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit 14194fe74bab38b89c144bd89656fbedca641d14 — the exact code this score is about.
  • Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-13a154b7f5d1.