ggml-org/Llama-macOS
48.9
Weak · 1 October 2026
13.2k
lines of production code
Swift
primary language
2
measurements over time
What this system is
Llama is a macOS menu-bar application that manages local LLM inference by handling model discovery, download, and server hosting. It provides system-wide prompt capture via a global input panel and allows users to compose and send API requests through a built-in builder. The system automatically manages the underlying llama.cpp engine, ensuring compatibility and secure network scoping for the local server.
Features
Introduce global input capture panel with system-wide hotkey
Users can now invoke a Spotlight-style floating panel from anywhere using a configurable global hotkey (none by default) or the "show global input" AppleScript command. The panel allows typing a prompt and selecting a target model via a filterable list that displays model names with parameter and quantization chips; the list is sorted by recently used models. Submitting the prompt sends it to the local web UI, and the feature is enabled for all users.
Llama/GlobalInput · high confidence
Introduce remote model catalog for Discover suggestions
The app now fetches a remote JSON catalog from llama.app to populate the in-app Discover section, replacing any previous hard-coded model lists. This change introduces a new Catalog module that parses featured families and sizes, filtering suggestions based on the device's memory budget and specific per-family or per-size memory caps defined in the catalog. Users will see recommended models with accurate size labels, quantization levels, and vision support indicators, with the app automatically selecting the highest-precision quant that fits the device's constraints.
Llama/Catalog · high confidence
New API request builder page with system-accent theming and improved UX
A new 'Build an API request' page (RequestBuilder) has been added to the Llama/Snippets section, allowing users to compose and send API requests directly from the app. The interface now uses the system's accent color for controls and highlights, replacing the previous hardcoded orange, and features neutral greys that adapt to any platform accent. The layout has been reorganized with controls on the left and the generated code snippet and response on the right, with 'Send' and 'Copy' actions moved below the request. The page provides clearer descriptions of options (what they are, not just what they send), disables unavailable features for the selected model with explanations, and shows progress states for loading and generating. It also fixes a sample image URL issue and correctly handles thinking and streaming behaviors per endpoint.
Llama/Snippets · high confidence
New shared utility library for common app functions
The Llama/Common module introduces a set of shared utilities to standardize core functionality across the application. This includes a Clipboard helper for consistent pasteboard operations, a FileManager extension for retrieving file sizes, and a comprehensive Formatters enum for handling byte, token, and memory display logic, including adaptive download readouts and metadata chip styling. Additionally, the module adds a QR code generator for displaying server addresses, a ModalPresentation helper that suppresses Sentry hang tracking during user interactions to prevent false positives, a ModelLogos mapper for brand-specific icons, and a centralized Notifications registry for app-wide state changes.
Llama/Common · high confidence
Support for installing models via llama:// deeplinks
Users can now install models directly from external links using the \llama://install?repo={org}/{repo}\[&quant={label}\]\ URL scheme. When a user clicks such a link, the app silently starts the download in the background and displays a menu-bar hint bubble to confirm the action, while in-app entry points (like the Discover catalog) use the same resolution logic without the extra notification. The system resolves the Hugging Face repository to find the appropriate GGUF file, optional mmproj sidecar, and MTP draft head, respecting the user's memory budget and handling errors like missing files or gated repos with clear, actionable alerts.
Llama/Deeplinks · high confidence
Removals
LlamaBarn app and core logic removed
The LlamaBarn application, including its SwiftUI menu bar interface, model catalog, server management, and download logic, has been deleted from the codebase.
LlamaBarn · high confidence
Removed Deno-based llama.cpp update script
The \scripts/update-llama-cpp.ts\ script, which previously used Deno to download and extract the latest llama.cpp release for macOS ARM64, has been removed. This change eliminates the automated mechanism for updating the local llama.cpp binaries and associated dynamic libraries (such as \libllama.dylib\ and \libggml-metal.dylib\) via the whitelist-based copy process.
scripts · high confidence
Behavioural changes
App rebranded to Llama with new identity, settings, and global input features
The application has been renamed from LlamaBarn to Llama, including the Xcode target, module, and source tree, with a migration path to carry over existing settings and staging files. The app now supports AppleScript automation via a new scripting definition, exposing commands to open the settings window (with optional tab selection) and the global-input capture panel. A new global-input capture panel is available for system-wide quick-capture of prompts to the web UI. The app enables launch-at-login by default on first launch and uses a more frequent update-check interval (every 2 hours via Sparkle). Outbound HTTP requests now include an explicit, stable User-Agent string to improve attribution in model download logs. The settings window has been simplified, and network access is scoped to the specific network it was granted on to prevent cross-network exposure.
Llama · high confidence
App renamed to Llama with automatic data migration and new launch settings
The application has been renamed from LlamaBarn to Llama, including its bundle identifier. To ensure a seamless upgrade, the app now automatically migrates your existing settings (such as HF cache location and context preferences) and in-progress downloads from the old identity to the new one on first launch. Additionally, the app is now configured to launch at login by default on its first run to keep it readily available in the menu bar, and it uses Metal's recommended working set size to more accurately determine available memory for model loading.
Llama/System · high confidence
App renamed to Llama with updated build configuration
The application has been renamed from LlamaBarn to Llama, affecting the target, scheme, module, and output bundle name. This change also updates the project's Xcode version compatibility to 2600, replaces the LaunchAtLogin dependency with Sentry for crash reporting, and simplifies the build process by removing the manual copy-files phase for llama-cpp binaries in favor of file-system synchronization.
Llama.xcodeproj · high confidence
Improved GGUF model detection and sidecar selection
The app now reads GGUF headers directly to detect multi-token prediction (MTP) heads and native context windows, allowing it to identify models that lack filename markers. Quantization naming is unified to match Hugging Face's canonical tags, ensuring consistent model identification. Sidecar selection for vision projectors (mmproj) and draft heads is improved to match any filename occurrence and pick the best bit-width match, while non-LLM files and known draft-head prefixes are excluded from main model selection.
Llama/HF · high confidence
Improved download reliability and live progress feedback
The download engine now provides a live transfer speed readout in the UI, calculated from a sliding time window of completed bytes. Resilience is enhanced by automatically pausing downloads when the network path is lost (e.g., wake-from-sleep) and auto-resuming them when connectivity returns, as well as by properly retiring download tasks that fail so they do not leave the model in a stuck 'downloading' state. Additionally, the system now scans for existing model cache files before starting a server to avoid redundant downloads.
Llama/Downloads · high confidence
New llama.cpp binary management and background update system
The app now manages its own llama.cpp CLI binary installation and updates. It resolves the binary from a managed path (\~/.llama-app/llama) or falls back to unmanaged locations like Homebrew, enforcing a minimum version (b9726) for unmanaged installs. The system supports background staging of the pinned target version (currently b11200) so that updates are applied at the next launch without interrupting the currently running server. It also handles version parsing for both older and newer llama.cpp output formats and ensures compatibility with specific features like multi-host binding (b11104).
Llama/Engine · high confidence
New model identification, compatibility, and memory budgeting system
This change introduces a new model identification system where every model uses a stable \{org}/{repo}:{TAG}\ ID, with display names and UI chips (quant, params, org prefix) derived from a parser that mirrors the llama.cpp WebUI. It adds a new ContextTier enum for context-length selection and a comprehensive compatibility check that estimates memory usage based on Metal's GPU working set and physical RAM, hiding context tiers the device cannot support. The system also adds support for Multi-Token Prediction (MTP) speculative decoding via embedded heads or sidecar files, and includes a DEBUG-only fixture for testing the installed-models list.
Llama/Models · high confidence
Redesigned Settings interface with sidebar navigation and enhanced controls
The Settings window has been restructured to include a persistent sidebar for navigation between sections (such as General, Network, and Advanced), replacing the previous layout. This update introduces a reusable 'PressableStyle' for consistent mouse-down feedback on chrome-less controls, a 'ShortcutRecorder' component for configuring global input hotkeys, and syntax-highlighted display for the server command. Additionally, the settings now support specifying a custom network bind address (including Tailscale IPs), allow enabling 'agent mode' and network access toggles with specific cautionary warnings, and provide a 'Restore Default' button for resetting individual settings.
Llama/Settings · high confidence
Redesigned menu bar interface with new model catalog and action rows
The menu bar interface has been completely redesigned to improve usability and visual consistency. A new \ActionItemView\ replaces standard buttons with labeled, icon-led rows that support inline two-step confirmations for destructive actions like deletion. The model discovery experience is enhanced with a \BrowseModelsRow\ for direct web catalog access and \CatalogItemView\ for curated suggestions, while the installed models list now features a \DisclosureRow\ to collapse long lists and an \EmptyInstalledRow\ placeholder. Server status and connectivity are clearer in the \HeaderView\, which displays a live status dot, the base URL with a copy button, and an inline QR code for easy mobile access. Additionally, \ExpandedModelDetailsView\ provides a segmented picker for context length that explicitly shows memory costs, and \CLISetupView\ offers clear feedback during engine installation or errors.
Llama/Menu · high confidence
Server port defaults to 9931 and network access is scoped to the current network
The server now listens on port 9931 by default to align with upstream changes, and network exposure is automatically disabled when the Mac joins a different network or when a stored Tailscale address becomes stale. This prevents the unauthenticated server from being inadvertently exposed on public or unfamiliar networks, while keeping localhost accessible for local use.
Llama/Server · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Score
- CAI 46 → 49 (+2.6)
- Rubric changed (rubric-2026.09.11 → rubric-2026.09.18) — scores are not directly comparable.
Lenses
- Code Health 95 → 95 (-0.0)
- Architecture 97 → 97 (+0.9)
- Maturity 56 → 56 (+0.0)
- Readiness 17 → 24 (+7.1)
- Security 100 → 100 (+0.0)
- Accessibility 68 → 68 (+0.0)
Resolved (4)
- Dependency hygiene PARTLY measured — SwiftPM pinning read, dependency currency NOT established
- Documentation: no architecture or design documentation (README.md)
- Duplicated block (6 lines × 2) (Llama/Server/LlamaServer.swift)
- Hotspot: Llama/HF/GGUFMetadata.swift (Llama/HF/GGUFMetadata.swift)
New (4)
- Change coupling: HFRepoResolver.swift ↔ HFCache.swift (Llama/Deeplinks/HFRepoResolver.swift)
- Duplicated block (6 lines × 2) (Llama/Models/SimulatedModels.swift)
- Outdated: sentry-cocoa
- Outdated: sparkle
Changes since last survey
- 11 commits — 11 feature/other, 0 fixes
By area
- Llama/Engine — 4 commits
- Llama/Menu — 2 commits
- Llama/GlobalInput — 1 commit
- Llama/HF — 1 commit
- Llama/Models — 1 commit
- Llama/Server — 1 commit
- Llama/Settings — 1 commit
Notable commits
- change: Add a custom web UI setting
- change: Build Hugging Face requests in one place
- change: Derive a model's display name from its id
- change: Describe a resolved llama binary with one enum
- change: Hide Unload when the model unloads
- change: Keep localhost reachable when bound to a specific address
- change: Move network address helpers out of LlamaServer
- change: Read the llama build number from its new version output
- change: Remove dead code and redundant plumbing
- change: Replace one-to-one notification relays with direct calls
- change: Update llama to b11200
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
ggml-org/Llama-macOS was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 1 October 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit c7fb61e3f4fd2cd8f28392d11f2714a98381753e — the exact code this score is about.
- Scored under rubric-2026.09.18 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-e569280dd5e2.