mudler/LocalAI
45.1
Weak · 24 September 2026
208.1k
lines of production code
Go
with JavaScript, Python, C++
4
measurements over time
What this system is
This system is a self-hosted, multi-modal AI inference platform that unifies text, audio, image, video, and 3D generation through a modular backend architecture. It supports a wide variety of model types and hardware accelerators via a plugin-based system of Go, C++, and Python services, all exposed through a standardized gRPC interface. The platform provides a comprehensive HTTP API compatible with OpenAI, Ollama, and Anthropic standards, alongside a modern React-based web UI for management and interaction. It also features distributed deployment capabilities, allowing for federated P2P networking, multi-node load balancing, and centralized agent job orchestration.
How it got here
2023–2024 — distributed architecture and web UI overhaul
52 changes.
This period focused on restructuring the system for distributed inference, introducing P2P networking, federated load balancing, and a standardized gRPC backend infrastructure. It also delivered a comprehensive web UI redesign using Alpine.js and added extensive support for new AI modalities, including agents, 3D generation, and various Python-based inference backends.
2025–2026 — Distributed infrastructure and React UI migration
144 changes.
This period focused on establishing a robust distributed architecture using NATS and PostgreSQL for job scheduling, state persistence, and cross-replica synchronization, alongside a comprehensive rewrite of the frontend to a React-based interface. It also involved expanding the backend ecosystem with numerous new gRPC services for AI models, audio processing, and biometrics, while hardening security through authentication, authorization, and OCI image verification.
Features
Add ACE-Step music generation backend
Introduces a new gRPC-based backend for ACE-Step 1.5 music generation, enabling users to generate audio from text prompts, lyrics, or metadata. The implementation includes a Python service (backend.py) that handles model loading, audio sample creation, and formatting, along with installation scripts, a Makefile for build/test management, and unit tests to verify health, model loading, and sound generation capabilities.
backend/python/ace-step · high confidence
Add AceStep C++ backend for text-to-music generation
This change introduces a new backend for LocalAI that integrates the AceStep C++ library to generate music from text prompts. The implementation includes a Go wrapper and C++ bindings that load multiple GGUF models (a language model, text encoder, DiT model, and VAE) to produce audio output. The backend supports various CPU instruction sets (AVX, AVX2, AVX512) on Linux and Metal on macOS, with build scripts and packaging logic to handle these variants. It exposes a gRPC interface for loading models and generating sound, allowing users to specify parameters like BPM, duration, and lyrics.
backend/go/acestep-cpp · high confidence
Add Anthropic Messages API endpoint with tool calling and thinking support
Introduces a new HTTP endpoint at /v1/messages that implements the Anthropic Messages API specification. This allows clients to interact with models using Anthropic's message format, including support for tool/function calling and interleaved 'thinking' (reasoning) blocks. The implementation converts incoming Anthropic requests to the internal OpenAI-compatible format, injects Model Context Protocol (MCP) tools when configured, and handles both streaming and non-streaming responses while correctly mapping Anthropic-specific features like reasoning content and tool use events.
core/http/endpoints/anthropic · high confidence
Add Apple Silicon MLX video generation backend
A new gRPC-based video generation backend is introduced for macOS on Apple Silicon, supporting LTX-2 and converted Wan2.1/Wan2.2 MLX checkpoints. The service exposes model loading, video generation, and health-check endpoints, enforcing platform requirements and rejecting unsupported features like audio conditioning. It is managed via standard Makefile targets (install, run, test) and includes unit tests for model classification and command building.
backend/python/mlx-video · high confidence
Add Bonsai backend build and runtime infrastructure
This change introduces the build and packaging system for the Bonsai backend, which uses a PrismML llama.cpp fork to support Q1\_0 (1-bit) and Q2\_0 (ternary) weight quantization. The new Makefile, apply-patches.sh, package.sh, and run.sh scripts in backend/cpp/bonsai handle cloning the fork, applying necessary compatibility patches, building the gRPC server for various CPU architectures (AVX2, AVX512, fallback, and CPU\_ALL\_VARIANTS), and packaging the resulting binaries and shared libraries for deployment.
backend/cpp/bonsai · high confidence
Add CED sound-event classification backend
A new sound-event classification backend (CED) is now available, allowing users to detect and tag audio content using the ced.cpp model. This change introduces the Go-based gRPC server implementation (main.go, goced.go) that binds to the libced.so shared library via purego, along with build scripts (Makefile, package.sh) and runtime helpers (run.sh) to manage the C++ dependency and packaging. Users can now load CED-compatible GGUF models to perform audio tagging via the SoundDetection gRPC endpoint.
backend/go/ced · high confidence
Add Chatterbox TTS backend with multilingual support and text chunking
Introduces a new gRPC-based TTS backend for Chatterbox, enabling text-to-speech generation via a dedicated server. This backend supports multilingual synthesis and automatically splits long input texts into manageable chunks to prevent processing errors, merging the resulting audio segments into a single output file. It includes installation scripts that handle dependency isolation to avoid conflicts, and provides a test suite to verify server startup, model loading, and TTS functionality.
backend/python/chatterbox · high confidence
Add ElevenLabs Sound Generation and TTS API endpoints
New HTTP handlers have been added to expose ElevenLabs-compatible audio generation capabilities. The \SoundGenerationEndpoint\ allows users to generate audio from text with support for parameters like duration, temperature, and instrumental mode, while the \TTSEndpoint\ provides a text-to-speech interface compatible with the ElevenLabs API structure. These endpoints integrate with the existing backend model loading and audio normalization logic to return generated audio files.
core/http/endpoints/elevenlabs · high confidence
Add Go client for the vector store API
A new \StoreClient\ implementation in Go has been added to the \core/clients\ package, enabling programmatic interaction with the store API. This client supports standard vector store operations including setting key-value pairs, retrieving values by keys, deleting keys, and finding similar vectors (top-k search) against a configurable base URL.
core/clients · high confidence
Add Jina-style reranking API endpoint
A new HTTP endpoint has been added to support the Jina reranker API specification, allowing users to rerank a list of documents based on their relevance to a given query. This change introduces the \JINARerankEndpoint\ handler which processes requests, validates parameters like \top\_n\, and delegates the actual reranking logic to the backend, returning results in the Jina-compatible response format.
core/http/endpoints/jina · high confidence
Add Kokoros TTS backend
Introduces a new Kokoros text-to-speech backend for the Rust backend area. This includes the build configuration (Makefile, build.rs) to compile the gRPC service from protobuf definitions, a packaging script (package.sh) that bundles the binary, dynamic libraries, espeak-ng data, and SSL certificates for portability, and a runtime script (run.sh) to execute the service with the correct library paths. The submodule for the Kokoros source code has been updated to commit 29e99ad.
backend/rust · high confidence
Add Liquid Audio backend for real-time speech and fine-tuning
A new Liquid Audio backend is introduced, wrapping the liquid-audio Python package to support chat, ASR, TTS, and real-time speech-to-speech (s2s) modes. The implementation includes a gRPC service that handles model loading with device selection (CUDA, MPS, or CPU), manages voice-specific prompts, and exposes a fine-tuning capability via a dedicated mode that allows training without loading inference weights. The entry also includes the necessary build scripts, installation logic requiring Python 3.12, and smoke tests to verify service health and basic functionality.
backend/python/liquid-audio · high confidence
Add LongCat Video backend for text-to-video and avatar generation
Introduces a new Python gRPC backend for Meituan's LongCat-Video and LongCat-Video-Avatar-1.5 models, enabling text-to-video, image-to-video, and audio-driven avatar generation. The backend includes a Makefile to fetch and patch the upstream source with a PyTorch SDPA attention fallback for systems lacking optional kernels, and supports configuration options for attention backends, resolution, and model variants.
backend/python/longcat-video · high confidence
Add MLX-Audio TTS backend
A new MLX-Audio backend has been added to the Python backends, providing Text-to-Speech functionality via a gRPC service. This implementation allows users to load MLX-based TTS models (such as Kokoro) and generate audio, supporting configuration options for voice, speed, and language. The addition includes the necessary backend logic, installation scripts, and integration tests to verify service startup, model loading, and audio generation capabilities.
backend/python/coqui, backend/python/kokoro, backend/python/mlx-audio · high confidence
Add MLX-VLM backend for multimodal vision-language model inference
Introduces a new MLX-VLM backend implementation that enables the platform to load and run multimodal vision-language models using the MLX framework. This backend supports gRPC-based model loading and prediction, including multimodal inputs (images and audio), tool-use parsing, and reasoning content extraction, along with the necessary installation, execution, and test scripts to integrate it into the existing backend infrastructure.
backend/python/mlx-vlm · high confidence
Add MOSS-TTS C++ text-to-speech backend
Users can now generate speech using the MOSS-TTS-Local (v1.5) model via a new C++ backend. This backend runs the model through moss-tts.cpp, producing 48 kHz stereo WAV output without requiring Python at inference time. It supports voice cloning via reference audio paths and allows configuration of the transformer, audio codec, and text tokenizer GGUFs. The backend includes build scripts for various CPU architectures (AVX, AVX2, AVX512) and GPU backends (CUDA, HIP, Metal, Vulkan, SYCL).
backend/go/moss-tts-cpp · high confidence
Add Magpie TTS C++ backend for multilingual text-to-speech
This location introduces the Go wrapper and build infrastructure for the Magpie TTS C++ backend, enabling text-to-speech synthesis using NVIDIA's Magpie TTS Multilingual 357M model via the magpie-tts.cpp library. The backend loads the model as a shared library (libgomagpiettscpp) and exposes a C-API interface for synthesis, supporting 5 baked voices (Aria, Jason, John, Leo, Sofia) and 12 languages. It outputs 22.05 kHz mono 16-bit PCM WAV files, with streaming support that emits a self-describing WAV header followed by PCM chunks. The build system (Makefile/CMakeLists.txt) handles GPU acceleration (CUDA, ROCm/HIP, Metal, Vulkan, SYCL) and CPU optimizations (AVX, AVX2, AVX512) via variant-specific shared libraries. The Go code handles voice/language resolution, case-insensitive matching, and error handling, while tests verify synthesis output and streaming behavior.
backend/go/magpie-tts-cpp · high confidence
Add Moonshine gRPC backend for faster transcription
Introduces a new Moonshine backend that exposes a gRPC service for audio transcription, allowing users to load models and transcribe audio files with timing segments. This includes the backend implementation, build/install scripts, and unit tests to verify server startup, model loading, and transcription accuracy.
backend/python/moonshine · high confidence
Add NATS JWT authentication and TLS/mTLS options for distributed mode
Introduces a new \pkg/natsauth\ package that enables secure, credential-based communication in distributed LocalAI deployments. Users can now configure NATS JWT authentication using account seeds and service user credentials, with options to enforce authentication (\RequireAuth\) and control worker JWT expiration (\WorkerJWTTTL\). The system automatically generates scoped, per-node worker JWTs with specific publish/subscribe permissions for backend and agent workers, ensuring least-privilege access. A validation step warns if distributed NATS is reachable without credentials, and a setup script (\scripts/nats-auth-setup.sh\) provides recommended service-user permissions to cover frontend publishing needs like prefix cache sync.
pkg/natsauth · high confidence
Add NVIDIA NeMo ASR backend with word-level timestamp support
Introduces a new gRPC-based backend for NVIDIA NeMo Automatic Speech Recognition, enabling users to transcribe audio using models like nvidia/parakeet-tdt-0.6b-v3. This backend supports word-level timestamps, allowing for granular timing information in transcription results, and includes installation, testing, and runtime scripts to facilitate deployment. The implementation handles model loading from local snapshots or Hugging Face, supports CUDA and MPS devices, and provides a health check endpoint for service monitoring.
backend/python/nemo · high confidence
Add NeuTTSAir gRPC backend for text-to-speech
Introduces a new gRPC-based backend service for NeuTTSAir, enabling text-to-speech generation with support for voice cloning via reference audio. The implementation includes a Python server (backend.py) that handles model loading and inference, utilizing PyTorch for device detection (CUDA, MPS, CPU) and integrating with the neutts-air library. The change also adds build and test infrastructure (Makefile, install.sh, run.sh, test.sh) to manage dependencies, start the service, and run unit tests for health checks, model loading, and TTS functionality.
backend/python/neutts · high confidence
Add OmniVoice C++ TTS backend with voice cloning and streaming
Introduces a new text-to-speech backend powered by the ServeurpersoCom/omnivoice.cpp library, accessible via a Go wrapper. This backend supports file-based synthesis and streaming audio output, and enables voice cloning (using reference audio) and voice design (using instructions). It includes build infrastructure for various CPU instruction sets (AVX, AVX2, AVX512) and GPU backends (CUDA, HIP, Metal, Vulkan, SYCL), along with e2e tests verifying WAV generation and streaming chunk delivery.
backend/go/omnivoice-cpp · high confidence
Add Pocket TTS backend support
A new Pocket TTS backend has been added to the system, providing a gRPC-based service for text-to-speech generation. This implementation includes a Python backend server that supports loading models on CPU, CUDA, or MPS devices, caching voice states, and generating audio output. The addition comes with a Makefile for build and run management, installation scripts handling dependencies (including specific configurations for Intel and L4T profiles), and a comprehensive test suite verifying server startup, model loading, and TTS generation with both HuggingFace voice URLs and default voices.
backend/python/pocket-tts · high confidence
Add Qwen3-ASR local speech recognition backend
Introduces a new gRPC-based backend for the Qwen3-ASR model, enabling local speech-to-text transcription. The implementation includes device selection logic that supports CUDA, Intel XPU, and MPS hardware, along with a language mapping layer that translates ISO 639-1 codes (e.g., 'de', 'zh') into the full language names required by the Qwen3-ASR library. The package comes with shell scripts for installation and execution, a Makefile for build management, and unit tests for device utility functions and end-to-end server health and model loading.
backend/python/qwen-asr · high confidence
Add SAM3-CPP vision backend for point, box, and text-prompted segmentation
This change introduces a new C++-based backend for the SAM3 segmentation model, enabling users to perform image segmentation using point prompts, box prompts, or text prompts (PCS mode). The backend wraps the PABannier/sam3.cpp library (version 416186c) and exposes it via a Go wrapper that communicates over gRPC. It supports multiple CPU instruction set variants (AVX, AVX2, AVX512) on Linux and includes Metal support for macOS. Users can now load SAM3 models and detect objects with associated bounding boxes, confidence scores, and PNG-encoded masks.
backend/go/sam3-cpp · high confidence
Add Silero VAD Go backend with macOS support
This change introduces a new Go-based backend for Silero Voice Activity Detection (VAD), providing a standalone gRPC server implementation. The backend now supports macOS (Darwin) in addition to Linux, handling architecture-specific ONNX runtime library packaging and dynamic library path configuration via the new Makefile and run.sh scripts. Users can now utilize this dedicated backend for VAD tasks, which loads the Silero detector and exposes speech segment detection capabilities through the standard gRPC interface.
backend/go/silero-vad · high confidence
Add Supertonic ONNX TTS backend
Users can now use the Supertonic ONNX Text-to-Speech backend for local voice synthesis. This new Go-based backend supports CPU and CUDA execution providers, enabling high-quality speech generation with configurable voice styles, languages, and synthesis parameters (steps, speed, silence). It includes streaming support and handles model loading via the standard LocalAI gRPC interface.
backend/go/supertonic · high confidence
Add VibeVoice backend with TTS and ASR support
Introduces a new gRPC-based backend for VibeVoice, enabling both Text-to-Speech (TTS) and Automatic Speech Recognition (ASR) capabilities. The backend loads models such as \microsoft/VibeVoice-Realtime-0.5B\ for TTS and \microsoft/VibeVoice-ASR\ for transcription, supporting streaming inference and configurable options like device selection (CUDA/MPS/CPU), voice presets, and generation parameters. Includes installation scripts, a Makefile for build/test/run workflows, and unit tests for the gRPC service endpoints.
backend/python/vibevoice · high confidence
Add VoxCPM TTS backend with streaming and voice cloning support
A new gRPC-based backend for the VoxCPM text-to-speech model has been added to the system. This integration enables users to generate audio using the VoxCPM1.5 model, featuring support for both standard and streaming generation modes, as well as voice cloning capabilities via prompt audio and text. The implementation includes a Python service that handles model loading on CUDA, MPS, or CPU devices, applies a specific patch to resolve PyTorch compatibility issues in the attention module, and pins setuptools to ensure stable dependency resolution. A comprehensive test suite is included to verify model loading and streaming functionality.
backend/python/voxcpm · high confidence
Add Voxtral audio transcription backend
Introduces a new backend for audio transcription using the Voxtral model. This change adds the Go wrapper, C/C++ integration code, and build scripts (CMake/Makefile) required to load Voxtral models and perform transcription via gRPC. It includes platform-specific optimizations, such as Metal support on macOS and AVX/AVX2 variants on Linux, along with packaging scripts to bundle the necessary shared libraries.
backend/go/voxtral · high confidence
Add Whisper-Medusa speech recognition backend
A new local AI backend for Whisper-Medusa speech recognition has been added to the system. This component implements a gRPC service that handles audio transcription, supporting automatic device selection (CUDA, MPS, or CPU), audio preprocessing (resampling to 16kHz, mono conversion), and configurable inference parameters such as language and exponential decay length penalty. The backend includes a default model reference ('aiola/whisper-medusa-linear-libri') and is packaged with installation, build, and test scripts to facilitate deployment.
backend/python/whisper-medusa · high confidence
Add WhisperX backend for transcription with speaker diarization
Introduces a new Python-based gRPC backend for WhisperX, enabling audio transcription with word-level timestamps, forced alignment, and speaker diarization. The implementation includes a dedicated server (\backend.py\) that loads WhisperX models, handles audio processing, and manages alignment caching. It supports device selection (CPU, CUDA, MPS) and requires an Hugging Face token (\HF\_TOKEN\) to enable the gated diarization pipeline, rejecting requests if the token is missing when diarization is requested. The change also adds build scripts (\install.sh\, \protogen.sh\, \run.sh\), a Makefile for managing the environment, and unit tests to verify server startup, model loading, and transcription functionality.
backend/python/whisperx · high confidence
Add faster-qwen3-tts voice cloning backend
A new gRPC-based backend for the Faster Qwen3-TTS model has been added to the system, enabling voice cloning capabilities. This component provides a CUDA-dependent inference service that loads the Qwen3-TTS model and generates audio from text using a reference voice sample, exposing Health, LoadModel, and TTS endpoints for integration.
backend/python/faster-qwen3-tts · high confidence
Add faster-whisper backend for speech-to-text
A new faster-whisper backend has been added to the speech-to-text capabilities, providing a gRPC-based service that leverages the faster-whisper library for transcription. This backend supports device selection (CPU, CUDA, and MPS) and enables word-level timestamp granularity in the transcription output. The implementation includes specific installation logic for ROCm environments via CTranslate2 wheels and enforces a 50MB gRPC message size limit.
backend/python/faster-whisper · high confidence
Add llama.cpp backend with speculative decoding and VRAM management
This change introduces a new GRPC-based backend for llama.cpp models, enabling users to run inference via the LocalAI server. The implementation supports speculative decoding through draft models, allowing for potential performance improvements, and includes a dedicated Free method to properly release GPU VRAM when models are unloaded. It also exposes various model loading options such as context size, LoRA adapters, and RoPE frequency settings.
backend/go/llm · high confidence
Add locate-anything-cpp backend for open-vocabulary object detection
Introduces a new C++-based backend that enables open-vocabulary object detection using the locate-anything.cpp library (ggml). This addition includes the Go gRPC handler for loading GGUF models and running detection, a CMake build system for compiling the native library with CPU feature variants (AVX, AVX2, AVX512) and macOS Metal support, and end-to-end smoke tests to verify the detection pipeline.
backend/go/locate-anything-cpp · high confidence
Add multilingual support with Indonesian, Korean, and Portuguese (Brazil)
The React UI now supports multiple languages, including newly added Indonesian (id), Korean (ko), and Portuguese (Brazil) (pt-BR), alongside existing English, Italian, Spanish, German, and Simplified Chinese. This change introduces the i18n infrastructure in core/http/react-ui/src/i18n/index.js, configuring i18next with lazy loading of translation files via HTTP backend and automatic language detection from the browser or local storage. Users can now switch interface languages, with translations loaded on demand for namespaces like common, nav, errors, auth, home, chat, studio, models, agents, skills, collections, biometrics, media, tools, admin, usage, and explorer.
core/http/react-ui/src/i18n · high confidence
Add prompt templates for Alpaca, Koala, Llama 2, Vicuna, and other models
New prompt templates have been added to the prompt-templates directory to support specific model formatting requirements. These include templates for Alpaca, Koala, Vicuna, WizardLM, and GPT4All-J, as well as a specific template for Llama 2 chat messages that handles system prompts and role-based content. A generic 'getting\_started' template is also included. These files define the input/output structures used by the application when interacting with these respective models.
prompt-templates · high confidence
Add rfdetr-cpp native object detection and segmentation backend
Introduces a new C++-based backend for LocalAI that provides object detection and instance segmentation capabilities using the rf-detr.cpp library. This change adds the Go gRPC server implementation (gorfdetrcpp.go, main.go) which loads the native shared library via purego and exposes detection endpoints. It also includes the build infrastructure (CMakeLists.txt, Makefile) to compile the backend with support for various CPU instruction sets (AVX, AVX2, AVX512) and GPU backends (CUDA, Metal, ROCm), along with smoke tests to verify detection and segmentation functionality.
backend/go/rfdetr-cpp · high confidence
Add trellis2cpp backend for 3D generation
This change introduces a new backend in \backend/go/trellis2cpp\ that integrates the \localai-org/trellis2cpp\ C++ library (pinned to commit \2f3e6e26edbbaaf8ce93d092f16f46968a366a6a\) to enable 3D model generation. The Go implementation provides gRPC bindings to the native library, handling model resolution for the required GGUF components, CPU variant selection (AVX/AVX2/AVX512/fallback), and GLB mesh parsing for print remeshing. It includes build scripts for packaging the shared libraries and tests for model path resolution and parameter mapping.
backend/go/trellis2cpp · high confidence
Add vLLM-Omni backend for multimodal generation
Introduces a new vLLM-Omni backend that enables multimodal capabilities including text-to-image, image editing, text-to-video, image-to-video, text generation with multimodal inputs, and text-to-speech. The backend exposes a gRPC interface, handles model loading and type detection, and includes installation scripts for various build profiles (CUDA 12, CUDA 13, ROCm, and L4T aarch64) along with unit tests for service startup and image generation.
backend/python/vllm-omni · high confidence
Added NATS payload types for distributed MCP tool execution
The \core/services/mcp\ package now includes \remote.go\, which defines the request and response structures for executing and discovering MCP tools via NATS in a distributed mode. These types (\MCPToolRequest\, \MCPToolResponse\, \MCPDiscoveryRequest\, \MCPDiscoveryResponse\) enable the serialization of model configurations and tool definitions between agents and workers, supporting the underlying infrastructure for remote tool execution.
core/services/mcp · high confidence
Added distributed routing infrastructure for node identification and prefix-cache awareness
This change introduces the \pkg/distributedhdr\ package to support distributed mode by enabling two key capabilities: first, it allows the system to record and expose the specific worker node that served a request via the \X-LocalAI-Node\ response header, using a race-safe context holder to bridge the router and HTTP middleware; second, it implements prefix-cache-aware routing by attaching and matching prompt prefix-hash chains to the request context. To support the latter, a new \pkg/radixtree\ package provides a concurrent, TTL-managed radix tree that maps prompt prefixes to node identifiers, allowing the router to select the most appropriate worker based on cache recency and hit depth.
pkg/distributedhdr · high confidence
Added internal versioning and user-agent identification
The internal package now exposes version metadata (Version and Commit) and a UserAgent function that constructs a client identity string for outbound requests. This string includes the LocalAI version (if available) and the OS/architecture platform, allowing registries and galleries to distinguish between source builds and released versions in logs. Tests verify that the user-agent correctly formats these details.
internal · high confidence
Added thread-safe generic map utility
A new \pkg/xsync\ package has been introduced, providing a \SyncedMap\ type that wraps a Go map with \sync.RWMutex\ to ensure safe concurrent access. This utility exposes methods for setting, getting, deleting, and iterating over key-value pairs, along with helpers to retrieve keys, values, length, and existence checks, accompanied by Ginkgo-based unit tests.
pkg/xsync · high confidence
Added vLLM streaming and tool-parser benchmark example
A new self-contained Python script and README have been added to the examples/vllm-bench directory to measure time-to-first-token (TTFT) for the vLLM backend's streaming path with a tool parser. This benchmark helps users verify that native parser-side streaming (parser.extract\_tool\_calls\_streaming) is working correctly, specifically ensuring that plain-text responses stream progressively rather than being buffered entirely until completion.
examples/vllm-bench · high confidence
Agent pool and job execution refactored for distributed mode with cross-replica state sync
The agent pool service now supports a distributed mode backed by PostgreSQL and NATS, abstracting local and distributed operations through new \AgentConfigBackend\ and \JobPersister\ interfaces. Agent configurations and job/task states are persisted to the database in distributed mode (while remaining file-backed in standalone mode), and agent tasks are kept consistent across replicas via a NATS-backed \SyncedMap\, ensuring that task creation, updates, and deletions are visible to all nodes. The service also introduces native Prometheus metrics for agent chat runs to track completion, errors, and duration, and includes a comprehensive test suite for the new job persistence and cross-replica synchronization logic.
core/services/agentpool · high confidence
Agent service restructured with distributed mode, KB citations, and timestamp fixes
The agent service in core/services/agents has been significantly restructured to support distributed execution and improved user experience. A new EventBridge component bridges agent events between NATS and SSE connections, enabling cross-instance real-time updates. Agent chat event timestamps are now emitted in Unix milliseconds (previously nanoseconds) to prevent display errors in the React UI. Knowledge Base responses now include source citations, with a new citations.go module handling the generation of clickable source links and deduplication. The service introduces configurable Knowledge Base modes (auto\_search, tools, both) and Skills modes (prompt, tools, both) via the AgentConfig, allowing users to control how these features are exposed. A static config metadata definition (configmeta.go) replaces the previous LocalAGI-dependent implementation for the UI form. The dispatcher now supports both local (goroutine-based) and distributed (NATS queue) execution modes.
core/services/agents · high confidence
Audio format identification and WAV parsing utilities
The audio package now includes utilities to identify audio file formats and parse WAV data. The \Identify\ function detects the audio type and MIME content type from a stream, while \NormalizeAudioFile\ ensures files have the correct extension and content type. Additionally, \ParseWAV\ and \StripWAVHeader\ provide robust parsing of WAV files by walking the RIFF structure, handling non-canonical layouts and metadata chunks to prevent audio artifacts.
pkg/audio · high confidence
Authenticated downloads from registries and private hosts
Downloads now support authentication for private registries and hosts. A new credentials system allows operators to configure bearer tokens, basic auth, or custom headers via a credentials file, which are automatically applied to outbound HTTP requests and OCI registry pulls. The system respects docker config credentials as a fallback, ensures secrets are never leaked in logs or errors, and supports dynamic secret resolution for rotated credentials.
pkg/downloader · high confidence
Automated inference defaults generation with HTTP redirect refusal
A new tool in core/config/gen\_inference\_defaults fetches and validates inference parameter defaults from the Unsloth repository, remapping field names (e.g., repetition\_penalty to repeat\_penalty) and filtering to a specific set of allowed fields (temperature, top\_p, top\_k, min\_p, repeat\_penalty, presence\_penalty) before writing the result to core/config/inference\_defaults.json. The tool uses the local HTTP client which now refuses outbound redirects, ensuring that the fetch operation is hardened against redirect-based attacks or unexpected location changes.
_core/config/gen\_inference\defaults · high confidence
Automatic chat context compression to prevent context overflow
Users can now enable automatic compression of long chat histories to keep conversations within the model's context window. The new \core/services/compression\ service monitors message token counts using a conservative estimate in \pkg/tokens\ and, when a configured threshold is reached, summarizes older conversation turns into a single system message while preserving recent messages and active tool-call chains. This feature includes configurable policies for trigger ratios, tail retention, and overflow recovery, along with OpenTelemetry metrics to track compression events and ratios.
core/services/compression, pkg/tokens · high confidence
Bounded backend trace history with audio quality metrics
The backend now captures and persists a bounded history of backend operations (LLM, TTS, transcription, etc.) in the admin Traces UI. To keep the UI responsive, trace payloads are capped (defaulting to 1 MB per trace body) and oversized string values in the Data field are truncated with a marker. Audio traces now include quality metrics (RMS, peak, DC offset) and a base64-encoded WAV snippet of the first 30 seconds, with the snippet dropped if it exceeds the body cap. Traces are persisted to disk and restored on restart, and non-finite float values (like -Inf dBFS for silence) are sanitized to ensure valid JSON responses.
core/trace · high confidence
Cloud proxy adds on-path PII redaction and secure backend forwarding
The cloud proxy now intercepts outbound API traffic to allowlisted hosts, performing request-body PII redaction via a NER engine before forwarding the request to the upstream. It also introduces a new gRPC-based backend forwarding path that instruments passthrough requests for the Traces UI. To support the interception, the proxy generates and caches its own TLS CA and leaf certificates, and the outbound client is hardened to refuse HTTP redirects, preventing potential credential leakage.
core/services/cloudproxy · high confidence
Configurable copy buffer size with context cancellation support
The \pkg/xio\ package now provides a \Copy\ function that allows users to configure the internal buffer size via the \WithBufferSize\ option, defaulting to 1 MiB instead of the standard library's smaller default. Additionally, this copy operation respects context cancellation, checking for cancellation before each read operation to allow for graceful interruption of data transfer.
pkg/xio · high confidence
Experimental tinygrad backend for LLMs and multimodal tasks
A new experimental backend powered by tinygrad has been added to the Python backends directory. It enables LLM text generation (supporting models like Qwen3, Llama 3.x, and GLM-4 via GGUF and HuggingFace safetensors), embeddings, Stable Diffusion 1.x image generation, and Whisper audio processing. The backend includes native tool-call extraction for various model families (Hermes, Llama 3, Mistral, Qwen 3) and handles device selection for CUDA, HIP, and CPU (via LLVM) environments.
backend/python/tinygrad · high confidence
Federated P2P mode with load balancing and worker targeting
The core/p2p package now includes a new federated server implementation that allows users to distribute requests across multiple P2P nodes. This change introduces a \FederatedServer\ that supports two routing strategies: random selection for basic distribution or least-connection load balancing to optimize resource usage. Users can also explicitly target a specific worker node via configuration. The implementation handles node discovery, connection proxying, and request tracking to ensure efficient load distribution within the P2P network.
core/p2p · high confidence
Formal verification specs for realtime and model-loader lifecycles
Added authoritative FizzBee formal specifications and a pinned FizzBee v0.5.2 toolchain to the \formal-verification\ directory. The specs cover the realtime API state machines (connection, turn detection, response coordination, conversation compaction, and TTS pipeline) and the shared model-loader shutdown lifecycle. These models serve as the source of truth for design verification, ensuring that implementation changes are checked against invariants such as single-flight compaction, idempotent teardown, and coupled speech/turn states.
formal-verification · high confidence
In-process LocalAI client for MCP tools
The LocalAI Assistant chat modality now uses an in-process client to call LocalAI services directly, avoiding the overhead and complexity of an HTTP loopback to the same process. This new client exposes tools for managing model aliases, scheduling configurations, gallery searches, and PII filtering, providing a more efficient and integrated experience for users interacting with the LocalAI Assistant.
pkg/mcp/localaitools/inproc · high confidence
Intelligent backend selection based on system capabilities
The system now detects hardware capabilities (GPU vendor, VRAM, CUDA/ROCm versions) to automatically select the most suitable backend engine (e.g., vLLM, SGLang, llama-cpp) and build tags. This ensures that models are installed and run using the optimal acceleration path for the host, such as preferring CUDA builds on NVIDIA GPUs or falling back to CPU-only builds when VRAM is insufficient. It also introduces a preference system that prioritizes specific engine types based on detected hardware, improving performance and compatibility out-of-the-box.
pkg/system · high confidence
Introduce C++ ds4 backend with DSML streaming parser and disk KV cache
Added a new C++ backend for the ds4 (DeepSeek V4 Flash) model that implements a streaming DSML parser to correctly handle thinking-mode reasoning and tool-call structures, a renderer for tool-call and manifest formatting, and a disk-based KV cache for session persistence. The backend is built via CMake and Makefile targets (grpc-server and ds4-worker) supporting CPU, CUDA, and Metal execution paths, and includes unit tests for the parser and generation-limit logic.
backend/cpp/ds4 · high confidence
Introduce CrispASR backend with multi-architecture support and word-level timestamps
The new CrispASR backend in \backend/go/crispasr\ provides a unified Go service for both automatic speech recognition (ASR) and text-to-speech (TTS). It bundles a C++ shim (\crispasr\_shim\) that links against the CrispASR library and GGML, enabling support for a wide range of models including parakeet, canary, qwen3, f5-tts, and piper. The backend is built with multi-architecture optimizations (AVX, AVX2, AVX512, fallback) and includes word-level timestamp accessors for supported models. It also handles TTS voice selection, voice cloning via reference WAVs, and correctly writes WAV files at the native sample rate for piper voices.
backend/go/crispasr · high confidence
Introduce Hugging Face model artifact materialization with robust concurrency and resumption
The \pkg/modelartifacts\ package now handles downloading and caching Hugging Face model snapshots. It supports bounded parallel downloads to speed up large repositories, while ensuring deterministic manifest ordering regardless of download completion order. To prevent data corruption on shared network filesystems (like CIFS) where file locks may fail, each download process stages files in a unique, writer-specific temporary tree; these are only promoted to the final location once fully verified. This design also enables safe resumption of interrupted downloads by allowing a new process to adopt and continue a previous writer's partial tree, and includes a background sweep to reclaim abandoned staging trees.
pkg/modelartifacts · high confidence
Introduce Kokoros gRPC backend with optional Bearer token authentication
Adds a new Rust-based gRPC backend for the Kokoros TTS engine, exposing standard backend operations such as health checks, model loading (with configurable language and speed), and text-to-speech inference that writes output as 16-bit PCM WAV files. The server listens on a configurable address (default localhost:50051) and supports optional Bearer token authentication via the LOCALAI\_GRPC\_AUTH\_TOKEN environment variable, intercepting requests to validate the authorization header before processing.
backend/rust/kokoros · high confidence
Introduce LocalAI Assistant MCP server with comprehensive admin tooling
This change adds the \pkg/mcp/localaitools\ package, which exposes LocalAI's administrative and management capabilities as a Model Context Protocol (MCP) server. It provides a unified \LocalAIClient\ interface that powers both in-process chat integration and a standalone CLI, enabling LLMs to manage models (install, delete, configure, alias), backends, distributed scheduling, VRAM budgets, voice profiles, branding, and system state. The package includes a full set of Data Transfer Objects (DTOs) for the LLM-visible wire format, embedded skill prompts for system instructions, and a test suite ensuring parity between the in-process and HTTP API client implementations.
pkg/mcp/localaitools · high confidence
Introduce LocalAI desktop launcher with auto-start, secure downloads, and scoped paths
This change adds the LocalAI desktop launcher (cmd/launcher/internal), a Fyne-based application that manages the LocalAI server lifecycle. The launcher automatically starts the server on launch to ensure the service is available immediately, and scopes generated-content and upload directories to the user's data folder to prevent permission errors on shared systems. It includes a robust release manager that bypasses GitHub API rate limits by using release redirects, supports resumable downloads with rate limiting, and enforces security by refusing HTTP redirects on outbound clients. The UI provides configuration for models/backends paths, environment variables, and logging, along with system tray integration for start/stop controls and update notifications.
cmd/launcher/internal · high confidence
Introduce LocalAI desktop launcher with system tray and welcome screen
The cmd/launcher directory now contains a new Fyne-based desktop application for LocalAI. This launcher provides a native UI that minimizes to a system tray icon instead of closing, handles graceful shutdown via signal handling, and displays a welcome window when the configuration option is enabled. The application is configured with the LocalAI branding, including a specific app ID (com.localai.launcher) and icon.
cmd/launcher · high confidence
Introduce LocalVQE backend for acoustic echo cancellation
Adds a new LocalVQE backend to the audio processing pipeline, enabling acoustic echo cancellation and noise gate features. This change introduces the Go-based gRPC server (golocalvqe.go, main.go) that binds to the upstream LocalVQE C++ library via purego, along with build infrastructure (Makefile, package.sh, run.sh) to compile and bundle the required shared libraries (liblocalvqe.so, libggml-\*.so) for Linux and macOS. The backend supports configuration via model options for noise gate thresholds and backend selection, and includes a test suite (localvqe\_test.go) to verify WAV parsing and option handling.
backend/go/localvqe · high confidence
Introduce MLX backend with thread-safe prompt caching
Adds a new MLX inference backend to the Python backend suite, enabling users to run models using Apple's MLX framework. This backend includes a thread-safe LRU prompt cache to improve performance by reusing KV cache states across concurrent requests, and supports standard sampling parameters including min\_p and top\_k. The implementation provides gRPC service methods for health checks, model loading, and text prediction, along with integration tests to verify concurrent request safety and cache reuse.
backend/python/mlx · high confidence
Introduce Open Responses API with WebSocket support, distributed store, and MCP integration
Adds a new /v1/responses endpoint implementing the Open Responses specification, including a WebSocket transport for streaming events. The implementation introduces a response store with configurable TTL, bounded resume buffers (8192 events / 64 MiB), and cross-replica synchronization via NATS to make responses visible and cancellable across distributed replicas. It also supports MCP tool injection from metadata, handles previous\_response\_id chaining for conversation history, and fixes message conversion to populate both Content and StringContent fields.
core/http/endpoints/openresponses · high confidence
Introduce P2P network explorer with local token database
The core/explorer package now includes a local JSON-based database (database.go) for persisting P2P network tokens along with their name, description, and cluster metadata, protected by file locking for safe concurrent access. This storage layer supports the new DiscoveryServer (discovery.go), which automatically discovers and syncs installed models across instances by connecting to P2P networks, tracking cluster workers, and removing tokens that exceed a configurable failure threshold. Comprehensive tests (database\_test.go, explorer\_suite\_test.go) verify the database's add, retrieve, delete, and persistence capabilities.
core/explorer · high confidence
Introduce Python backend template with standardized build and runtime scripts
A new common template for Python-based backends has been added, providing a standardized set of shell scripts (install.sh, run.sh, test.sh, protogen.sh) and a Makefile to handle backend lifecycle operations. This template sources shared utilities from the common library to manage dependency installation (including specific workarounds for Intel pip index issues), gRPC code generation, backend startup, and unit testing, ensuring consistent behavior across different Python backend implementations.
backend/python/common/template · high confidence
Introduce Qwen-TTS backend for text-to-speech generation
Adds a new gRPC-based backend for the Qwen3-TTS model, enabling text-to-speech synthesis. The implementation includes a Python service (backend.py) that handles model loading, device detection (CUDA/MPS/CPU), and voice configuration via options. It also provides supporting scripts for installation (install.sh), running (run.sh), and testing (test.sh, test.py) to ensure the service operates correctly.
backend/python/qwen-tts · high confidence
Introduce TurboQuant llama.cpp fork backend
Adds a new TurboQuant backend (a llama.cpp fork) alongside the existing standard llama.cpp backend. This includes build infrastructure (Makefile, patch scripts) to compile the fork with specific CPU variants (AVX2, AVX512, etc.) and a single-build CPU-all-variants target, as well as packaging and runtime scripts to bundle the binary and required GPU/CPU libraries.
backend/cpp/turboquant · high confidence
Introduce advisory lock service with PostgreSQL and SQLite support
Added a new advisory lock service in core/services/advisorylock that provides cross-process coordination for PostgreSQL (using pg\_advisory\_lock) and in-process serialization for SQLite. The implementation includes a leader election loop (RunLeaderLoop) for periodic tasks, context-aware lock acquisition with timeout handling to prevent hangs, and a suite of tests covering both dialects and key generation logic.
core/services/advisorylock · high confidence
Introduce biometrics UI components for face and voice analysis
The biometrics section of the React UI now includes a suite of new components to support face and voice identification workflows. Users can capture or upload media via the MediaInput component, which supports file uploads, clipboard pasting for images, and live webcam/microphone capture. Enrollment management is handled by EnrollmentList, allowing users to view and delete enrolled subjects. Visual feedback is provided through BoundingBoxCanvas for face detection overlays, DistributionBars for classification probabilities, MatchGauge for verification distance metrics, and EmbeddingInspector for visualizing raw vector data. Audio analysis is supported by WaveformStrip for displaying audio peaks and segments.
core/http/react-ui/src/components/biometrics · high confidence
Introduce cloud-proxy backend for OpenAI and Anthropic integration
A new LocalAI backend binary, cloud-proxy, is added to forward model traffic to external HTTP providers. It supports two modes: passthrough, which preserves the client wire format end-to-end, and translate, which converts internal proto requests to the provider's wire format for OpenAI and Anthropic. The translate mode handles specific provider requirements, such as hoisting system messages, mapping tool calls, and injecting Anthropic prompt-cache breakpoints. The backend also enforces security by refusing HTTP redirects on outbound clients to prevent API key leakage.
backend/go/cloud-proxy · high confidence
Introduce comprehensive HTTP authentication and authorization system
This change introduces a new authentication and authorization framework for the HTTP layer. It adds support for user accounts, session-based login (including GitHub and OIDC providers), and API key management (generation, validation, revocation, and expiration). A new middleware enforces authentication by default on all routes, with exemptions for public endpoints and legacy API keys. It also implements a feature-gating system that allows administrators to enable or disable specific capabilities (like chat, images, audio, etc.) per user or role, and adds CSRF protection for browser-based sessions. The system uses a database (PostgreSQL or SQLite) to store users, sessions, API keys, permissions, and usage records.
core/http/auth · high confidence
Introduce distributed job scheduling with NATS and PostgreSQL persistence
The jobs service now supports distributed execution across multiple instances. Jobs and tasks are persisted in PostgreSQL using GORM models (TaskRecord, JobRecord) with JSON columns for complex parameters, and schema migrations are protected by advisory locks to prevent race conditions. A Dispatcher component coordinates work by publishing job events to a NATS queue, allowing agent workers to consume and process jobs independently of the frontend. The system also includes an SSE bridge for real-time job progress streaming, cron-based task scheduling with time-based due checks, and comprehensive round-trip conversion and persistence tests.
core/services/jobs · high confidence
Introduce fish-speech TTS backend with platform-specific build fixes
Adds a new fish-speech backend for text-to-speech, including a gRPC server implementation, installation scripts, and integration tests. The backend includes specific fixes to ensure compatibility across platforms: it allows the \invalid\_reference\_casting\ lint to enable building on macOS (Darwin), preserves ROCm PyTorch dependencies instead of replacing them with CUDA wheels, patches PyTorch/torchaudio versions for CUDA 13 on ARM64, and sets the \TRITON\_PTXAS\_PATH\ to support CUDA toolkit usage. It also replaces \torchaudio\ with \soundfile\ for reference audio loading to support Linux ARM64 environments lacking \torchcodec\ wheels.
backend/python/fish-speech · high confidence
Introduce grammar-based function calling with crash protection
This change adds a new grammar generation system in \pkg/functions/grammars\ to support structured function calling, including a standard JSON schema converter and a specific converter for the llama3.1 schema format. To prevent process crashes from malformed inputs, the implementation enforces a maximum recursion depth of 256 and detects cyclic \$ref\ references in JSON schemas, returning a clear error instead of exhausting the stack. The update also includes a test suite to verify grammar generation and the robustness of these safety guards.
pkg/functions/grammars · high confidence
Introduce ik-llama-cpp gRPC backend with multimodal support
Adds a new C++ gRPC backend server for the ik\_llama.cpp fork, enabling users to run inference via the LocalAI gRPC protocol. The implementation includes build scripts (CMakeLists.txt, Makefile) and a packaging script to handle dependencies and GPU libraries, while the server code integrates the mtmd library to support multimodal inputs (images/audio) and applies patches to ensure compatibility with the upstream ik\_llama.cpp API.
backend/cpp/ik-llama-cpp · high confidence
Introduce local-store as an in-process vector search backend
A new local-store backend has been added, providing an in-process vector storage solution exposed via gRPC. This component implements a sorted parallel-slice store with binary-search lookups and linear-scan top-K retrieval, supporting Set, Get, Delete, and Find operations. It includes strict validation for key dimensions and namespace prefixes to prevent cross-contamination, and features an optimized fast path for normalized vectors. The package includes build scripts, a clean target, and a comprehensive test suite covering edge cases like dimension mismatches and partial deletions.
backend/go/local-store · high confidence
Introduce managed voice cloning profiles with multi-reference support
Users can now create and manage persistent voice-cloning profiles that store reference audio clips and transcripts for use in text-to-speech synthesis. This change adds a new \voiceprofile\ service that handles the creation, storage, and leasing of these profiles, enforcing constraints such as a 50 MiB size limit, PCM WAV format, and mandatory consent confirmation. The implementation supports multi-reference personalities, allowing profiles to contain multiple ordered audio/transcript pairs, and ensures data integrity through atomic operations and secure file permissions.
core/services/voiceprofile · high confidence
Introduce native C++ audio.cpp backend with gRPC server
Adds a new native C++ backend for audio.cpp, built as a gRPC server that handles audio input/output, capability routing, and model loading. The implementation includes a pinned version of the audio.cpp repository, a CMake build system with support for CUDA, ROCm, Vulkan, and Metal backends, and specific handling for protobuf ABI compatibility. The backend provides functions for reading and writing WAV files, converting between audio formats, and routing requests based on model capabilities.
backend/cpp/audio-cpp · high confidence
Introduce parakeet-cpp backend with dynamic batching and multilingual support
Added a new parakeet-cpp backend for audio transcription that supports dynamic batching of concurrent requests to improve GPU throughput, multilingual streaming with per-request language targeting, and real-time utterance boundary detection for accurate turn-taking in live feeds.
backend/go/parakeet-cpp · high confidence
Introduce rerankers gRPC backend with top\_n support
Adds a new Python-based gRPC backend for reranking models, including the server implementation, installation/run scripts, and unit tests. The backend exposes Health, LoadModel, and Rerank services, allowing users to load reranker models and retrieve ranked document lists. It supports the top\_n parameter to limit the number of returned results, and correctly handles cases where top\_n is omitted or set to zero by returning all ranked documents.
backend/python/rerankers · high confidence
Introduce sglang as a new gRPC backend
Adds a new Python-based gRPC backend for sglang, mirroring the existing vLLM backend structure. The implementation includes a \backend.py\ servicer that wraps sglang's async Engine API, supporting engine argument validation against \ServerArgs\ (handling both dataclass and msgspec-based configurations in sglang \>= 0.5.20), tool call parsing, and reasoning content extraction. It also provides build, run, and test scripts (\install.sh\, \run.sh\, \test.sh\) to handle dependencies, environment setup (including CPU-specific libraries like libiomp5), and unit testing for the backend's helper functions.
backend/python/sglang · high confidence
Introduce stablediffusion-ggml backend with VAE tiling and video generation support
The stablediffusion-ggml backend is now available as a standalone Go-based component, replacing the previous integrated approach. This backend leverages the leejet/stable-diffusion.cpp library (updated to version c92d73c) and exposes new capabilities including video generation (GenVideo) and configurable VAE tiling (allowing users to set tile sizes and overlap to manage memory usage). The build system supports multiple CPU instruction set variants (AVX, AVX2, AVX512) and GPU backends (CUDA, ROCm, Metal, Vulkan, SYCL), with automatic runtime selection of the appropriate library variant based on the host CPU. The backend is packaged with its own run script and library dependencies, ensuring it can operate independently.
backend/go/stablediffusion-ggml · high confidence
Introduce standalone Go-based Whisper backend with streaming and VAD
The whisper backend is now a separate Go service (backend/go/whisper) that wraps whisper.cpp via purego instead of cgo. This new backend adds streaming transcription (emitting segment deltas progressively), voice activity detection (VAD), and client-request cancellation support. It also ships multiple CPU-optimized shared libraries (AVX, AVX2, AVX512, and a fallback) on Linux, with Darwin using a single fallback variant. The service listens on a configurable address (honoring both the -addr flag and a positional argument) and includes tests for cancellation and streaming behavior.
backend/go/whisper · high confidence
Introduce standalone Piper TTS backend
Adds a new, standalone Piper text-to-speech backend that can be built and packaged independently from the main LocalAI binary. This change introduces a dedicated Go implementation (main.go, piper.go) that exposes a gRPC service for TTS operations, along with a Makefile to manage the go-piper dependency and espeak-ng data, and packaging scripts to bundle the binary with its required shared libraries for runtime execution.
backend/go/piper · high confidence
Introduce structured gRPC error signals for model and capability states
Added a new \pkg/grpc/grpcerrors\ package that defines canonical error types and detection helpers for cross-language gRPC communication. This allows the router to reliably distinguish between specific backend states—such as a model not being loaded (\ModelNotLoaded\), a model identity mismatch (\ModelMismatch\), or unsupported transcription capabilities (\LiveTranscriptionUnsupported\, \StreamTranscriptionUnsupported\)—by using typed gRPC status codes and sentinels rather than fragile string matching. This ensures that routing decisions, such as retrying or dropping stale replicas, are based on accurate semantic signals from backends.
pkg/grpc/grpcerrors · high confidence
Introduce swappable face-recognition registry with replay-safe enrollment
The face-recognition service now uses a new Registry interface backed by LocalAI's in-memory gRPC vector store, enabling 1:N identification, embedding registration, and deletion. The implementation derives enrollment IDs deterministically from the embedding vector, so re-registering the same face (or replaying saved enrollments after a restart) preserves the original identity instead of creating duplicates. The registry is designed to be swappable; a persistent PostgreSQL/pgvector backend is planned for production use, but the current in-memory store loses data on restart.
core/services/facerecognition · high confidence
Introduce vLLM Python backend with multi-platform support and testing
This change adds the vLLM Python backend to the system, providing a new inference path that supports CPU, CUDA, ROCm, Intel XPU, and Apple Silicon (Metal/MLX) hardware profiles. The backend includes a gRPC servicer (backend.py) that handles model loading, text generation, and tool-call parsing, along with installation scripts (install.sh) that manage environment setup and dependency resolution for each platform. A packaging script (package.sh) bundles necessary system libraries and a C++ toolchain for CPU profiles to ensure portability. The entry also includes a comprehensive test suite (test.py) verifying gRPC service startup, model loading, sampling parameters, and message parsing, alongside standard build/run/test Makefile targets.
backend/python/vllm · high confidence
Introduce vector store client for key-value operations
Added a new gRPC client in pkg/store that enables setting, getting, deleting, and finding similar keys in a vector store backend. The client handles serialization of float32 keys and byte values, using a 'store://' namespace prefix to prevent model autoloaders from incorrectly binding vector store loads to LLM backends.
pkg/store · high confidence
Introduce vllm-cpp backend for local inference
Adds a new LocalAI backend that wraps the vllm.cpp engine via a stable C ABI, enabling text generation and MiniMax-H3 video+audio generation without Python at inference time. The backend supports GGUF and safetensors models, implements engine-side chat templating with tool calling, and includes a build system that compiles the engine with CUDA, Metal (with an optional MLX GEMM provider on macOS), or Vulkan backends.
backend/go/vllm-cpp · high confidence
Introduce voice recognition registry for speaker identification
Added a new \voicerecognition\ service package that provides a swappable registry for storing speaker embeddings and performing 1:N voice identification. The implementation uses LocalAI's in-memory gRPC backend, allowing users to register speaker profiles, identify speakers from audio probes via similarity search, and remove registrations. This service mirrors the existing face recognition architecture to enable a shared generic biometric registry in the future, though current data is lost on restart.
core/services/voicerecognition · high confidence
Introduces a reflection-based model configuration metadata system for the UI editor
A new \core/config/meta\ package automatically generates structured metadata for the \ModelConfig\ struct by walking its fields via reflection, inferring UI types (e.g., int, bool, string-list), and grouping them into ordered sections (General, LLM, Parameters, etc.). This metadata is enriched by a registry of overrides that provide human-readable labels, descriptions, specific UI components (like \select\ or \code-editor\), and constraints (min/max/step). The system supports dynamic autocomplete providers for fields like backends and models, and static option lists for things like quantization and pipeline types. This change provides the foundational data structure that powers the interactive model configuration editor with autocomplete and proper field validation.
core/config/meta · high confidence
Introduces prefix-cache-aware routing for distributed mode
Adds a new prefix-cache routing subsystem that extracts a deterministic hash chain from the prompt head and tracks which replica served which prefix in an in-memory radix tree. The system provides a pluggable routing pipeline with configurable filters, weighted scorers, and pickers, allowing requests to be routed to the replica holding the longest matching prefix (hot match) while falling back to load-balanced cold placement when no match exists or the matched replica is unavailable. It includes configuration for thresholds, TTLs, and scorer weights, a pressure-tracking mechanism to signal cache saturation for autoscaling, and a reported-index provider that can be populated by backend residency events instead of request observations.
core/services/nodes/prefixcache · high confidence
Introduction of base gRPC backend classes for model backends
New base classes (\Base\ and \SingleThread\) have been added to the gRPC package to serve as the foundation for all model backends. The \Base\ class implements the full gRPC service interface with stub methods that return 'unimplemented' errors, while \SingleThread\ extends this to provide mutex-based locking for backends that do not support concurrent requests. This change establishes a standardized structure for backend implementations to inherit from, ensuring consistent handling of model loading, prediction, streaming, and status reporting across the system.
pkg/grpc/base · high confidence
Model configuration administration service with distributed revision consistency
A new \modeladmin\ service has been introduced to manage the configuration of installed models, providing a unified API for reading, patching, and mutating model YAML files (including enable/disable and pin/unpin actions) while ensuring data integrity through atomic file writes and mutation rollbacks. To support distributed deployments, the service implements a revision lifecycle that publishes authoritative configuration hashes to the node registry, allowing replicas to reconcile state via cache-invalidation events and automatically resync stored revisions at startup to prevent models from becoming unroutable due to revision drift.
core/services/modeladmin · high confidence
NATS messaging layer adds distributed install progress, cancellation, and security features
The core messaging service now supports distributed deployments with a new BackendInstallProgressEvent that publishes worker install phases (resolving, downloading, extracting, starting) to NATS, allowing the UI to show real-time progress. It introduces a CancelRegistry for managing cancellable operations and new NATS subjects for cancelling in-flight responses and backend upgrades. Security is enhanced with support for NATS JWT authentication (including dynamic providers) and TLS/mTLS configuration. The messaging client also validates subscription permissions synchronously to prevent silent failures and sanitizes NATS subject tokens to handle reserved characters safely.
core/services/messaging · high confidence
New C++ bindings for Stable Diffusion generation
The backend now includes a new C++ interface (gosd.cpp/gosd.h) that exposes Stable Diffusion image and video generation capabilities to the Go layer. This addition introduces functions for creating and configuring image generation parameters (prompts, dimensions, seed), loading models, and generating images, as well as video generation parameters and generation functions. It also implements embedding and LoRA discovery from directories and parsing of LoRA tags from prompts.
backend/go/stablediffusion-ggml/cpp · high confidence
New C++ privacy-filter backend for PII/NER detection
A new standalone C++ gRPC backend has been added to handle Personally Identifiable Information (PII) and Named Entity Recognition (NER) filtering. This backend wraps the privacy-filter.cpp GGML engine, replacing the previous llama.cpp-patched path for this specific model family. It supports CPU, CUDA, and Vulkan execution, manages model loading and token classification via gRPC, and includes logic to reject requests routed to the wrong model identity in distributed deployments.
backend/cpp/privacy-filter · high confidence
New C++/ggml backend for Depth Anything 3 depth estimation
Adds a new C++/ggml-based backend for Depth Anything 3, enabling users to run depth estimation and camera pose inference locally via a gRPC service. The backend supports multiple CPU instruction set variants (AVX, AVX2, AVX512, fallback) and includes a nested metric model mode for higher-precision depth maps. It exposes two primary capabilities: generating min-max-normalized grayscale depth images and predicting depth statistics with camera extrinsics/intrinsics. The implementation uses purego to dynamically load the C++ shared library and handles model loading via a two-file pair (anyview + metric) when specified.
backend/go/depth-anything-cpp · high confidence
New CLI commands for standalone agents, distributed workers, backend management, and benchmarking
The CLI now includes several new commands to manage and operate LocalAI components outside the main server process. Users can run standalone agents via \localai agent run\ (loading configuration from a JSON file or the agent pool registry) and list registered agents with \localai agent list\. A new \localai agent-worker\ command allows nodes to register with the frontend, authenticate via NATS JWT/mTLS, and execute agent jobs in distributed mode. Backend management is now accessible through \localai backends\ (list, install, uninstall, upgrade), supporting gallery-based installations and integrity checks. Additionally, a \localai benchmark\ command enables users to measure text inference performance (latency, throughput) against a running LocalAI server, with output available in table or JSON format.
core/cli · high confidence
New Go-based face-detect backend replaces Python implementation
A new Go backend for face detection and verification has been added to replace the previous Python-based insightface implementation. This change introduces a self-contained gRPC service that loads the upstream face-detect.cpp library via purego, providing face detection, embedding generation, and biometric verification capabilities. The backend supports various hardware acceleration backends (CPU, CUDA, ROCm, Vulkan, Metal) and includes a build system to compile the C++ library and package it with necessary runtime dependencies.
backend/go/face-detect · high confidence
New HTTP client for LocalAI MCP tools
Added a new \httpapi\ package that provides a REST-based client for the LocalAI MCP server, enabling it to control a remote LocalAI instance over stdio. This client implements the \LocalAIClient\ interface by mapping MCP tool calls to LocalAI's admin REST endpoints, covering model management (listing, installing, deleting), gallery search, scheduling administration, and other administrative functions.
pkg/mcp/localaitools/httpapi · high confidence
New HTTP middleware layer for routing, compression, and security
This change introduces a comprehensive set of new HTTP middleware components in the core request pipeline. It adds an admission control middleware that enforces per-model concurrency limits and returns HTTP 503 with a Retry-After header when capacity is reached. A new compression middleware enables gzip responses for static assets and JSON APIs while explicitly excluding streaming endpoints (SSE, WebSockets) to prevent latency issues. The pipeline now includes robust base URL and path prefix handling that correctly resolves reverse-proxy configurations (including Caddy's path stripping) while validating the X-Forwarded-Prefix header to prevent open-redirect vulnerabilities. Additionally, a new node header middleware stamps the X-LocalAI-Node response header with the ID of the distributed worker that served the request, using a per-request holder to ensure correct attribution in concurrent multi-replica scenarios. Context compression middleware is also added to transform chat requests before they reach the backend.
core/http/middleware · high confidence
New HTTP route registration files for agents, Anthropic, auth, and distributed cluster capabilities
The \core/http/routes\ package now includes dedicated registration files for several new and refactored capabilities. \agents.go\ registers the full agent pool API (CRUD, chat, skills, collections) behind a readiness middleware. \anthropic.go\ implements the Anthropic Messages API compatibility layer with request context and PII filtering. \auth.go\ and its tests introduce a comprehensive authentication system with local/OIDC/GitHub providers, session management, and strict field-tampering protection on user profiles. \cluster\_capabilities.go\ and \cluster\_memory.go\ provide distributed-mode logic for backend discovery and model sizing, allowing controllers to query worker node capabilities and memory budgets instead of relying solely on local hardware.
core/http/routes · high confidence
New MLX Distributed Inference Backend
Adds a new MLX distributed inference backend for LocalAI, enabling multi-GPU inference via pipeline parallelism. The backend supports two distributed modes—Ring (using MLX hostfiles) and Jaccl (using RDMA devices)—and includes a coordinator to broadcast commands and tokens across ranks. It also features a thread-safe LRU prompt cache for single-node efficiency and auto-parallel sharding of model layers across available devices.
backend/python/mlx-distributed · high confidence
New NVIDIA NeMo-Speech.cpp backend for speech processing
Added a new backend implementation for NVIDIA NeMo-Speech.cpp, enabling local speech recognition (ASR), speaker diarization, text-to-speech (TTS), and neural machine translation (NMT). This change introduces the Go bindings (abi.go), build infrastructure (Makefile), and core logic (asr.go, asr\_stream.go) required to integrate the C++ library, along with comprehensive test coverage (abi\_test.go, asr\_test.go, asr\_stream\_test.go) to validate the ABI and streaming behavior.
backend/go/nemo-speech-cpp · high confidence
New Opus audio codec backend for WebRTC
A new Opus audio codec backend has been added to the Go backend directory, enabling audio encoding and decoding for the WebRTC feature. This backend implements a gRPC service that wraps the libopus library (via a C shim and purego) to handle PCM-to-Opus encoding and Opus-to-PCM decoding. It includes logic to refuse loading foreign models, ensuring it is only used for its specific audio codec role, and supports session-based decoder caching for stateful decoding. The package includes build scripts and packaging logic to bundle the necessary shared libraries for both Linux and macOS.
backend/go/opus · high confidence
New PEG-based parser for chat messages and tool calls
A new PEG (Parsing Expression Grammar) parser has been added to the functions package to handle chat message parsing. This parser supports extracting reasoning content, plain text, and structured tool calls from model responses. It includes a fluent builder API for defining grammar rules and specific support for parsing tool calls in various formats, including OpenAI-style JSON, Cohere-style actions, and function-as-key JSON structures. The implementation also handles partial input streams and provides utilities for normalizing quotes to valid JSON.
pkg/functions/peg · high confidence
New Qwen3-TTS C++ backend with voice cloning and streaming
A new text-to-speech backend is available for Qwen3-TTS GGUF models, powered by the qwentts.cpp library. This backend supports 24 kHz speech generation, streaming audio output, named speakers, voice design, and reference-audio voice cloning. It includes a Go wrapper and C++ bridge, with build scripts that fetch the upstream library (held at commit 35ebe537 to avoid a synthesis hang) and compile CPU variants (AVX, AVX2, AVX512, fallback) and GPU backends (CUDA, HIP, Metal, Vulkan, SYCL). Users can install models like qwen3-tts-cpp and configure voice cloning via the tts.audio\_path setting or per-request voice references.
backend/go/qwen3-tts-cpp · high confidence
New React UI component library for the frontend
The frontend has been rebuilt in React, introducing a new set of UI components in the \core/http/react-ui/src/components\ directory. This includes a kebab-menu ActionMenu for row actions, an AmbiguityAlert to resolve import backend conflicts inline, and specialized viewers for 3D animations and code artifacts. The update also brings a ChatsMenu for conversation management, a ClientMCPDropdown for browser-side MCP server configuration, and a ConfigFieldRenderer that dynamically renders complex configuration fields using CodeMirror and other primitives.
core/http/react-ui/src/components · high confidence
New React UI hooks for chat, media, and distributed mode
The React UI now includes a suite of new hooks in the \src/hooks\ directory to support the frontend migration and new capabilities. \useChat\ and \useAgentChat\ manage chat history and conversation state with local storage persistence, while \use3DHistory\ and \useMediaHistory\ handle media generation records, with 3D history using IndexedDB to support large GLB blobs. \useMCPClient\ enables Model Context Protocol tool integration, and \useDistributedMode\ probes the backend to detect cluster topology. Additional hooks like \useAudioPeaks\ for waveform visualization, \usePolling\ for visibility-aware data fetching, and \useCodeMirror\ for editor integration provide the foundation for the new UI features.
core/http/react-ui/src/hooks · high confidence
New React-based UI shell with split-pane navigation and audio visualizations
The frontend has been replaced with a new React application that introduces a split-pane layout for resource management (Models, Backends, etc.), replacing the previous table-based views with a scannable rail and detail pane. The navigation is now organized into 'Build' and 'Operate' consoles with a collapsible secondary rail, and the app includes a mobile header with hamburger navigation. New audio capabilities include a waveform player with click-to-seek and a spectrogram visualization for audio transforms. The UI also features a new footer, mobile responsiveness, and a robust routing system with code-splitting and automatic chunk-reload on deployment.
core/http/react-ui/src · high confidence
New React-based admin pages for account, activity, and agent management
The admin interface now includes dedicated React pages for managing user accounts (profile and security settings), viewing system activity (model/backend installs and cluster operations), and administering agents (creating tasks, viewing job details, and monitoring agent status). These pages replace or supplement previous UI implementations with a modern React component structure, supporting features like SSE-based agent chat, job history filtering, and agent task execution.
core/http/react-ui/src/pages · high confidence
New Sherpa-ONNX backend for speech recognition, voice activity detection, and text-to-speech
A new Sherpa-ONNX backend has been added to the system, providing a unified Go-based implementation for offline and online speech recognition (ASR), voice activity detection (VAD), and text-to-speech (TTS). This backend supports multiple model families, including Whisper, Paraformer, SenseVoice, and Omnilingual for transcription, Silero for VAD, and VITS, Piper, and Kokoro for TTS. It also includes offline speaker diarization capabilities. The implementation uses purego bindings to interface with the Sherpa-ONNX C API, allowing for efficient execution without cgo overhead, and is packaged with the necessary shared libraries for deployment.
backend/go/sherpa-onnx · high confidence
New WebUI static assets and chat functionality
The WebUI now includes a new set of static assets to support a redesigned interface and enhanced chat capabilities. This adds a comprehensive CSS animation system (\animations.css\) for UI transitions, a component style library (\components.css\) defining buttons, cards, and inputs, and a new SVG favicon. Additionally, a new \chat.js\ file implements the client-side logic for the chat interface, including features for managing chat history, auto-saving conversations to local storage, and handling streaming responses.
core/http/static · high confidence
New agent management and collection endpoints with per-user isolation
This change introduces a new set of HTTP endpoints for managing AI agents, agent jobs, skills, and vector collections, all located in the \core/http/endpoints/localai\ package. The implementation adds handlers for creating, listing, updating, and deleting agents (\agents.go\), managing agent tasks and job executions (\agent\_jobs.go\), handling skill definitions (\agent\_skills.go\), and performing CRUD operations on vector collections including uploads and searches (\agent\_collections.go\). A key behavioral addition is the \AgentResponsesInterceptor\ (\agent\_responses.go\), which routes \/v1/responses\ requests to local or distributed agents based on the model name. The endpoints enforce strict per-user data isolation, allowing admins and specific service accounts to impersonate users or aggregate data across all users via \?all\_users=true\ and \?user\_id=\ query parameters. The code also includes a fix for URL-decoding path parameters to handle legacy API key prefixes in collection names, and adds corresponding unit tests to verify the isolation and decoding logic.
core/http/endpoints/localai · high confidence
New application core with distributed, P2P, and PII middleware support
The core/application package has been rewritten to introduce a centralized Application struct that manages the full lifecycle of new capabilities. This includes wired support for distributed mode (NATS, S3, node registry), P2P service discovery for federated and MLX workers, and a cloudproxy MITM listener for PII filtering. The new core also introduces an agent job service, a config file watcher for runtime settings, and a robust readiness probe that ensures the application is fully ready before accepting traffic.
core/application · high confidence
New audio conversion and resampling utilities
Added new helper functions in the sound package to convert between float32 and int16 PCM formats, including clamping for float-to-int16 conversion, and to resample int16 audio buffers using linear interpolation. These utilities support changing sample rates (e.g., 48kHz to 16kHz) and include RMS calculation for audio analysis.
pkg/sound · high confidence
New backend and control-plane database monitoring metrics
The monitoring service now exposes new observability capabilities for backend stability and distributed cluster health. It introduces a BackendMonitorService that samples local backend process resource usage (CPU, memory) via gopsutil, providing a fallback when RPC status checks fail. Additionally, it registers OpenTelemetry gauges for the PostgreSQL control-plane database, tracking the oldest transaction snapshot age, the longest open transaction duration, and the dead-to-live tuple ratio on registry tables to help detect vacuum lag and database wedging before they cause outages.
core/services/monitoring · high confidence
New devcontainer lifecycle scripts and utility functions
The development container environment now includes dedicated lifecycle scripts and helper utilities to streamline setup and customization. A new \postcreate.sh\ script handles initial repository cloning or fetching and allows for custom post-creation logic via \/devcontainer-customization/postcreate.sh\. A \poststart.sh\ script ensures generated source files are present by running \make prepare\ on container start, also supporting custom pre-start logic. Additionally, \utils.sh\ provides helper functions for configuring git user/remote settings and securely setting up SSH keys from a customization directory, enhancing the flexibility and ease of use for developers working in the devcontainer.
.devcontainer-scripts · high confidence
New distributed storage layer with S3 support and local caching
The storage service now supports a distributed mode where files are stored in an S3-compatible object store (AWS S3 or MinIO) with a local cache on each node, while single-node mode continues to use the local filesystem. This change introduces a unified FileManager that handles uploads with progress tracking, downloads with singleflight deduplication to prevent thundering-herd issues, and automatic cleanup of ephemeral keys to prevent storage leaks. The implementation includes a new ObjectStore interface, S3 and Filesystem implementations, and comprehensive tests for upload, download caching, and deletion behaviors.
core/services/storage · high confidence
New distributed worker commands for vLLM, MLX, and DS4
The CLI now exposes dedicated subcommands to launch distributed inference workers: \vllm\ for vLLM data-parallel follower processes, \mlx-distributed\ for MLX distributed workers, and \ds4-distributed\ for DS4 layer-split workers. These commands handle backend discovery and installation from the gallery, accept extra arguments (e.g., \--llama-cpp-args\), and support node registration with labels (such as \vllm-follower\) for visibility in the admin UI. The vLLM follower specifically self-registers as an agent-type node to appear in the UI while scoping regular model placement away from it.
core/cli/worker · high confidence
New insightface backend for face recognition and liveness detection
A new Python-based face recognition backend is available in the insightface module, offering 1:1 verification, 1:N identification, face detection, embedding generation, and age/gender analysis. It supports two interchangeable engines: the default insightface engine (using buffalo\_l/s/m/sc and antelopev2 model packs) and an Apache-2.0 licensed onnx\_direct engine (using OpenCV Zoo YuNet and SFace models). Additionally, antispoofing (liveness) detection is now supported via the Silent-Face MiniFASNet ensemble, configurable through specific ONNX model paths and thresholds.
backend/python/insightface · high confidence
New llama.cpp-based quantization backend for GGUF model conversion
A new gRPC-based backend has been added to handle model quantization using the llama.cpp toolchain. This component allows users to download HuggingFace models, convert them to the GGUF format, and apply quantization (such as q4\_k\_m) to reduce model size. The backend manages the installation of necessary dependencies, including the \convert\_hf\_to\_gguf.py\ script and the \llama-quantize\ binary, and exposes a \StartQuantization\ API for initiating jobs with progress tracking. A corresponding test suite verifies the end-to-end workflow of downloading, converting, and quantizing a small model.
backend/python/llama-cpp-quantization · high confidence
New model importers for ACE-Step, Bark, Chatterbox, Coqui, Depth Anything, Diffusers, and DeepSeek V4
The gallery importer system now includes dedicated handlers for several new model types, allowing automatic detection and configuration for ACE-Step music generation, Suno's Bark TTS, Resemble AI's Chatterbox, Coqui TTS, ByteDance's Depth Anything 3 (via GGUF), Hugging Face Diffusers pipelines, and DeepSeek V4 Flash (ds4). These importers recognize models via Hugging Face repository metadata, URI patterns, or explicit backend preferences, and generate the necessary configuration files to route them to the correct inference backends.
core/gallery/importers · high confidence
New moss-transcribe-cpp backend for offline transcription and diarization
This change introduces the \moss-transcribe-cpp\ backend, enabling LocalAI to perform offline audio transcription, speaker diarization, and timestamping using the \moss-transcribe.cpp\ engine. The backend loads a GGUF model via a shared C library (\libmoss-transcribe.so\) and exposes a gRPC service that accepts audio files, normalizes them to 16 kHz mono WAV, and returns structured transcript segments with speaker labels and nanosecond-precision timestamps. It includes build automation via a Makefile to clone and compile the upstream C++ source at a pinned commit, packaging scripts to bundle the binary and required system/GPU libraries, and a Go-based test suite covering ABI validation, load validation, and transcript parsing logic.
backend/go/moss-transcribe-cpp · high confidence
New routing subsystem: per-model concurrency control, usage billing, and KNN corpus management
This change introduces a new routing module in core/services/routing that adds three key capabilities. First, it implements per-model concurrency limiting (admission), which acquires a slot before request processing and returns a 503 with a Retry-After header when limits are reached, preventing request pile-up. Second, it provides a flexible usage billing system (billing) that tracks token counts and costs, supporting both persistent storage via GORM (when authentication is enabled) and an in-memory ring buffer fallback for single-user deployments, with Prometheus metrics for real-time monitoring. Third, it adds a KNN classifier corpus manager (corpus) that persists and indexes labelled exemplar texts for model routing, handling embedding caching, re-embedding on model changes, and duplicate detection. These components work together to provide better control over model usage, accurate billing/usage tracking, and intelligent request routing based on content similarity.
core/services/routing · high confidence
New schema definitions for Anthropic, 3D, and agent job capabilities
This change introduces a comprehensive set of new Go schema types in core/schema to support recently added product capabilities. It adds the Anthropic Messages API schema (anthropic.go), enabling support for tool use, extended thinking, and streaming events. It defines the data structures for 3D asset generation and animation (animation.go, localai.go), including inputs, parameters, and validation logic. It also introduces the schema for agent jobs (agent\_jobs.go), defining tasks, jobs, execution requests, and webhook configurations. Additionally, it includes schemas for audio transformations (audio\_transform.go), diarization (diarization.go), ElevenLabs sound generation (elevenlabs.go), fine-tuning jobs (finetune.go), Jina reranking (jina.go), and updates the core message schema (message.go) to handle reasoning content and tool calls. A JSON schema for gallery model specifications (gallery-model.schema.json) is also added to define the structure for model gallery entries.
core/schema · high confidence
New shared Python backend infrastructure for authentication, model identity, and process safety
This change introduces a new \backend/python/common\ library that standardizes core behaviors across all LocalAI Python backends. It adds gRPC bearer token authentication (controlled by \LOCALAI\_GRPC\_AUTH\_TOKEN\), model-identity enforcement to prevent serving the wrong model in distributed setups, and a parent-death watcher to automatically terminate orphaned backend processes. It also provides shared utilities for parsing model references, handling chat templates and tool calls, and a bash library (\libbackend.sh\) for portable Python installation and backend lifecycle management.
backend/python/common · high confidence
New speaker-recognition backend for voice verification and analysis
A new Python-based gRPC backend has been added to handle speaker recognition, providing 1:1 voice verification, embedding extraction, and demographic analysis (age, gender, emotion). It supports two engine backends: a default SpeechBrain engine using the ECAPA-TDNN model for high-accuracy verification, and an OnnxDirect engine for CPU-friendly inference with pre-exported ONNX models. The service exposes endpoints for verifying speaker identity, extracting speaker embeddings, and analyzing voice demographics, with optional age/gender and emotion models that can be configured via gallery entries or environment variables.
backend/python/speaker-recognition · high confidence
New unified Transformers backend with multi-hardware and modality support
The \backend/python/transformers\ directory now contains a standalone, unified backend implementation that consolidates previous separate backends (such as SentenceTransformers and MusicGen) into a single service. This backend supports a wide range of Hugging Face model types, including causal language models, feature extraction, audio generation (MusicGen), and text-to-speech. It provides hardware acceleration for CUDA, Apple MPS, Intel XPU, and OpenVINO, and includes specific optimizations like 4-bit/8-bit quantization via BitsAndBytes and Intel extensions. The entry also introduces the necessary build infrastructure (Makefile, install/run scripts) and a comprehensive test suite covering model loading, embedding generation, and audio/TTS capabilities.
backend/python/transformers · high confidence
New utility libraries for JSON handling, safe goroutines, and URL sanitization
This change introduces three new utility packages to the codebase. The \dbutil\ package provides \MarshalJSON\ and \UnmarshalJSON\ functions that safely handle JSON serialization for database storage, treating nil, empty, or null-equivalent values as empty strings. The \concurrency\ package adds a \SafeGo\ function that launches goroutines with automatic panic recovery, logging any panics with stack traces via the xlog library. The \sanitize\ package includes a \URL\ function that masks sensitive userinfo (username and password) in URL strings, returning '\\\*' if parsing fails.
core/services/dbutil, pkg/concurrency, pkg/sanitize · high confidence
New vibevoice-cpp backend with streaming TTS and GPU support
This location introduces the build system, Go wrapper, and tests for the new vibevoice-cpp backend. It adds a CMake build configuration and Makefile that fetch and compile the vibevoice.cpp library, supporting CPU variants (AVX, AVX2, AVX512) and GPU backends (CUDA, ROCm/HIP, Metal, Vulkan). The Go code provides a purego binding to the C API, enabling both batch TTS and true streaming TTS (TTSStream) that delivers audio incrementally, along with ASR capabilities. The package includes integration and unit tests to validate the streaming callback framing and backend semantics.
backend/go/vibevoice-cpp · high confidence
New voice-detect backend for speaker recognition and analysis
This change introduces a new C++-based voice-detect backend (replacing the previous Python insightface/speaker-recognition implementation) that provides speaker embedding, speaker verification, and demographic analysis (age, gender, emotion). The Go implementation in this directory bridges to the upstream voice-detect.cpp library via purego, exposing gRPC endpoints for VoiceEmbed, VoiceVerify, and VoiceAnalyze. It supports configurable verification thresholds (defaulting to 0.25 to match previous behavior) and allows users to specify thread budgets. The backend is packaged with its required shared libraries and GPU runtime dependencies (CUDA, ROCm, etc.) to ensure self-contained deployment.
backend/go/voice-detect · high confidence
OCI backend images are verified using keyless Cosign signatures
LocalAI now verifies the integrity and origin of downloaded backend images by checking for a Sigstore bundle attached as an OCI referrer. The new \pkg/oci/cosignverify\ package discovers these bundles via the OCI 1.1 referrers API (with fallback for registries that only support the legacy referrers-tag scheme) and validates them against a configurable policy, including checks for the signing identity (e.g., GitHub Actions OIDC), transparency log inclusion, and a \NotBefore\ cutoff time to invalidate older signatures. This ensures that only images signed by trusted CI pipelines are installed, protecting against registry compromise.
pkg/oci/cosignverify · high confidence
OCI image and artifact support with resilient downloads
The \pkg/oci\ package now provides core capabilities for pulling and extracting OCI images and artifacts, enabling features like gallery-based model distribution and Ollama model integration. To ensure reliability on unstable networks, layer downloads now automatically retry on transient errors and resume interrupted transfers using HTTP Range requests, which is critical for large images behind expiring pre-signed URLs. The artifact puller enforces strict security by validating artifact types, capping layer counts and total sizes, and preventing path traversal or symlink-based directory escapes. Additionally, extraction logic now gracefully handles filesystems that do not support symlinks by falling back to copying linked files as regular copies.
pkg/oci · high confidence
Ollama API compatibility layer for model inference and metadata
This change introduces a new Ollama-compatible HTTP API in the LocalAI server, enabling clients to interact with LocalAI using standard Ollama endpoints. The implementation includes handlers for chat (/api/chat), generation (/api/generate), and embeddings (/api/embed), which translate Ollama requests into the internal OpenAI-compatible format. It also adds model management endpoints (/api/tags, /api/show, /api/ps, /api/version, and root health check) that expose model metadata, including capabilities (e.g., embedding, completion, vision, tools, thinking), parameter size, and quantization level derived from model configurations. Additionally, the change includes logic to safely apply Ollama-specific options (like context size) with clamping to prevent integer overflow and denial-of-service risks, and supports the :latest tag for model lookups.
core/http/endpoints/ollama · high confidence
Per-node VRAM allocation budget and improved GPU memory detection
Users can now set a hard cap on VRAM usage for model allocation via the LOCALAI\_VRAM\_BUDGET configuration option (or the Settings page), accepting percentages (e.g., "80%") or absolute sizes (e.g., "12GB"). This prevents over-provisioning on multi-GPU hosts by ensuring models only fit within the specified budget. Additionally, GPU memory reporting is more accurate: AMD APUs now correctly include GTT in total VRAM, Intel iGPUs are filtered to avoid double-counting system RAM, and live usage is tracked via DRM fdinfo for better visibility.
pkg/xsysinfo · high confidence
Persistent state for distributed fine-tune, quantization, gallery, and skills operations
The distributed service now persists fine-tune jobs, quantization jobs, gallery operations, and skill metadata in PostgreSQL via new \FineTuneStore\, \QuantStore\, \GalleryStore\, and \SkillStore\ components. This ensures that in-flight operations (such as model installs and cancellations) and job states survive frontend replica restarts and are synchronized across the cluster, rather than being lost in memory.
core/services/distributed · high confidence
Placeholder directory for custom CA certificates
A new empty directory named custom-ca-certs has been added to the repository structure, containing only a .keep file to ensure the directory persists in version control. This establishes a location for future custom CA certificate configurations, though no actual certificate handling logic or user-facing functionality is implemented in this change.
custom-ca-certs · low confidence
Preload models from URIs at startup
Users can now specify model URIs (such as Hugging Face URLs or embedded YAML configurations) to be automatically downloaded and installed during application startup. The new \InstallModels\ function in the startup package handles resolving these URIs, downloading the necessary artifacts, and registering them with the model loader, ensuring models are available before the service becomes fully operational.
core/startup · high confidence
Quantization jobs now sync across distributed replicas
The quantization service now supports distributed mode, keeping job state consistent across multiple server instances via NATS and an optional PostgreSQL store. Users running multiple replicas will see quantization jobs created on any node reflected in real-time on all others, with job progress updates and status changes propagated automatically. In standalone mode, the service continues to persist jobs to disk as before, ensuring backward compatibility for single-instance deployments.
core/services/quantization · high confidence
React UI migration: new utility layer for API, artifacts, and configuration
The React UI has been migrated to a new codebase, introducing a dedicated \utils\ directory that centralizes core frontend logic. This includes a comprehensive API client (\api.js\) with typed methods for models, backends, chat, and MCP operations, alongside a centralized endpoint configuration (\config.js\). New capabilities include GLB animation parsing (\animationGlb.js\) for 3D content, code and media artifact extraction from chat messages (\artifacts.js\), and robust clipboard handling that works in non-secure contexts (\clipboard.js\). The migration also brings a new CodeMirror editor theme (\cmTheme.js\) and YAML autocomplete engine (\cmYamlComplete.js\), along with utility functions for entity grouping, FFT, and timestamp normalization.
core/http/react-ui/src/utils · high confidence
React UI project scaffolding with i18n, linting, and test tooling
The React UI location now includes the foundational configuration files for the frontend application. This adds a Vite build configuration that sets a relative base path to correctly handle reverse-proxy subpaths (like X-Forwarded-Prefix) and supports code-splitting. It also introduces ESLint rules for React hooks and refresh, an i18next parser configuration defining supported locales (en, it, es, de, zh-CN, id), and scripts for managing locale translations and enforcing an inline-style baseline. Additionally, a Playwright configuration is added for end-to-end testing, and an .nycrc.json file is created to manage code coverage reporting.
core/http/react-ui · high confidence
Redesigned Nodes page with fleet overview, capacity gauges, and local machine view
The Nodes page has been restructured to provide a comprehensive fleet overview and detailed node management. A new Cluster Overview displays fleet health and capacity gauges for VRAM, RAM, CPU, and disk, alongside an attention queue for nodes needing approval or showing unhealthy status. The view distinguishes between distributed clusters and single-node installations: distributed mode shows a Node Fleet Table with selection, grouping, and capacity details, while single-node mode presents a Local Machine View with a host-specific memory share bar and a Local Running Models table. New components include a Model Inspector for replica placement details, a Capacity Editor for per-model replica limits and VRAM budgets, and a KeyValueChips component for node labeling with suggestions.
core/http/react-ui/src/components/nodes · high confidence
Scheduling rules can now be keyed by model aliases
Operators can now create node scheduling rules using an alias name (e.g., 'production') that maps to a physical model (e.g., 'qwen3'). The system resolves the alias to determine the actual target model for placement, allowing rules to persist even if the alias is repointed to a different model. The scheduler enforces that only one rule governs a single physical model to prevent conflicts, and the node registry now exposes the resolved target model alongside the operator-defined rule name.
core/services/nodes · high confidence
Support for distributed skill management with PostgreSQL sync
Agents can now operate in a distributed mode where skill metadata is synchronized to a PostgreSQL database while the full skill content remains on the filesystem. The new \DistributedManager\ in \core/services/skills\ handles this by writing metadata to the database for cross-worker visibility (used for listing skills) while reading full content from the filesystem. A \FilesystemManager\ is retained for standalone, single-instance deployments. This change introduces the underlying service layer for distributed skill sharing, ensuring that skills created by one agent are discoverable by others via the database.
core/services/skills · high confidence
Usage page now breaks down traffic by source and API key with interactive charts
The Usage page now provides a detailed breakdown of token and request usage by source (Web UI, Legacy, and API keys). A new Sources tab displays a SourceMixRibbon showing the percentage share of each source class, a SourceTimeChart visualizing usage trends for the top seven sources over time, and a sortable, searchable SourcesTable listing individual API keys with their token counts, request counts, and last-used timestamps. In the admin view, the table also attributes Web UI and Legacy traffic to specific user accounts, allowing administrators to see which users are generating unkeyed traffic.
core/http/react-ui/src/pages/Usage · high confidence
WebUI redesign with new pages and Alpine.js frontend
The WebUI has been redesigned using Alpine.js (dropping HTMX) and features a new minimalistic theme. This update introduces several new pages: a dedicated 404 error page, a Backend Management gallery for discovering and installing backends, and a comprehensive Agent Jobs system for creating, scheduling, and monitoring tasks with execution traces. Existing pages like Chat, Index, and Models have been completely restyled and enhanced with new capabilities such as PDF upload support, model size estimation, and improved navigation.
core/http/views · high confidence
Worker registration client and NATS credential manager for distributed mode
Added a new worker registry package that provides a shared HTTP client for worker node registration, heartbeating, draining, and deregistration against the LocalAI frontend. This includes a NATS credential manager that handles acquiring credentials (waiting through admin approval if required) and automatically refreshing JWTs before they expire, ensuring long-running workers maintain valid authentication for NATS messaging.
core/cli/workerregistry · high confidence
Removals
Deprecation of legacy Alpine.js asset downloader
The tool previously used to download static assets for the legacy Alpine.js UI has been marked as deprecated. This component is no longer needed as the new React UI bundles all dependencies via npm, and the file will be removed once the legacy UI is fully retired.
_core/dependencies\manager · high confidence
Security
Centralized base64 handling and hardened URL validation
The \pkg/utils\ package now centralizes base64 data-URI and URL processing. \GetContentURIAsBase64\ handles both local data URIs (including those with extra MIME parameters like \data:audio/webm;codecs=opus;base64,...\) and remote HTTP(S) URLs, downloading and encoding remote content in memory. Remote URL fetching is protected by \ValidateExternalURL\, which blocks SSRF by rejecting loopback, private, link-local, and cloud metadata addresses, and the download client is configured to refuse redirects. Additionally, \pkg/utils/untar.go\ now validates archive member paths and rejects hardlinks that escape the extraction root, and \pkg/utils/ffmpeg.go\ introduces \AudioToWavPreservingShape\ to transcode audio while keeping the original sample rate and channel count, alongside a hardened \ffmpegCommand\ that clears the environment.
pkg/utils · high confidence
Secure file operations with race-safe removal and symlink protection
The \pkg/safefile\ package now provides platform-specific implementations for safe file reading and removal. On Unix systems, \ReadRegularAt\ uses \openat\ with \O\_NOFOLLOW\ to prevent symlink attacks and reject non-regular files (like FIFOs) without blocking. \RemoveExact\ performs race-safe removal of files and their sidecars using directory file descriptors, ensuring that symbolic links are not followed and that replacement targets are not accidentally deleted during concurrent operations. A portable fallback is provided for non-Unix platforms, which rejects symbolic links and non-regular files but fails closed for exact removal due to lack of component-relative no-follow operations. Comprehensive tests verify these security properties.
pkg/safefile · high confidence
gRPC backend security and reliability hardening
The gRPC transport layer now enforces strict model identity checks across all inference modalities (image, video, 3D, audio, etc.), rejecting requests that target a model different from the one currently loaded to prevent cross-model data leakage. It also introduces bearer token authentication for distributed backends, requiring a valid token and rejecting raw or empty Bearer values. Additionally, the client now uses a counter-based busy state to correctly track parallel in-flight requests, and supports rich tool-call deltas and usage metadata through the new AIModelRich interface.
pkg/grpc · high confidence
Architecture
Refactored backend build system with dedicated Dockerfiles and pre-built CI base images
The backend build infrastructure has been restructured to decouple individual C++ and Go backends from the monolithic build process. New dedicated Dockerfiles (e.g., Dockerfile.audio-cpp, Dockerfile.llama-cpp, Dockerfile.bonsai) now define isolated build environments for specific backends, supporting multiple hardware targets including CPU, CUDA, ROCm, and Vulkan. A new \Dockerfile.base-grpc-builder\ establishes a standardized pre-built CI base image (\quay.io/go-skynet/ci-cache:base-grpc-\*\) containing shared dependencies like gRPC and CMake, which backend-specific Dockerfiles can leverage to ensure bit-equivalent builds between local development and CI. This change improves build isolation, caching efficiency via ccache, and maintainability by allowing each backend to specify its own dependencies and build scripts.
backend · high confidence
Behavioural changes
Backend gallery integrity, distributed capability resolution, and variant selection
The gallery now enforces backend integrity by requiring SHA256 checksums or OCI signature verification (cosign) for installations, with strict mode blocking unverified backends. In distributed deployments, backend discovery and model sizing now correctly account for worker node capabilities and memory rather than being limited to the controller's hardware, ensuring GPU-only backends are visible and sized appropriately. Additionally, the gallery supports model variant collapsing to reduce listing clutter, provides detailed variant views (including memory footprint, quantization, and features) for auto-selection, and introduces backend versioning with metadata tracking for upgrades.
core/gallery · high confidence
Deterministic model auto-detection and per-replica backend logging
The model loader now uses a deterministic, type-filtered algorithm to auto-detect backends for GGUF models, ensuring that incompatible backends (like audio codecs) are excluded and llama-cpp is prioritized. Additionally, a new backend log store aggregates logs across distributed model replicas, allowing operators to view a unified log stream for a model regardless of how many instances are running.
pkg/model · high confidence
Extracted transport-agnostic replica selection policy
The replica picking logic has been extracted into a new, standalone package (pkg/clusterrouting) to be shared between the NATS distributed mode and the p2p federation server. This change introduces a unified selection policy that prioritizes replicas with the least in-flight requests, uses the oldest last-used timestamp as a tiebreaker to ensure round-robin distribution, and falls back to available VRAM for cold-start scenarios. The new package includes a test suite verifying this tiered precedence and stability.
pkg/clusterrouting · high confidence
HTTP server migration to Echo framework with enhanced security and performance
The HTTP server implementation has been migrated to the Echo framework, introducing several behavioral changes for users. Security is strengthened by defaulting to authentication for all HTTP routes, with a strict allowlist for public endpoints (health checks, login, static assets) and improved CSRF protection for multipart requests. Performance is improved through gzip response compression and immutable caching (one-year TTL) for content-hashed React UI assets. The server now handles model loading states more gracefully by returning 503 with Retry-After headers for both failed loads (cooldown) and in-progress loads, and respects the X-Forwarded-Prefix header for correct URL generation behind reverse proxies.
core/http · high confidence
Improved chat streaming reliability and spec compliance
The chat completion endpoint now correctly handles context overflow errors by returning a 400 Bad Request instead of a 500 error, allowing clients to adjust their context window. Streaming behavior has been fixed to comply with the OpenAI specification: usage data is no longer incorrectly included in intermediate chunks, and the final usage trailer is now properly emitted when requested. Additionally, reasoning content is now correctly separated from tool calls and text, preventing reasoning tags from leaking into the final response content.
core/http/endpoints/openai · high confidence
Introduce dynamic pipeline loader for automatic diffusers support
The diffusers backend now uses a dynamic loader that automatically discovers and loads any Hugging Face diffusers pipeline at runtime, eliminating the need for hardcoded per-pipeline conditionals. Users can now load new or custom pipelines (e.g., Flux, Wan, Sana) without code changes, resolving models by class name, task alias (like 'text-to-image'), or model ID. The backend also supports single-file model loading, optional prompt weighting via Compel and sd\_embed, and auto-detects CUDA availability instead of defaulting to CPU.
backend/python/diffusers · high confidence
Introduces React context providers for branding, operations, and system state
The React UI now uses dedicated context providers to manage global state, replacing previous ad-hoc patterns. BrandingContext fetches instance-specific branding (name, tagline, logos) from the backend and updates the document title and favicon, falling back to defaults if the API is unreachable. OperationsContext consolidates polling for active operations into a single shared poller to reduce API load, while also tracking cancelled operations and computing transfer rates/ETAs. OperateSummaryContext aggregates system health signals (upgrades, node status, resources, traces) with a slower polling interval to populate the Operate console's attention rail. FormContext provides read-only form data to nested field editors, and ThemeContext manages the dark/light theme preference persisted in localStorage.
core/http/react-ui/src/contexts · high confidence
Local AI CLI entry point restructured with enhanced logging and credential support
The local-ai command-line interface now initializes a new xlog-based logger that supports configurable log formats and optional log deduplication for terminal output. The startup process explicitly loads environment variables from multiple standard locations (including user-specific config paths) and introduces a credentials file mechanism to authenticate registries and downloads before running commands. Additionally, the CLI now passes its model definition to the completion command to enable dynamic shell completion script generation for bash, zsh, and fish.
cmd/local-ai · high confidence
MCP connection timeouts and distributed tool execution
MCP server connections now respect a bounded timeout, preventing the UI from hanging indefinitely when a server is unreachable or misbehaving. The system also introduces a ToolExecutor abstraction that supports both local in-process sessions and distributed execution via NATS, and adds an in-memory LocalAI Assistant server that exposes admin tools directly within the chat session.
core/http/endpoints/mcp · high confidence
More accurate model size and VRAM estimates with persistent caching
The VRAM estimation logic in the gallery now reports the size of the largest GGUF quantization variant for a model instead of summing all files in the repository, preventing inflated size displays. Estimates are now computed using GGUF architecture metadata (such as layer count and embedding dimensions) for greater precision, and the system persists remote probe results to disk so that size and metadata lookups are reused across server restarts, speeding up startup and reducing redundant network requests.
pkg/vram · high confidence
New CLI context configuration for logging and credentials
The CLI now exposes a dedicated context struct that allows users to configure logging behavior and authentication sources. Users can set the log level (error, warn, info, debug, trace) and format (default, text, json) via environment variables or flags, and enable deduplication of consecutive identical log lines. Additionally, a credentials file path can be specified to authenticate against private registries, galleries, and download hosts.
core/cli/context · high confidence
New Hugging Face API client with hardened HTTP behavior
A new Hugging Face API client has been introduced in \pkg/huggingface-api\ to handle model searches and snapshot resolutions. The client is configured to refuse HTTP redirects on outbound requests, enhancing security by preventing potential redirect-based attacks. It includes robust retry logic for transient errors and rate limits, and supports resolving model snapshots with immutable revision pinning, file filtering, and pagination handling.
pkg/huggingface-api · high confidence
New PEG-based tool call parser with auto-detection and backend integration
The function call parsing subsystem has been significantly refactored to support a new PEG-based parser that handles various XML-style tool call formats (such as Qwen3-Coder, Functionary, and GLM-4.5) and integrates with the C++ backend's auto-detected format markers. This change introduces a new \ToolCallsFromChatDeltas\ function to extract tool calls directly from C++ autoparser chat deltas, bypassing Go-side parsing when backend markers are available. It also adds robust validation for auto-detected XML tool call names to prevent false positives (e.g., from Hermes-style JSON blocks) and supports mixed grammars, parallel calls, and custom XML formats via the \FunctionsConfig\ structure.
pkg/functions · high confidence
New backend execution layer with distributed tracing and context propagation
The core/backend package has been refactored to introduce a new execution layer that standardizes model loading, slot acquisition, and distributed tracing across all inference types (LLM, image, audio, video, etc.). This change ensures that per-request context (including the X-LocalAI-Node header for distributed routing) is correctly propagated from the HTTP handler to the backend router, fixing previous issues where context was dropped. It also adds comprehensive backend tracing for observability, allowing users to track request durations, errors, and metadata for all backend operations.
core/backend · high confidence
New config package with runtime settings and backend capabilities
The core/config package has been restructured to introduce a single source of truth for default values and a new runtime settings system. This change adds a comprehensive backend capabilities registry that maps usecases (such as chat, vision, TTS, and video) to specific gRPC methods, ensuring accurate backend detection and routing. It also introduces VRAM-aware context sizing that automatically caps the default context window based on per-device memory to prevent load failures, and adds a configurable disk-headroom check for distributed mode. Additionally, the application configuration now supports runtime updates via a new ApplyRuntimeSettings mechanism, allowing operators to modify settings like watchdog behavior and concurrency without restarting the service.
core/config · high confidence
New llama.cpp gRPC backend with unified build and compatibility fixes
The llama.cpp backend is now built as a standalone gRPC server (grpc-server) with a dedicated CMakeLists.txt and Makefile, replacing the previous embedded build approach. This new build system introduces a single-build CPU variant (CPU\_ALL\_VARIANTS) that produces one binary with runtime-dynamic CPU backends, replacing the previous multi-binary AVX/AVX2/AVX512/fallback builds. The backend now includes robust compatibility handling for upstream llama.cpp refactors, automatically detecting the renamed llama-common library and conditionally including split source files (server-chat.cpp, server-schema.cpp, server-stream.cpp) to support both current and older fork versions. Additionally, it introduces message content normalization to prevent JSON-sniffing errors with plain-text prompts, adds gRPC bearer token authentication for distributed mode, and includes unit tests for message handling, stream peers, and model load errors.
backend/cpp/llama-cpp · high confidence
New template engine with caching, multimodal support, and structured data injection
The template system in core/templates has been rewritten to include a caching layer (cache.go) that loads and compiles templates once for reuse, improving performance. The new evaluator (evaluator.go) introduces structured data injection for prompts, allowing templates to access specific fields like SystemPrompt, ReasoningEffort, and Metadata, as well as detailed ChatMessage data including Role, FunctionCall, and ToolCalls. This enables more precise control over how function calls and tool responses are formatted in chat histories. Additionally, a new multimodal templating component (multimodal.go) handles the insertion of media markers for images, audio, and video, ensuring that multimodal content is correctly represented in the final prompt sent to the model. Tests confirm the new behavior for both standard chat and embedding scenarios.
core/templates · high confidence
New web UI design system and navigation structure
The web interface now uses a new design system built on Tailwind CSS and Alpine.js, replacing the previous styling and scripting approach. This change introduces a new sidebar navigation layout with sections for Home, Install Models, Chat, Images, Video, TTS, Sound, Talk, Agent Jobs, Traces, and System tools, along with a dedicated explorer navbar. Users benefit from a persistent dark/light theme toggle that respects system preferences and persists via local storage, as well as a global operations status bar that provides real-time progress updates and cancellation capabilities for background tasks like model imports or deletions.
core/http/views/partials · high confidence
Refactored Docker build system with shared scripts and configurable apt mirrors
The Docker build infrastructure has been restructured to use shared shell scripts (install-base-deps, apt-mirror, and specific compile/build-target scripts for llama-cpp, turboquant, bonsai, and ik-llama-cpp) instead of inline RUN blocks. This ensures bit-equivalence between local from-source builds and CI prebuilt images. A new apt-mirror script allows routing package downloads through alternate Ubuntu mirrors via environment variables to mitigate archive outages. Build logic now centralizes decisions on CPU variant compilation (e.g., using a single CPU\_ALL\_VARIANTS build or a fallback) based on architecture and backend type, fixing issues where GPU builds unnecessarily compiled CPU variants or where SYCL/ROCm builds exceeded time limits.
.docker · high confidence
TRL fine-tuning backend introduces GRPO inline reward security fix
The new TRL fine-tuning backend in backend/python/trl adds support for SFT, DPO, GRPO, and other training methods via a gRPC interface, including built-in reward functions and a Makefile/installation script. Crucially, inline reward code execution (GRPO) is disabled by default to prevent remote code execution; it now requires the LOCALAI\_TRL\_ALLOW\_INLINE\_REWARD environment variable to be explicitly set to a truthy value, with tests verifying this opt-in gate.
backend/python/trl · high confidence
Updated API documentation with 3D generation and agent job endpoints
The Swagger API documentation has been regenerated to reflect recent backend changes. The documentation now includes new endpoints for 3D asset management, specifically POST /3d/animate for creating animations, POST /3d/generations for creating assets from images, and POST /3d/remesh for remeshing existing models. Additionally, the API reference now documents the agent job management endpoints, including GET /api/agent/jobs for listing jobs and POST /api/agent/jobs/execute for triggering job execution. The generated files (docs.go, swagger.json, swagger.yaml, and embed.go) have been updated to ensure the interactive documentation matches the current server implementation.
swagger · high confidence
Website content and structure updated for 4.8 release
The website now includes the 4.8 release post, updated benchmark data for DeepSeek and Laguna, and new demo clips. The ecosystem band has been rewritten to remove incorrect coverage links, and the hero and gallery video clips have been re-recorded to reflect the 4.8 UI.
website · high confidence
macOS app notarization and offline verification support
The macOS distribution process now includes code signing and notarization for both the application bundle and the disk image. A new script (contrib/macos/sign-and-notarize.sh) and entitlements file handle the signing process, while the notarization step specifically staples the Apple notarization ticket to the .app bundle itself, not just the DMG. This ensures the application can verify its notarization status offline, preventing launch failures on machines without internet access or behind firewalls.
contrib · high confidence
Fixes
8 commits (5 fixes) fixing pkg/reasoning
A fix in pkg/reasoning — 8 commits (5 fixs), 6 files.
pkg/reasoning · medium confidence · unverified
Fix HIP compilation and remove D512 turbo quantization support
This update fixes compilation failures on HIP-enabled builds by patching the CUDA event creation to use \cudaEventCreateWithFlags\ instead of the unsupported \cudaEventCreate\. It also removes the D512 flash attention implementations for the TURBO2\_0, TURBO3\_0, and TURBO4\_0 quantization types, effectively disabling these specific quantization configurations in the CUDA backend.
backend/cpp/turboquant/patches · high confidence
Gallery operations now persist cancellable state and survive restarts
In-flight model and backend installations are now correctly marked as cancellable in the persistent gallery store, ensuring that the cancel button remains available in the UI after a service restart. Previously, the cancellable flag was lost during persistence, causing in-flight operations to become orphaned and uncancelable until a stale-operation reaper timed them out. This fix also ensures that cancelled operations are persisted as terminal, preventing them from being re-hydrated as active on restart, and adds tests to verify this behavior across restarts.
core/services/galleryop · high confidence
Test coverage
Added e2e test coverage for the React UI; Added standalone C++ unit test runner script; New build scripts and regression tests for macOS backends, health checks, and GPU library packaging; New end-to-end test suites for AIO, backends, and UI; Worker test suite for address resolution, auth logic, and backend lifecycle.
Dependencies
Add dependency requirements for the Ace-Step music generation backend
Added Python dependency manifests (requirements files) for the new Ace-Step backend, covering CPU, CUDA 12/13, ROCm, Intel XPU, L4T13, and MPS hardware targets. These files define the necessary packages for the backend to function, including PyTorch variants, transformers, diffusers, and gradio.
(dependencies) · high confidence
Update ggerganov/llama.cpp
The ggerganov/llama.cpp submodule has been updated to the latest commit, bringing in upstream improvements and bug fixes.
(repo-wide) · high confidence
Vendored frontend libraries added to static assets
The static assets directory now includes vendored copies of Alpine.js (v3.13.10), CodeMirror (including its CSS and an auto-refresh extension), enabling the web UI to function without external CDN dependencies. These files are newly added to the repository to support the static embedding of frontend resources.
core/http/static/assets · high confidence
Housekeeping
Placeholder file added to configuration directory
A .keep file has been added to the configuration directory to ensure the directory is tracked by version control.
configuration · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Score
- CAI 44 → 45 (+0.6)
- Rubric changed (rubric-2026.08.19 → rubric-2026.09.15) — scores are not directly comparable.
Lenses
- Code Health 36 → 39 (+3.4)
- Architecture 92 → 57 (-34.5)
- Maturity 78 → 78 (+0.0)
- Readiness 46 → 43 (-2.5)
- Security 56 → 67 (+10.2)
- Domain Modelling 61 → 65 (+3.8)
- Event-Driven 88 → 88 (+0.1)
- Accessibility 48 → 47 (-0.6)
Resolved (502)
- (anonymous) (cognitive 155) (core/http/static/chat.js)
- (anonymous) (cognitive 162) (core/http/static/chat.js)
- (anonymous) (cognitive 20) (core/http/react-ui/src/hooks/useGalleryEnrichment.js)
- (anonymous) (cognitive 313) (core/http/react-ui/src/hooks/useChat.js)
- (anonymous) (cognitive 47) (core/http/react-ui/src/utils/cmYamlComplete.js)
- (anonymous) (cognitive 68) (core/http/static/chat.js)
- (anonymous) (cyclomatic 134) (core/http/react-ui/src/hooks/useChat.js)
- (anonymous) (cyclomatic 27) (core/http/react-ui/src/utils/cmYamlComplete.js)
- (anonymous) (cyclomatic 47) (core/http/static/chat.js)
- (anonymous) (cyclomatic 50) (core/http/static/chat.js)
- (anonymous) (cyclomatic 53) (core/http/static/chat.js)
- ConfigService.EditYAML (cognitive 31) (core/services/modeladmin/config.go)
- ConfigService.EditYAML (cyclomatic 24) (core/services/modeladmin/config.go)
- Coverage not included — suite not readable by the collector
- Critical CVE: [GHSA redacted] (backend/python/sglang/requirements-cublas12-after.txt)
- Critical CVE: [GHSA redacted] (backend/python/coqui/requirements-cpu.txt)
- Critical CVE: [GHSA redacted] (go.mod)
- Dependency hygiene not measured — dependency manifest found but not parsed for hygiene
- Duplicated block (10 lines × 2) (backend/go/moss-tts-cpp/gomossttscpp.go)
- Duplicated block (10 lines × 2) (backend/go/omnivoice-cpp/audio.go)
- …and 482 more
New (2359)
- (anonymous) (cognitive 90) (core/http/static/chat.js)
- (anonymous) (cyclomatic 51) (core/http/static/chat.js)
- (anonymous)::add (cognitive 27) (core/http/static/chat.js)
- (anonymous)::add (cyclomatic 23) (core/http/static/chat.js)
- (anonymous)::updateTokenUsage (cognitive 17) (core/http/static/chat.js)
- (anonymous)::updateTokenUsage (cyclomatic 17) (core/http/static/chat.js)
- Account.ApiKeysTab (cognitive 18) (core/http/react-ui/src/pages/Account.jsx)
- Account.ProfileTab (cognitive 16) (core/http/react-ui/src/pages/Account.jsx)
- Account.ProfileTab (cyclomatic 18) (core/http/react-ui/src/pages/Account.jsx)
- ActionMenu.ActionMenu (cognitive 35) (core/http/react-ui/src/components/ActionMenu.jsx)
- ActionMenu.ActionMenu (cyclomatic 41) (core/http/react-ui/src/components/ActionMenu.jsx)
- Activity.Activity (cognitive 20) (core/http/react-ui/src/pages/Activity.jsx)
- Activity.Activity (cyclomatic 25) (core/http/react-ui/src/pages/Activity.jsx)
- AgentChat.AgentChat (cognitive 113) (core/http/react-ui/src/pages/AgentChat.jsx)
- AgentChat.AgentChat (cyclomatic 131) (core/http/react-ui/src/pages/AgentChat.jsx)
- AgentCreate.AgentCreate (cognitive 172) (core/http/react-ui/src/pages/AgentCreate.jsx)
- AgentCreate.AgentCreate (cyclomatic 140) (core/http/react-ui/src/pages/AgentCreate.jsx)
- AgentCreate.ConfigForm (cognitive 19) (core/http/react-ui/src/pages/AgentCreate.jsx)
- AgentCreate.ConfigForm (cyclomatic 21) (core/http/react-ui/src/pages/AgentCreate.jsx)
- AgentCreate.FormField (cognitive 20) (core/http/react-ui/src/pages/AgentCreate.jsx)
- …and 2339 more
Changes since last survey
- 300 commits — 218 feature/other, 82 fixes
By area
- backend/cpp — 64 commits
- backend/go — 56 commits
- backend/python — 37 commits
- gallery/index.yaml — 35 commits
- core/http — 31 commits
- docs/content — 29 commits
- core/services — 15 commits
- (root) — 5 commits
- core/gallery — 4 commits
- pkg/model — 3 commits
- swagger/docs.go — 3 commits
- .github/workflows — 2 commits
- backend/Dockerfile.audio-cpp — 2 commits
- core/config — 2 commits
- pkg/xsysinfo — 2 commits
- .agents/backend-signing.md — 1 commit
- .github/bump_ctranslate2_rocm_wheel.sh — 1 commit
- core/cli — 1 commit
- docs/data — 1 commit
- docs/themes — 1 commit
Notable commits
- fix: Update containers.md to fix podman image qualification (#11749)
- fix: fix(agent-ui): keep chat open for status
- fix: fix(agents): pass pool limits to the LocalAGI agent pool
- fix: fix(api): report an alias's target in /v1/models/capabilities (#12183)
- fix: fix(audio-cpp): bundle rocRoller for ROCm
- fix: fix(audio-cpp): install hipBLAS headers
- fix: fix(audio-cpp): install rocBLAS headers
- fix: fix(audio-cpp): normalize HIP target list
- fix: fix(auth): bypass API-key auth for CORS preflight (OPTIONS) requests (#11113)
- fix: fix(auth): require validated header credentials for CSRF exemption (#12185)
- fix: fix(backends): bound temporary scratch files (#11941)
- fix: fix(backends): honor enable_thinking=false in mlx and vllm-omni (#11962)
- fix: fix(backends): preserve an explicit seed of 0 in sglang and vllm
- fix: fix(ci): sign backends in the format we verify (#12166)
- fix: fix(config): register original diffusers config field
- fix: fix(cosignverify): find a bundle the index entry describes badly (#12165)
- fix: fix(deps): upgrade path-to-regexp to 8.4.0 ([CVE redacted])
- fix: fix(detection): avoid temporary image files (#11938)
- fix: fix(diffusers): auto-detect CUDA instead of defaulting to CPU (#11891)
- fix: fix(diffusers): forward original config for single files
- …and 280 more
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
mudler/LocalAI was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 24 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 81eaca8768287ce380ea548bcef0b3bf7eab24d1 — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-5f8d0eb43fd7.