exo-explore/exo
38.8
Weak · 26 September 2026
43.3k
lines of production code
Python
with TypeScript, Swift
5
measurements over time
What this system is
EXO is a distributed AI inference cluster that enables multiple devices to collaboratively run large language and image generation models. It manages the lifecycle of model instances across nodes, handling tasks like tensor and pipeline parallelism, prefill/decode disaggregation, and resource-aware placement. The system provides a macOS-native dashboard and API server for monitoring cluster topology, managing downloads, and interacting with models via standard OpenAI, Claude, and Ollama interfaces.
How it got here
2024–2025 — Distributed inference architecture and macOS client
41 changes.
The project established a new distributed inference engine by replacing legacy networking with Zenoh and implementing a master-worker architecture with intelligent placement and pipeline parallelism. This backend overhaul was accompanied by the development of a comprehensive macOS application and web dashboard for cluster management, monitoring, and user interaction.
2026 — API standardization and image generation support
27 changes.
This period focused on establishing a standardized API layer with adapters for OpenAI, Claude, and Ollama formats, alongside introducing initial image generation capabilities via the M-Flux engine. Significant infrastructure work included implementing prefill/decode disaggregation for MLX models, enhancing the download subsystem with offline mode support, and expanding the dashboard with new traces and integrations pages.
Features
Add example configurations for local exo cluster integration
Added example configuration files for integrating with a local exo cluster. The new \claude\_code.sh\ script provides a shell environment to run Claude Code against a local endpoint using the GPT OSS 120B model, while \opencode.json\ supplies the corresponding JSON configuration for the OpenCode editor, defining the local API endpoint, model details, and context/output limits.
_tmp/config\examples · high confidence
Added PyInstaller build specification for Exo
The repository now includes a PyInstaller specification file (exo.spec) that defines how to package the Exo application into a standalone executable. This configuration ensures the inclusion of the dashboard build assets, resource files, shared model directories, and necessary MLX libraries, while also handling platform-specific requirements such as the macmon binary on macOS.
packaging/pyinstaller · high confidence
Added type stubs for mflux model backends and core infrastructure
This update adds comprehensive Python type stubs (.pyi) for the mflux library, covering the core model infrastructure and several new model backends. The stubs define the interfaces for model configuration, weight loading and mapping, LoRA support, tokenizers, VAE tiling, and schedulers. Additionally, they introduce type definitions for the new DepthPro depth-estimation model, the FIBO model, and the FIBO-VLM vision-language model, ensuring static type checking and IDE support for these components.
.typings · high confidence
Custom macOS DMG installer with guided drag-to-Applications layout
The macOS packaging process now generates a polished DMG installer for EXO that features a custom background image and a specific window layout. The installer window displays the application icon on the left and an alias to the Applications folder on the right, connected by a curved arrow graphic to guide users in dragging the app to install. This change introduces the \create-dmg.sh\ build script and a Python helper (\generate-background.py\) to render the visual assets, replacing any previous default or unstyled DMG generation.
packaging/dmg · high confidence
EXO uninstaller script added with option to preserve models
A new uninstall script for EXO has been added, allowing users to cleanly remove system components such as the LaunchDaemon, network setup scripts, logs, and the EXO data directory. The script now supports a --keep-models flag, which preserves the \~/.exo/models directory during uninstallation, giving users control over whether their model data is retained or removed.
app/EXO · high confidence
Initial release of the EXO dashboard with chat, topology, and model management
This change introduces the core dashboard application for the EXO distributed AI cluster. It establishes the main layout with connection status banners and toast notifications, and implements the primary home page featuring a real-time cluster topology graph, a chat interface with model selection and message history, and a model picker modal. The UI includes support for mobile sidebars, download progress tracking, and specific hardware warnings for RDMA and macOS version mismatches.
dashboard/src/routes · high confidence
Initial repository scaffolding and developer guidelines
The repository has been initialized with foundational configuration and documentation files, including a Nix flake for the development environment, a justfile for build and lint commands, and a .python-version file specifying Python 3.13. To guide contributors and AI coding agents, the project includes AGENTS.md, CLAUDE.md, .cursorrules, and .clauderules, which enforce strict typing, pure functions, and Pydantic usage. Additionally, a README.md, PLATFORMS.md, and CONTRIBUTING.md have been added to document features, supported hardware (Apple Silicon, Linux), and contribution workflows.
(repo-wide) · high confidence
Initial support for FLUX and Qwen image generation models
The image generation engine now supports FLUX (including Schnell, Dev, and Kontext variants) and Qwen-Image models. This change introduces a new model adapter registry in the \src/exo/worker/engines/image/models\ package, allowing users to generate images using these specific model families. The implementation includes dedicated adapters for handling model-specific requirements, such as FLUX's dual text encoders and Qwen's attention masks, along with configuration for block types and synchronization steps.
src/exo/worker/engines/image/models · high confidence
Initial support for image generation via the M-Flux engine
The worker now includes a new image generation engine located in \src/exo/worker/engines/image\. This addition introduces the \DistributedImageModel\ class, which handles model loading, sharding, and diffusion execution, along with a \generate\_image\ function that processes generation tasks. Users can now generate images using the M-Flux backend, with support for configurable quality levels (low, medium, high), custom dimensions, seed control, and optional streaming of partial results.
src/exo/worker/engines/image · high confidence
Introduce Sparkle framework and commit Xcode project to version control
The EXO Xcode project is now tracked in the repository, enabling developers to build the app directly from source. This change adds the Sparkle framework as a dependency, which provides automatic update capabilities for the Mac app. Additionally, the commit scheme includes environment variables for AWS credentials, indicating integration with a bug-reporting or analytics service.
app/EXO/EXO.xcodeproj · high confidence
Introduce exo-bench for standardized inference throughput and energy measurement
The \bench\ area now includes a comprehensive benchmarking suite (\exo\_bench.py\) that measures inference throughput (prefill and generation tokens per second) and system-level resource consumption (GPU utilization, temperature, and power draw) for an exo cluster. The tool sends prompts to a dedicated \/bench/chat/completions\ endpoint, which disables the KV prefix cache by default and bans EOS tokens to ensure consistent, fair timing measurements. It supports concurrent request batching, configurable warmup runs, and prefix cache mode for specific optimization tracking. Results are output as JSON containing per-request statistics, cluster state, and time-series system metrics, with energy computed via trapezoidal integration of power samples. The suite is configured via TOML files (e.g., \bench.toml\, \single-m3-ultra.toml\) and includes specific benchmarks for disaggregated prefill-decode workloads (\prefill\_decode\_bench.py\) and tool-calling evaluation scenarios (\scenarios.toml\).
bench · high confidence
Introduce macOS menu bar application with cluster management and network diagnostics
The EXO macOS app is now available as a menu bar interface that manages the EXO process lifecycle, displays cluster topology and resource usage, and provides settings for custom namespaces, Hugging Face tokens, and environment variables. It includes a local network access checker that warns users when device discovery is blocked by macOS permissions and offers a direct link to System Settings, alongside a Thunderbolt Bridge service to prevent network loops. The app also supports automatic updates via Sparkle, launch-at-login registration, and a dedicated bug report window that uploads logs via presigned URLs.
app/EXO/EXO · high confidence
Introduces MLX-based image generation engine with pipeline support
The worker now includes a new engine architecture for MLX-based image generation, featuring a base engine interface and a builder pattern to manage connection and loading phases. This update adds a dedicated image pipeline with a DiffusionRunner, block wrappers for joint transformer blocks, and a specialized ImagePatchKVCache to optimize key-value caching for image patches during generation. The worker's main loop and planning logic have been extended to handle ImageEdits and ImageGeneration tasks, including buffering input image chunks and managing an image cache for efficient processing.
src/exo/worker · high confidence
Introduces structured type definitions for the distributed inference system
This change establishes the core type system in src/exo/shared/types, defining the data structures that govern how the system communicates and manages state. It introduces a Backend enum to distinguish between MlxMetal, MlxCpu, MlxCuda, and Vllm execution environments, and defines granular chunk types (TokenChunk, ImageChunk, ToolCallChunk, ErrorChunk, PrefillProgressChunk) to support streaming responses, image generation, tool usage, and progress tracking. The diff also adds command definitions for text and image generation, instance lifecycle management, and downloads, alongside a comprehensive State model that tracks node identities, memory/disk usage, network interfaces, and custom model cards. Additionally, it provides utility types for memory management, system profiling, and unique identifiers (NodeId, ModelId, CommandId) to ensure consistent data handling across the distributed cluster.
src/exo/shared/types · high confidence
New API adapters for Claude, Ollama, and OpenAI Responses formats
The API layer now includes dedicated adapters for Claude Messages, Ollama Chat/Generate, and OpenAI Responses protocols, allowing users to interact with the model using these specific interface standards. These new modules handle the conversion of incoming requests from each format into the internal text generation parameters and translate internal model chunks back into the respective response formats, including support for multimodal inputs (images) and tool calls where applicable.
src/exo/api/adapters · high confidence
New EXO macOS services for network configuration, state polling, and bug reporting
The app now includes a suite of new background services to improve stability and user experience. NetworkSetupHelper installs a LaunchDaemon to dynamically manage network locations and disable Thunderbolt Bridge to prevent packet storms, while ThunderboltBridgeDetector and ThunderboltBridgeService actively monitor for and prompt users to break detected bridge loops. ClusterStateService polls the local API for cluster state using an ephemeral, non-caching URLSession to prevent excessive disk I/O and memory usage. LocalNetworkChecker verifies actual local network connectivity via mDNS rather than relying solely on system settings. NetworkStatusService aggregates RDMA and interface status for the dashboard. Finally, BugReportService collects logs and state data and uploads them using presigned URLs.
app/EXO/EXO/Services · high confidence
New Integrations dashboard page with configuration helpers
A new Integrations page has been added to the dashboard, providing a centralized interface for configuring and connecting various AI models and tools. The page dynamically detects running models and offers pre-configured setup snippets for Anthropic Claude (including shell commands and JSON settings), OpenCode, and OpenAI Codex. It also includes state variables and default model assignments for Pi and OpenClaw integrations, streamlining the process of setting up these external services within the application.
dashboard/src/routes/integrations · high confidence
New LLM inference output parsing and chunking module
A new \llm\_inference\ package has been introduced to handle the parsing of model outputs and conversion into standardized response chunks. This module implements specific parsers for DeepSeek V3.2, DeepSeek V4, and GPT-OSS models, including logic to correctly count reasoning tokens for models that use thinking tags. It also handles tool call parsing and ensures that error finish reasons are properly propagated as error chunks to the user.
_src/exo/worker/runner/llm\inference · high confidence
New MLX generator engine with remote prefill and pipeline parallel support
This change introduces a new MLX-based text generation engine located in src/exo/worker/engines/mlx/generator. The engine implements core generation logic in generate.py, including support for pipeline parallel prefilling (overlapping stages across ranks) and remote prefilling via remote\_prefill.py (fetching KV cache chunks from a disaggregated server). It integrates with MLX's stream\_generate and logits processors, handles vision model embedding injection, and manages KV cache snapshots. This is a new component replacing or supplementing previous generation mechanisms.
src/exo/worker/engines/mlx/generator · high confidence
New Rust-based Python bindings for networking and system utilities
The \exo\_rs\ module now provides Python bindings for core networking and system functionality, replacing the previous \exo\_pyo3\_bindings\ crate. Users can now access \NetworkingHandle\ for managing GossipSub subscriptions and publishing messages, as well as \Pidfile\ for creating and locking process ID files. This change introduces a new Rust implementation for these features, including async support via PyO3 and Zenoh for the networking layer, and adds corresponding Python tests to validate the integration.
rust · high confidence
New Traces dashboard with deletion and Perfetto integration
The dashboard now includes a dedicated Traces page that lists all available traces, allowing users to select and delete multiple traces at once. Individual trace detail pages provide performance statistics broken down by phase and rank, along with options to download the raw trace data or open it directly in the Perfetto UI for advanced visualization.
dashboard/src/routes/traces · high confidence
New benchmarking, evaluation, and API type definitions
This change introduces new tooling and type definitions for the exo API. It adds benchmarking scripts (\bench/eval\_tool\_calls.py\, \bench/exo\_eval.py\) to evaluate model performance on tool calling and standard benchmarks (GPQA, MMLU, AIME, HumanEval, LiveCodeBench). It also defines new Pydantic types for the OpenAI Responses API (\src/exo/api/types/openai\_responses.py\), Claude Messages API (\src/exo/api/types/claude\_api.py\), and Ollama API (\src/exo/api/types/ollama\_api.py\), expanding the supported API interfaces. Additionally, it includes a new \DownloadCoordinator\ for managing model downloads and an \ImageStore\ for handling image data in the master node.
exo · high confidence
New centralized state management and UI stores for the dashboard
The dashboard now uses a new set of Svelte stores to manage application state and UI interactions. The central app store handles topology data, chat state, and download progress, while dedicated stores manage user preferences like favorite and recently launched models (persisted to localStorage) and a global toast notification system for user feedback.
dashboard/src/lib/stores · high confidence
New dashboard view models for instances, nodes, and topology
The app now exposes structured view models for the dashboard UI: InstanceViewModel aggregates download progress, node names, and chat tasks for each model instance, while NodeViewModel exposes per-node hardware metrics (RAM, GPU/CPU usage, temperature, power) and device icons. Additionally, TopologyViewModel presents the cluster network graph, ensuring the local node is always positioned at the top of the list for easier identification.
app/EXO/EXO/ViewModels · high confidence
New integration test infrastructure and CLI tooling
This change introduces a new set of integration test scripts and a Python-based CLI tool (\exo\_tools\) to support testing and interacting with the exo cluster. The \tmp/old\_tests\ directory now contains shell scripts (\auto\_bench.sh\, \eval\_tool\_calls.sh\) and a Python helper (\get\_all\_models\_on\_cluster.py\) for automating benchmark runs, evaluating tool-call capabilities, and querying cluster state via Tailscale. Additionally, the \tools/src/exo\_tools\ package provides an HTTP client for the exo API (including streaming chat completions and state retrieval) and a harness for managing instance lifecycles, waiting for readiness, and handling model resolution.
_tmp/old\tests, tools · high confidence
New macOS native UI components for settings, bug reporting, and onboarding
The app introduces several new SwiftUI views and window controllers for macOS: a dedicated Bug Report window for submitting diagnostics, a First Launch Popout that guides new users to the dashboard, a native Settings window with tabs for General, Model, Advanced, Environment, and About sections, and new list/detail views for instances, nodes, and cluster topology.
app/EXO/EXO/Views · high confidence
New network reachability and macOS system info gathering utilities
The info\_gatherer module now includes dedicated utilities for checking node reachability via HTTP probes against the configured API port and for collecting detailed macOS system information. The new net\_profile.py component asynchronously probes network interfaces to verify node identity and reachability, yielding results as they complete. The system\_info.py component provides functions to retrieve the OS version, macOS build number, friendly computer name, and detailed network interface types (distinguishing Wi-Fi, Ethernet, and Thunderbolt) by parsing macOS-specific system commands. These changes support more robust cluster discovery and node identification on macOS.
_src/exo/utils/info\gatherer · high confidence
New scripts for model card maintenance and cluster-wide model distribution
Added two new utility scripts to the scripts directory. The fetch\_kv\_heads.py script automates the process of fetching the num\_key\_value\_heads parameter from HuggingFace model configurations and updating the corresponding local TOML model card files, supporting both incremental updates for missing values and full overwrites. The download\_model\_to\_cluster.py script provides a way to bypass the standard placement logic to download a specified model to every node in an exo cluster simultaneously, handling model registration, topology discovery, and progress polling until all nodes report completion.
scripts · high confidence
New temporary utility scripts for model management, deployment, and security testing
Added a collection of temporary scripts in the tmp directory to support model operations and infrastructure. gen\_card.py automates the generation of inference model cards from Hugging Face. quantize\_and\_upload.py handles downloading, quantizing, and uploading models (including Qwen, FLUX, and FIBO variants) to Hugging Face. run\_exo\_on.sh facilitates deploying the EXO binary to remote hosts via SSH, using Zenoh for networking and checking availability on port 52415. run\_llm.py and run\_llm.sh provide client-side tools to stream chat completions from the local server. set\_rdma\_network\_config.sh configures macOS network settings for RDMA/Thunderbolt. Finally, test\_trust\_remote\_code\_attack.sh verifies that custom models are safely loaded with trust\_remote\_code disabled to prevent remote code execution.
tmp · high confidence
New topic-based routing system with network integration
The \src/exo/routing\ module has been introduced to provide a structured, topic-based messaging layer. This includes a \Router\ class that manages \TopicRouter\ instances, allowing components to subscribe to and publish messages on specific topics. The system integrates with the underlying networking layer (via \exo\_rs.NetworkingHandle\) to handle serialization, deserialization, and network transmission of messages, supporting policies like 'Always' or 'Minimal' publishing. It also handles connection state updates through \ConnectionMessage\ types.
src/exo/routing · high confidence
New utility modules for process management, power profiling, and event logging
The src/exo/utils package introduces several new capabilities: an AsyncProcess class for managing background processes with captured stdio; a PowerSampler that tracks and reports energy usage separately for prefill and generation phases; a DiskEventLog for persistent, compressed event storage; and supporting utilities for reactive variables, keyed backoff, and dashboard resource discovery.
src/exo/utils · high confidence
Prefill/Decode disaggregation support for MLX models
Users can now enable prefill/decode disaggregation for MLX-based models by setting the ENABLE\_DISAGGREGATION cluster flag. This change introduces a new Advanced dashboard page (dashboard/src/routes/advanced) to manage this configuration, alongside the backend implementation in src/exo/worker/disaggregated and src/exo/worker/engines/mlx/disaggregated. The backend adds a TCP-based protocol and server to stream KV-cache chunks between nodes, an MLX-specific adapter to serialize/deserialize cache states, and a prefill service to handle remote prefill requests. Comprehensive end-to-end and unit tests are included to verify the protocol and adapter logic.
dashboard/src/routes/advanced, src/exo/worker/disaggregated, src/exo/worker/engines/mlx/disaggregated · high confidence
Structured runner diagnostics for Metal GPU and Ring transport errors
The runner now includes a diagnostics module that automatically parses stderr output to identify and classify specific failure modes, such as Metal GPU timeouts and Ring transport/socket errors. This allows the system to provide structured, actionable error information rather than raw logs when these specific runtime issues occur.
src/exo/worker/runner · high confidence
Removals
Removal of initial networking and orchestration scaffolding
The initial scaffolding for distributed inference has been removed, specifically deleting the \networking\ and \orchestration\ modules. This eliminates the previous gRPC-based peer discovery, communication, and node management implementation, including the \GRPCDiscovery\, \GRPCServer\, \StandardNode\, and associated Protocol Buffer definitions. Users relying on this early distributed networking layer will no longer have access to these components.
networking, orchestration · high confidence
Removal of legacy MLX sharded inference components
The previous implementation of the MLX-based sharded inference engine has been removed. Specifically, the \InferenceEngine\ abstract base class, the \MLXFixedShardInferenceEngine\ concrete implementation, and the \Shard\ data class have been deleted from the codebase. This eliminates the fixed-shard processing logic and the associated shard definition structure that were previously used for local model inference.
inference · high confidence
Architecture
Core shared infrastructure and event-driven state management
The src/exo/shared package has been restructured to provide the foundational components for the application's state machine and configuration. This includes a new event-sourcing engine (apply.py) that processes domain events like task creation, deletion, and status updates to mutate the central State object, ensuring consistent state transitions. Configuration and file-system paths are now standardized via constants.py, adopting the XDG Base Directory Specification on Linux while respecting the EXO\HOME environment variable, and defining dedicated directories for logs, models, and caches. Additionally, the package introduces a new asynchronous election mechanism (election.py) for master node selection within the cluster, a robust logging setup (logging.py) with log rotation and compression, and an empty \\init\\_.py to formalize the package structure.
src/exo/shared · high confidence
Behavioural changes
API server implementation and SSE keep-alive support
The API server module in src/exo/api has been implemented, introducing a new FastAPI-based HTTP interface that handles model inference, chat completions, and image generation. A key behavioral change is the addition of Server-Sent Events (SSE) keep-alive functionality via the new keepalive.py module, which prevents client timeouts during long-running prefill operations by periodically sending keep-alive messages. The main.py module integrates various adapters (OpenAI, Ollama, Claude, Responses) and manages instance lifecycles, downloads, and tracing, while types/\_\init\\_.py exports the necessary data structures for these interactions.
src/exo/api · high confidence
Added multiprocessing support for frozen PyInstaller applications
The exo package now supports running as a frozen PyInstaller bundle with multiprocessing enabled. A new entry point in \_\main\\_.py handles the -c flag for inline code execution, which is required by Python's multiprocessing helper processes, and calls freeze\_support() to ensure compatibility with bundled executables.
src/exo · high confidence
Dashboard UI overhaul with EXO Command Center theme and new chat capabilities
The dashboard has been visually redesigned with a new 'EXO Command Center' dark theme featuring yellow accents, CRT scanline effects, and a monospace font. Functionally, the chat interface now supports file attachments (images, text, PDFs) with drag-and-drop and paste handling, and includes a new model selector with smart auto-selection based on memory and task type. The chat experience also features uncertainty visualization via token-level heatmaps, a prefill progress bar with ETA, and the ability to cancel generation during the prefill phase.
dashboard/src/lib/components · high confidence
Download system refactored with explicit offline mode and improved reliability
The download subsystem has been restructured to support an explicit offline mode, allowing the application to function in air-gapped environments without relying on flaky internet checks. The new implementation includes resumable downloads, better handling of Hugging Face rate limits, and fixes for crashes during first start in offline mode. Users will experience more reliable model downloads, with improved progress reporting and the ability to pre-download models to custom directories.
src/exo/download · high confidence
Enable MLX CPU builds and update Apple SDK to 26.2 on Nix
The Nix build environment now supports building MLX on CPU-only systems by enabling the MLX\_BUILD\_CPU option. Additionally, the Apple SDK overlay has been updated to version 26.2, and a custom Metal toolchain is provided to ensure compatibility with the newer SDK versions during compilation.
nix · high confidence
Git LFS hooks added to enforce LFS usage
The repository now includes Git hooks (post-checkout, post-commit, post-merge, pre-push) that automatically invoke Git LFS commands. This ensures that large files tracked by Git LFS are properly managed during standard Git operations, and provides a clear error message if the 'git-lfs' tool is not installed on the user's system.
.githooks · high confidence
Improved context window handling for YARN-based models via RoPE patching
The MLX engine now applies a patch to the underlying \mlx\_lm\ library to align its YARN RoPE (Rotary Positional Embedding) implementation with vLLM's inverse-frequency blending formula. This change ensures better compatibility and stability for models using YARN scaling (such as Qwen3.5), specifically preventing infinite generation loops and memory leaks that could occur with the default implementation. The patch is applied automatically when the MLX engine initializes.
src/exo/worker/engines/mlx/patches · high confidence
Introduce Nix-based build system and update API proxy port for the dashboard
The dashboard now uses a Nix build process (via dream2nix) to manage dependencies and generate the static site, which significantly improves build stability by decoupling the dependency resolution from source code changes. Additionally, the development server's API proxy configuration has been updated to forward requests to the backend service on port 52415, replacing the previous port 8000.
dashboard · high confidence
MLX engine refactored with distributed pipeline support and DeepSeek-V4 tool-calling fixes
The MLX engine has been restructured to support distributed pipeline execution and tensor sharding across multiple devices, introducing custom layer wrappers (PipelineFirstLayer, PipelineLastLayer) to handle inter-rank communication and cache synchronization during prefill and decode phases. This change also adds robust support for DeepSeek-V4 and V3.2 models, including specific DSML encoding/decoding logic to correctly handle tool calls and interleaved reasoning blocks, while fixing race conditions in distributed initialization and memory leaks in the KV prefix cache.
src/exo/worker/engines/mlx · high confidence
Networking stack migrated from libp2p to Zenoh
The Rust networking module has been replaced with a new implementation based on the Zenoh protocol. This change introduces a custom IPv6 multicast-based discovery mechanism (in \discovery.rs\) to locate peers and a compatibility layer (in \swarm.rs\) that maps the previous pub/sub and subscription operations to Zenoh's publish/subscribe and liveliness APIs. Users will now connect to peers via TCP using Zenoh's internal routing, replacing the previous libp2p-based transport and discovery logic.
rust/networking · high confidence
New cluster state model with granular node data and Thunderbolt bridge support
The app introduces a new \ClusterState\ model that replaces the previous \nodeProfiles\ structure with a more granular approach, separating node identity, memory, system metrics, and Thunderbolt bridge status into distinct fields (\nodeIdentities\, \nodeMemory\, \nodeSystem\, \nodeThunderboltBridge\). This change enables the app to display detailed hardware information (such as GPU usage, temperature, and power) and specific Thunderbolt Bridge loop detection status for each node, while maintaining backward compatibility by merging these granular states back into \NodeProfile\ objects where needed.
app/EXO/EXO/Models · high confidence
New master node orchestrator with disk-based event logging and intelligent instance placement
The master node component has been restructured into a new modular architecture (src/exo/master) that introduces a central orchestrator for managing model instances and tasks. This change implements disk-based event logging to replace unbounded in-memory storage, ensuring durability and preventing memory exhaustion during long-running sessions. The master now features an intelligent placement engine that assigns model shards to nodes based on available memory, network topology (including RDMA and Ethernet prioritization), and download progress, supporting both tensor and pipeline parallelism for models like DeepSeek and Gemma 4. Additionally, the system now supports image generation tasks and seamless chat UX with auto model selection, while improving load balancing by routing requests based on in-flight tasks rather than completed ones.
src/exo/master · high confidence
New typed worker state and response models
The worker type system has been replaced with a new set of Pydantic models that define precise states and data structures for inference. Runner lifecycle is now explicitly tracked through a tagged union of states (Idle, Connecting, Connected, Loading, Loaded, WarmingUp, Ready, Running, ShuttingDown, Shutdown, Failed), allowing consumers to distinguish between connection phases and actual execution. Shard metadata now supports Tensor, Pipeline, and CFG parallelism, with specific fields for layer ranges and configuration ranks. Runner responses are strictly typed to include text generation with token-level logprobs, image generation (including partial streaming), tool calls, and prefill progress, ensuring that usage stats and generation details are consistently structured across all output types.
src/exo/shared/types/worker · high confidence
Nix-based Python environment with CUDA and platform-specific build support
The Python build environment is now managed via a new \python/parts.nix\ file using \uv2nix\ and \pyproject-nix\. This change introduces support for Python 3.13 and integrates CUDA libraries (including CUDA 13) for Linux builds, while applying specific build fixes and dependencies for macOS (Darwin) using Apple SDK 26. It also ensures type stubs for the \exo-rs\ package are included to support static type checking with basedpyright.
python · high confidence
PDF support and Safari compatibility fix in dashboard file handling
The dashboard now supports uploading and processing PDF files alongside images, text, and audio. To ensure this works on Safari, a polyfill for ReadableStream's async iterator is included to fix PDF text extraction failures, and the pdfjs-dist worker is configured via CDN.
dashboard/src/lib/types · high confidence
Redesigned downloads page with model×node table view
The downloads interface has been completely redesigned to present download status in a model-by-node table layout. This new view allows users to see the state of downloads (completed, downloading, pending, failed) across multiple nodes simultaneously, displaying detailed progress metrics such as percentage, speed, and estimated time of arrival. It also integrates disk usage information for each node and provides controls to pause, resume, or delete active downloads directly from the UI.
dashboard/src/routes/downloads · high confidence
Test coverage
Added comprehensive test suite for MLX worker engine; Added test coverage for API error handling, command cancellation, streaming shapes, and Claude/Responses adapters; Added test coverage for shared utilities; Added test infrastructure and test plan for the worker module; Added tests for MLX batch generation equivalence and logprob extraction; Added tests for async process, channels, event log, PID file, power sampling, and tagged unions; Added tests for download cancellation, offline mode, and model directory management; Added tests for state application logic; Added tests for the OrderedEventBuffer component; Added unit test infrastructure for the worker module; Added unit tests for image generation cancellation logic; Added unit tests for runner output parsing, event ordering, and supervisor error handling; Added unit tests for the worker plan module; Initial test suite for master node logic, placement, and topology; Integration test infrastructure and test suite.
Dependencies
Initial dependency manifests for Rust, Python, and Dashboard workspaces
This change introduces the foundational dependency configuration files for the project's multi-language structure. It adds a Rust workspace (Cargo.toml/Cargo.lock) that establishes the \exo\_rs\ and \networking\ crates, utilizing Zenoh for networking and PyO3 for Python bindings. It also adds Python project manifests (\pyproject.toml\) for the main application, benchmarking tool, and shared tools, defining dependencies such as FastAPI, MLX, and Transformers. Additionally, it includes the Svelte-based dashboard configuration (\package.json\/\package-lock.json\) and the macOS app's Swift package resolution file.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Score
- CAI 33 → 39 (+6.2)
- Rubric changed (rubric-2026.08.15 → rubric-2026.09.15) — scores are not directly comparable.
Lenses
- Code Health 38 → 33 (-5.1)
- Architecture 96 → 94 (-1.0)
- Maturity 54 → 65 (+11.0)
- Readiness 18 → 30 (+11.6)
- Security 39 → 51 (+11.9)
- Accessibility 58 (new)
Resolved (89)
- Coverage not measured — test suite did not build
- Dimension evaluation failed
- Duplicated block (10 lines × 2) (src/exo/utils/channels.py)
- Duplicated block (10 lines × 3) (src/exo/api/main.py)
- Duplicated block (11 lines × 2) (src/exo/worker/engines/image/models/qwen/adapter.py)
- Duplicated block (11 lines × 2) (src/exo/worker/engines/image/models/qwen/edit_adapter.py)
- Duplicated block (12 lines × 2) (src/exo/api/tests/test_claude_tool_use.py)
- Duplicated block (12 lines × 2) (src/exo/worker/runner/llm_inference/batch_generator.py)
- Duplicated block (13 lines × 2) (src/exo/api/main.py)
- Duplicated block (13 lines × 2) (src/exo/worker/engines/image/models/flux/wrappers.py)
- Duplicated block (13 lines × 2) (src/exo/worker/engines/image/pipeline/runner.py)
- Duplicated block (13 lines × 3) (src/exo/worker/engines/mlx/auto_parallel.py)
- Duplicated block (14 lines × 2) (src/exo/download/coordinator.py)
- Duplicated block (14 lines × 2) (src/exo/worker/tests/unittests/test_runner/test_finish_reason_sse.py)
- Duplicated block (15 lines × 2) (src/exo/worker/tests/unittests/test_runner/test_finish_reason_sse.py)
- Duplicated block (15 lines × 2) (src/exo/worker/tests/unittests/test_runner/test_parse_tool_calls.py)
- Duplicated block (17 lines × 2) (src/exo/worker/tests/unittests/test_mlx/test_kv_prefix_cache.py)
- Duplicated block (21 lines × 3) (src/exo/master/main.py)
- Duplicated block (22 lines × 2) (src/exo/worker/tests/unittests/test_runner/test_dsml_e2e.py)
- Duplicated block (8 lines × 2) (src/exo/download/tests/test_rate_limit_handling.py)
- …and 69 more
New (355)
- API._apply_state (cognitive 21) (src/exo/api/main.py)
- API._collect_image_chunks (cognitive 34) (src/exo/api/main.py)
- API._collect_image_chunks (cyclomatic 18) (src/exo/api/main.py)
- API._collect_text_generation_with_stats (cognitive 16) (src/exo/api/main.py)
- API._generate_image_stream (cognitive 27) (src/exo/api/main.py)
- API.get_placement_previews (cognitive 31) (src/exo/api/main.py)
- API.get_placement_previews (cyclomatic 17) (src/exo/api/main.py)
- AppStore.editImage (cognitive 34) (dashboard/src/lib/stores/app.svelte.ts)
- AppStore.editImage (cyclomatic 33) (dashboard/src/lib/stores/app.svelte.ts)
- AppStore.generateImage (cognitive 55) (dashboard/src/lib/stores/app.svelte.ts)
- AppStore.generateImage (cyclomatic 38) (dashboard/src/lib/stores/app.svelte.ts)
- AppStore.parseSSEStream (cognitive 49) (dashboard/src/lib/stores/app.svelte.ts)
- AppStore.parseSSEStream (cyclomatic 20) (dashboard/src/lib/stores/app.svelte.ts)
- AppStore.refreshConversationModelFromInstances (cognitive 18) (dashboard/src/lib/stores/app.svelte.ts)
- AppStore.refreshConversationModelFromInstances (cyclomatic 16) (dashboard/src/lib/stores/app.svelte.ts)
- AppStore.regenerateChatCompletion (cognitive 24) (dashboard/src/lib/stores/app.svelte.ts)
- AppStore.regenerateChatCompletion (cyclomatic 22) (dashboard/src/lib/stores/app.svelte.ts)
- AppStore.regenerateFromToken (cognitive 44) (dashboard/src/lib/stores/app.svelte.ts)
- AppStore.regenerateFromToken (cyclomatic 36) (dashboard/src/lib/stores/app.svelte.ts)
- AppStore.sendMessage (cognitive 105) (dashboard/src/lib/stores/app.svelte.ts)
- …and 335 more
Changes since last survey
- 1 commits — 1 feature/other, 0 fixes
By area
- (root) — 1 commit
Notable commits
- change: docs: request the mlx extra in the documented setup commands (#2245)
Architecture
- Containers 0 added · 0 removed · contexts 1 added · 0 removed · edges 0 added · 0 removed
Added bounded contexts (1)
- exo
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
exo-explore/exo was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 26 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 21a54c5ea0230a3bec1e1a786d200126c7e34ec6 — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-09659c52afae.