Skip to content
CAI
Software that uses CAICheck a score

mlc-ai/mlc-llm

60.6

Adequate · 19 September 2026

57.6k

lines of production code

Python

with C++, C

1

measurement over time

CAI band scale
CAI lens gauges

What this system is

This system is MLC LLM, a framework for compiling and running large language models across diverse hardware backends including CUDA, Metal, Vulkan, WebGPU, Android, and iOS. It provides a unified serving engine that exposes an OpenAI-compatible API for text generation, embeddings, and multimodal interactions, while supporting advanced inference techniques like speculative decoding and grammar-constrained generation. The project includes native mobile applications for Android and iOS, a WebAssembly runtime for browser-based execution, and comprehensive CLI tools for model conversion, quantization, and benchmarking.

How it got here

2023 — Platform consolidation and mobile expansion

25 changes.

The project consolidated its build infrastructure by migrating to scikit-build-core and establishing comprehensive CI pipelines, while simultaneously removing legacy C++ chat modules and custom tokenizers. This period focused on delivering native mobile applications for Android and iOS, introducing a formal Python package with dedicated CLI tools for model delivery and serving, and expanding REST API examples for broader integration.

2024–2026 — MLC LLM Python package and cross-platform serving

52 changes.

This period focused on establishing the MLC LLM Python package and a modular C++ serving engine, introducing comprehensive support for disaggregated inference, speculative decoding, and grammar-constrained generation. It expanded cross-platform capabilities by releasing native iOS and Android SDKs with OpenAI-compatible APIs, while adding WebAssembly runtime infrastructure and extensive test coverage for the new architecture.

Features

Add GSM8K and MMLU evaluation benchmarks

New standalone evaluation scripts have been added for the GSM8K (grade-school math) and MMLU (massive multitask language understanding) benchmarks. The GSM8K evaluator (\gsm8k.py\) supports few-shot learning with chain-of-thought reasoning and flexible answer extraction to assess mathematical problem-solving capabilities. The MMLU evaluator (\mmlu.py\) assesses knowledge across 57 diverse academic and professional subjects by evaluating multiple-choice questions using token log-probabilities to determine the most likely answer choice. Both tools allow users to benchmark model performance on these specific datasets via command-line arguments.

_python/mlc\llm/bench/evaluation · high confidence

Add Node.js and Python REST API access examples

New example scripts have been added for the REST API in both Node.js and Python, allowing users to interact with the mlc\_llm service. The Node.js examples (sample\_client.js, sample\_openai.js, sample\_langchain.ts) demonstrate chat completions (streaming and non-streaming), legacy completions, and LangchainJS integration, requiring Node v18.17.x. The Python examples (sample\_client.py, sample\_openai.py, sample\_langchain.py) provide similar functionality using the requests and openai libraries, plus Langchain features like retrieval-augmented generation (RAG) with Chroma and MLCEmbeddings. Resource files (linux.txt, state\_of\_the\_union.txt) are included to support the document question-answering demos.

examples/rest · high confidence

Add Objective-C wrapper for JSON FFI LLM engine

Introduced a new Objective-C interface (JSONFFIEngine) in the iOS MLCSwift module that exposes the underlying C++ JSON FFI engine to Swift. This wrapper provides methods to initialize the background engine with a Metal device, reload engine configurations, unload resources, reset state, and manage chat completions and background loops, enabling Swift applications to interact with the LLM backend via the updated JSON FFI convention.

ios/MLCSwift/Sources/ObjC · high confidence

Add hero section to the website

The website now features a new hero section on the landing page, including a heading, a link to the GitHub repository, and a 'Get Started' call-to-action. This section also displays a project workflow diagram and includes hover effects for the interactive links.

_site/\includes · high confidence

Add minimal iOS MLCEngine example app

A new minimal iOS application (MLCEngineExample) has been added to demonstrate interaction with the MLC Engine. The app loads a packaged model (Llama-3-8B-Instruct) and performs a chat completion using the OpenAI-style async API, streaming the response text and usage statistics to the user interface. It includes necessary entitlements for extended virtual addressing and memory limits, and serves as a quick testing scaffold for the MLCSwift bindings.

ios/MLCEngineExample/MLCEngineExample · high confidence

Centralized quantization registry and FP8 support for Mixtral experts

The quantization module now uses a centralized registry to manage quantization configurations, exposing a unified factory interface via \make\_quantization\_functions\ that allows model loaders to select from group, FT, AWQ, per-tensor, and block-scale quantization strategies. This change introduces native FP8 support for Mixtral experts, enabling per-tensor quantization with both inference and max-calibration modes, and adds new quantization presets such as \e4m3\_e4m3\_f16\ and \fp8\_e4m3fn\_bf16\_block\_scale\ to the available configuration options.

_python/mlc\llm/quantization · high confidence

Expanded model support with new weight loaders

The model loader module now includes support for a wide range of new architectures, including Baichuan, BERT, ChatGLM3, Cohere (Aya-23), DeepSeek, DeepSeek-V2, Eagle, Gemma, Gemma2, Gemma3, GPT-2, GPTBigCode, GPT-J, GPT-NeoX, InternLM, InternLM2, Llama, Llama4, and LLaVA. These additions provide specific HuggingFace and AWQ weight mapping logic, enabling users to load and run these models within the MLC LLM framework.

_python/mlc\llm/model · high confidence

Experimental embedding API for encoder models

Added a new \mlc\_llm.contrib.embeddings\ module providing experimental support for running encoder-based embedding models. This includes an \MLCEmbeddings\ class that wraps the MLC runtime to perform batched text embedding, along with a LangChain-compatible adapter (\openai.py\) that allows these local embeddings to be used with standard LangChain tools.

_python/mlc\llm/contrib · high confidence

Initial release of the MLC LLM Python package

This change introduces the \mlc\llm\ Python package, establishing the core runtime and command-line interface for MLC LLM. It provides the main entry point via \\\main\\_.py\ with subcommands for compiling, converting weights, generating configs, chatting, serving, packaging, calibrating, and routing. The package exposes \MLCEngine\ and \AsyncMLCEngine\ for model inference, handles library loading and path resolution in \base.py\ and \libinfo.py\, and registers global functions for distributed socket sessions and CUDA profiling.

_python/mlc\llm · high confidence

Initial release of the MLCChat Android application

The MLCChat Android app is now available, providing a native interface for running local large language models on Android devices. The application features a modern UI built with Jetpack Compose and Material Design 3, supporting both light and dark themes. Users can manage a list of pre-configured models (including Phi-3.5, Qwen3, Gemma-2, Llama-3.2, and Mistral-7B) directly from the app, with capabilities to download, pause, and delete model weights. The chat interface supports markdown rendering for assistant responses and allows users to toggle between raw text and formatted views. Additionally, the app supports multimodal interactions by allowing users to pick images from the gallery or take photos with the camera to send to the model.

android/MLCChat · high confidence

Interactive CMake configuration script for build backends

The build system now includes a Python script (cmake/gen\_cmake\_config.py) that interactively prompts users to select enabled hardware backends (CUDA, OpenCL, ROCm, Vulkan, Metal, CUTLASS, CUBLAS, and OpenCL Host Ptr) and generates a config.cmake file. The script automatically enables Thrust when CUDA is selected and handles parent-child backend dependencies, simplifying the initial setup process for developers.

cmake · high confidence

Introduce C++ serving engine with disaggregation, grammar, and speculative decoding support

The \cpp/serve\ directory now contains the core C++ implementation of the MLC LLM serving engine, replacing or supplementing previous Python-side logic. This change introduces structured configuration classes (\EngineConfig\, \GenerationConfig\, \DebugConfig\, \ResponseFormat\) that allow users to control inference parameters such as temperature, top\_p, penalties, and response formats (including JSON schema). It adds support for disaggregated serving via \DisaggConfig\ (handling \prepare\_receive\, \remote\_send\, and \start\_generation\ request kinds with KV window management) and grammar-constrained generation with configurable execution modes (\jump\_forward\ or \constraint\). The engine also implements speculative decoding infrastructure, including a \DraftTokenWorkspaceManager\ for managing draft token states and hidden states, and integrates with the prefix cache for efficient request handling. Data structures for multi-modal inputs (\TextData\, \TokenData\, \ImageData\) and sampling results (\SampleResult\ with logprobs) are defined, enabling detailed output streaming and debugging via \EventTraceRecorder\.

cpp/serve · high confidence

Introduce JSON FFI Engine for Android LLM inference

The Android MLC library now includes a new JSON FFI engine implementation (JSONFFIEngine) that exposes core LLM capabilities such as model initialization, reloading, unloading, and chat completion via JSON-based function calls. This component bridges the Android application layer with the underlying TVM runtime, enabling background processing and streaming responses through a callback mechanism, effectively replacing or supplementing previous engine integration methods.

android/mlc4j/src/main · high confidence

Introduce JSON FFI Engine with OpenAI-compatible API, vision, and function calling support

This change adds the JSON FFI Engine (\cpp/json\_ffi\), a new C++ component that exposes an OpenAI-compatible chat completion interface over JSON. It includes a conversation template system for prompt construction, support for multimodal inputs via base64 image loading and CLIP preprocessing, and function calling capabilities with tool definitions and arguments. The engine wraps the existing \ThreadedEngine\ to handle request scheduling, generation configuration, and streaming responses, enabling external clients to interact with MLC LLM models using standard JSON protocols.

_cpp/json\ffi · high confidence

Introduce JSON FFI engine with OpenAI-compatible streaming interface

A new JSON FFI engine module has been added to provide a pure string-based interface for the MLC LLM Engine, primarily for testing and internal use. This module exposes a \JSONFFIEngine\ that implements an OpenAI-compatible streaming API, handling chat completions via a synchronous queue and background threads. It supports standard parameters such as model selection, temperature, top\_p, tools, and stream options, while strictly requiring streaming mode. The implementation ensures correct serialization of request data using Pydantic's \by\_alias\ option and manages background engine loops for efficient processing.

_python/mlc\_llm/json\ffi · high confidence

Introduce MLC LLM Python package with CLI entrypoint

The Python package is now installable via pip as 'mlc\_llm', providing a formal distribution structure. This includes a console script entry point 'mlc\_llm' that invokes the main module, enabling direct command-line usage. The setup script handles binary library inclusion for non-Conda builds and parses requirements from requirements.txt.

python · high confidence

Introduce MLC LLM benchmarking submodule

A new benchmarking tool is added under \python/mlc\_llm/bench\ to evaluate LLM inference performance. It supports multiple API backends (including OpenAI v1/completions and v1/chat/completions) and various datasets (ShareGPT, LooGLE, WildChat, AzureLLMInference, LLMPerf, JSON schema, and prefix-cache). The tool allows users to configure request rates, concurrent requests, and engine parameters, and outputs detailed metrics including latency, throughput, and server-side statistics to CSV and debug logs.

_python/mlc\llm/bench · high confidence

Introduce MLCEngine with OpenAI-compatible chat API in Swift

The iOS Swift SDK now provides a new MLCEngine class that exposes a chat completion interface aligned with the OpenAI API style. This engine accepts structured chat messages and returns streaming responses via AsyncStream, supporting standard parameters such as model, temperature, top\_p, tools, and stream options. The implementation handles background execution and JSON serialization internally, allowing iOS developers to integrate large language model inference using familiar OpenAI-style request and response structures.

ios/MLCSwift/Sources/Swift · high confidence

Introduce centralized support utilities for MLC LLM

This change introduces a new \python/mlc\_llm/support\ package containing shared utilities for the MLC LLM Python package. It adds auto-detection logic for local devices (auto\_device.py) and compilation targets (auto\_target.py), including support for WebGPU subgroups and multi-arch builds. It also provides helpers for automatically detecting model configurations, weight paths and formats (auto\_config.py, auto\_weight.py), and a robust download cache system with configurable policies (download\_cache.py, constants.py). Additional utilities include argument parsing enhancements, logging, progress bar integration, random seed management, and pre-sharding support for tensor parallelism.

_python/mlc\llm/support · high confidence

Introduce dedicated CPU and GPU samplers for text generation

The serving runtime now uses distinct CPU and GPU sampler implementations to handle token generation. The CPU sampler provides optimized top-p sampling with early-exit logic and pivot-based renormalization, while the GPU sampler leverages FlashInfer for parallel sampling on CUDA devices and supports Metal and Vulkan backends. Both samplers integrate with speculative decoding, allowing draft token verification to run on the same device as the main model, and include support for NVTX benchmarking on CUDA.

cpp/serve/sampler · high confidence

Introduce initial CI infrastructure for building and testing MLC LLM

This change adds the foundational CI pipeline for the project, introducing a Jenkins-based workflow defined in jenkinsfile.groovy. The pipeline now supports building MLC LLM runtime libraries and wheels for CUDA, Metal, and Vulkan backends, as well as running unit tests and model compilation checks for CUDA, Metal, Vulkan, and WASM. A new bash helper script (ci/bash.sh) standardizes Docker container execution, handling GPU passthrough (NVIDIA/Vulkan) and environment variables, while a conda environment file (build-environment.yaml) pins Python 3.11 and build dependencies like CMake and Vulkan headers to ensure reproducible builds.

ci · high confidence

Introduce mlc4j build system for Android

This change introduces the \mlc4j\ module, providing a new CMake build configuration and a Python preparation script (\prepare\_libs.py\) to build the MLC LLM and tvm4j libraries for Android. The build system configures the compilation of \tvm4j\_runtime\_packed\ and \tvm4j\_core\, linking against MLC LLM static libraries and tokenizers, while the Python script automates the CMake generation, build, and installation steps for the Android NDK environment.

android/mlc4j · high confidence

Introduce modular neural network components for LLMs

The \python/mlc\_llm/nn\ package now provides core building blocks for defining large language models, including a \PagedKVCache\ for efficient attention computation, a \MixtralExperts\ module for Mixture-of-Experts layers, and an \RNNState\ class for state-space models. These modules expose the necessary interfaces for model architectures to utilize optimized caching, expert routing, and recurrent state management within the MLC LLM framework.

_python/mlc\llm/nn · high confidence

Introduce multi-GPU loader with pre-sharding support

Added new C++ implementation files (\builtin.cc\ and \multi\_gpu\_loader.cc\) to the \cpp/multi\_gpu\ module, enabling loading-time parameter sharding for pipeline parallelism. This change introduces functions to dispatch operations across GPU groups and handles the reception, preprocessing, and scattering of model parameters across multiple workers, allowing models to be loaded and sharded efficiently across a multi-GPU cluster.

_cpp/multi\gpu · high confidence

Introduce programmable router for microserving endpoints

Added a new Router component in the python/mlc\_llm/router module that dispatches OpenAI API requests to multiple underlying engine endpoints. The router supports 'disagg' and 'round-robin' modes, manages server lifecycle via PopenServer, and handles request translation and response streaming to the configured hosts and ports.

_python/mlc\llm/router · high confidence

Introduce structured Python interface modules for MLC LLM

The \python/mlc\_llm/interface\ package has been reorganized into dedicated modules (\calibrate\, \gen\_config\, \help\, \package\, \router\, \serve\) that expose specific capabilities. Users can now access a dedicated calibration entrypoint for quantizing models, a configuration generator for creating \mlc-chat-config.json\ files, and structured CLI help documentation. The package also provides explicit entrypoints for packaging models for mobile platforms (iOS, Android), running a router for disaggregated or round-robin inference, and serving models via a FastAPI-based server with OpenAI-compatible endpoints, embedding support, and API key authentication.

_python/mlc\llm/interface · high confidence

Introduces new compiler passes for GPU sampling, logit processing, and CUDA graph support

The \python/mlc\_llm/compiler\_pass\ module now includes a suite of new compiler passes that attach specialized GPU kernels and metadata to the compiled IRModule. These passes enable GPU-accelerated sampling (including multinomial sampling, top-p renormalization, and batch verification), in-place logit processing (bias, penalty, and bitmask application) for both CPU and GPU targets, and two-stage softmax with temperature for numerical stability. Additionally, the module adds support for CUDA graph capture by attaching initialization functions and symbolic capture hints, while also providing passes to dispatch Triton-based FP8 matmul kernels and manage KV cache creation strategies (TIR vs. FlashInfer).

_python/mlc\_llm/compiler\pass · high confidence

Introduction of AppViewModel for model and chat state management

The Android app now utilizes a new AppViewModel class to manage application configuration, model states, and chat interactions. This component handles loading app settings from local storage or assets, managing the lifecycle of downloaded models (including deletion and configuration updates), and coordinating with the MLCEngine for chat operations. It introduces structured state management for model lists and chat states, replacing previous ad-hoc handling of these concerns within the UI or other modules.

app · high confidence

Introduction of MLC LLM Protocol Module

A new protocol module has been added to the MLC LLM package, establishing a dedicated location for Pydantic model definitions used in API entry points and configuration. This change introduces a structured convention for organizing these models, separating those appearing in API endpoints from general configuration classes, thereby laying the groundwork for the API layer of the renamed MLC LLM library.

_python/mlc\llm/protocol · high confidence

Introduction of MLC WASM runtime library pack

A new source file, mlc\_wasm\_runtime.cc, has been added to the web/emcc directory to serve as the MLC WASM runtime library pack. This file establishes the compilation configuration for the TVM logging system and defines the COMPILE\_MLC\_WASM\_RUNTIME flag to ensure that unsupported code is excluded from the resulting binary.

web/emcc · high confidence

New C++ tokenizer and streaming components

The \cpp/tokenizers\ directory now includes the core C++ implementations for tokenizer loading, text streaming, and stop-string handling. \tokenizers.cc\ and \tokenizers.h\ provide the \Tokenizer\ class, which supports loading HuggingFace, SentencePiece, and ByteLevelBPE formats (prioritizing HuggingFace/SentencePiece) and exposes encoding, decoding, and vocabulary utilities. \streamer.cc\ and \streamer.h\ introduce \TextStreamer\ for streaming UTF-8-valid decoded text chunks and \StopStrHandler\ for detecting generation stop conditions using KMP partial match tables. These components are registered via the TVM FFI reflection system for use by the serving engine.

cpp/tokenizers · high confidence

New CLI entrypoints for calibration, device checking, and model metadata inspection

The CLI module now includes dedicated command-line entrypoints for new capabilities: \calibrate.py\ adds an FP8 calibration tool that accepts model, dataset, and output arguments to generate calibration data; \check\_device.py\ provides a utility to verify the existence of available devices (CPU/GPU) and report their IDs; and \model\_metadata.py\ allows users to inspect compiled model libraries to report memory usage (parameters and KV cache) and other metadata. These additions complement the existing chat, compile, and serve commands by providing specialized tools for model preparation and diagnostics.

_python/mlc\llm/cli · high confidence

New CLI tools for model delivery, library compilation, and serving

The \mlc\_llm\ package now includes dedicated command-line interfaces for managing model artifacts and running inference. The \delivery\ CLI automates downloading models from Hugging Face, generating configuration, converting weights, and uploading the resulting packages. The \lib\_delivery\ CLI handles continuous compilation of model libraries for various devices and quantizations based on a JSON specification. Additionally, the \serve\ CLI provides a direct entry point to start an inference server with configurable engine parameters, while the \chat\ CLI offers an interactive text-based chat interface. These tools streamline the workflow from model preparation to deployment.

_mlc\llm · high confidence

New Kotlin JSON FFI Engine for Android Chat

The mlc4j Android library now includes a new MLCEngine implementation that communicates with the underlying model via a JSON FFI interface. This change introduces a Kotlin-native engine class that manages background workers for inference and streaming, exposing a chat API compatible with the OpenAI protocol (including support for tools, streaming responses, and usage statistics). This provides a modern, structured way to integrate LLM inference into Android applications using Kotlin coroutines and serialization, replacing or supplementing previous integration methods.

mlc4j · high confidence

New MLC-LLM WebAssembly runtime build infrastructure

The web directory now includes a Makefile, a dependency preparation script (prep\_emcc\_deps.sh), and documentation to build the MLC-LLM WebAssembly runtime. This setup compiles the C++ runtime source into a WebAssembly binary (mlc\_wasm\_runtime.wasm) using Emscripten, enabling WebLLM and other runtimes to directly reuse MLC-LLM source code at runtime.

web · high confidence

New MLCEngineExample Android app using JSON FFI

A new Android example application (MLCEngineExample) has been added to demonstrate the JSON FFI Engine. The app uses Jetpack Compose and Material 3 for its UI, loads the phi-2 model via the MLCEngine API, and streams completions using the OpenAI-compatible protocol. It also includes a build script (bundle\_weight.py) to install the APK and push model weights to a device.

android/MLCEngineExample · high confidence

New Python examples for MLC Engine and Microserving

Added two new Python example scripts: \sample\_mlc\_engine.py\ demonstrates how to initialize and use the \MLCEngine\ for chat completions via the OpenAI API, and \custom\_router.py\ provides a template for implementing a custom router in the Microserving architecture, showing how to handle request translation and coordinate between prefill and decode engines.

examples/python · high confidence

New conversation templates for Llama 4, Qwen 3.5, Ministral 3, and DeepSeek-V3

This update adds conversation templates for several new and updated models, including Llama 4, Qwen 3.5, Ministral 3, and DeepSeek-V3. It also introduces templates for Qwen 3, which includes a new \strip\_reasoning\_in\_history\ option to handle thinking traces, and Hermes 3 for Llama 3.1. Additionally, templates for LLM-jp, Nemotron, and various other models have been added or updated to ensure correct formatting and stop tokens.

_python/mlc\_llm/conversation\template · high confidence

The website now features a new hero section that includes a prominent GitHub link button and a demo container. The styling is fully responsive, with the hero image and layout enlarging on small screens and adapting to breakpoints at 640px, 768px, and 1024px to optimize the viewing experience across different device sizes.

site/assets · high confidence

New low-level operator library for attention, MoE, and speculative decoding

The \python/mlc\_llm/op\ package now provides a comprehensive set of low-level operators for the compiler. This includes an attention module that integrates FlashInfer for high-performance inference with fallback paths, and Mixture-of-Experts (MoE) operators for gating, top-k selection, and quantized GEMV/GEMM using CUTLASS and Triton. Additionally, it introduces batch speculative decoding verification, top-p sampling pivots, and pipeline parallelism boundary markers, enabling advanced model architectures and decoding strategies.

_python/mlc\llm/op · high confidence

New modular server entrypoints with OpenAI compatibility, metrics, and debugging tools

The server now exposes a structured set of API entrypoints in the \python/mlc\_llm/serve/entrypoints\ module. This includes an OpenAI-compatible interface (\openai\entrypoints\) supporting \/v1/completions\, \/v1/models\, and \/v1/embeddings\ (with dimension truncation and base64 encoding), protected by optional API key authentication. Additionally, new endpoints are available for operational visibility and control: \/metrics\ for Prometheus-style metrics, \/debug/\\ routes for event tracing, CUDA profiling, engine metrics dumping, and engine reset, and \/microserving/\*\ endpoints to support disaggregated serving workflows (prep\_recv, remote\_send, start\_generate).

_python/mlc\llm/serve/entrypoints · high confidence

New modular serving engine with embedding support and prefix caching

The serving subsystem has been reorganized into a new \mlc\_llm.serve\ package, introducing \AsyncMLCEngine\ and \MLCEngine\ as the primary interfaces for text generation, alongside a new \AsyncEmbeddingEngine\ for encoder and decoder embedding models. This update adds support for the \/v1/embeddings\ endpoint, implements a \PagedRadixTree\ for efficient prefix caching, and exposes an \EventTraceRecorder\ for request-level performance tracing. The change also includes a synchronous \SyncMLCEngine\ for debugging and refactors the internal FFI APIs to align with the new TVM FFI structure.

_python/mlc\llm/serve · high confidence

New server subprocess management and global context infrastructure

The server module now includes a \PopenServer\ class that launches the MLC LLM server in a background subprocess, allowing for easier debugging and isolated execution with configurable host, port, and engine settings. Additionally, a \ServerContext\ class has been introduced to manage the global state of running models and embedding engines, providing methods to add, retrieve, and terminate engines while ensuring only one context exists at a time.

_python/mlc\llm/serve/server · high confidence

New support utilities for encoding, JSON parsing, and compatibility shims

The cpp/support directory now includes a suite of new utility headers and source files to enhance core functionality and maintain compatibility. This adds a \Result\ class for structured error handling, a \DynamicBitset\ for runtime-sized bit manipulation, and comprehensive UTF-8 and escape-sequence encoding/decoding tools. JSON interaction is improved with a \json\_parser.h\ that provides safe lookup methods returning \Result\ types. To support recent TVM runtime and FFI refactors, compatibility shims are introduced in \module\_vtable.h\ and \threading\_backend.h\, preserving existing call sites for module vtables and parallel execution. Additional utilities include a \ProgressBar\ for status updates, a \RandomGenerator\, file loading helpers, and VLM-specific image processing functions for resizing and padding.

cpp/support · high confidence

Removals

Removal of legacy C++ CLI and ChatModule

The legacy C++ command-line interface (\cli\_main.cc\) and the core chat implementation (\llm\_chat.cc\) have been removed from the codebase. This change eliminates the previous C++-based chat module, which is no longer supported, aligning with the project's shift toward Python-based CLI tools and the new SLM architecture.

cpp · high confidence

Removed custom C++ tokenizer wrapper

The custom C++ wrapper for the tokenizer library, including its C bindings and CMake build configuration, has been removed from the 3rdparty directory. This eliminates the local Rust-based tokenizer implementation and its associated build logic, indicating a shift away from this specific integration point.

3rdparty/tokenizers-cpp · high confidence

Architecture

Refactored engine execution into a modular action-based architecture

The engine's serving loop has been restructured from a monolithic implementation into a modular system of distinct 'EngineAction' objects (prefill, decode, draft, verify, jump-forward). This change introduces a factory function that dynamically assembles the correct sequence of actions based on the engine configuration, enabling support for multiple speculative decoding modes (Eagle, Medusa, and automatic adaptive decoding) and grammar-based jump-forward decoding within a unified execution pipeline.

_cpp/serve/engine\actions · high confidence

Behavioural changes

Adopts scikit-build-core for Python packaging and introduces pre-commit linting

The project has migrated its Python build system from the legacy setup.py to scikit-build-core, enabling modern wheel building and aligning with the CMake-based build infrastructure. To support this transition, the old build.py and setup.py scripts have been removed. Additionally, a pre-commit configuration has been added to enforce code quality using ruff for Python linting/formatting, clang-format for C/C++, and yamllint/taplo for configuration files, ensuring consistent code style across the repository.

(repo-wide) · high confidence

Android TVM runtime logging redirected to Android logcat

The MLC4J Android runtime now routes TVM log messages directly to the Android logcat system. This is achieved by overriding the internal logging implementation in the C++ runtime layer to use \\_\_android\_log\_write\, ensuring that debug and fatal logs from the ML inference engine are visible in standard Android device logs.

android/mlc4j/src/cpp · high confidence

Introduce structured C++ model metadata abstraction

The \cpp/metadata\ location now provides a dedicated C++ abstraction for model metadata, replacing previous ad-hoc handling with a structured \ModelMetadata\ class. This new component parses model configuration from JSON and TVM modules, exposing specific fields such as \max\_batch\_size\, \context\_window\_size\, \sliding\_window\_size\, \attention\_sink\_size\, and \kv\_state\_kind\ (supporting \kv\_cache\, \rnn\_state\, \none\, and \hybrid\ states). It also captures parameter details including dynamic shapes, preprocessing functions, pipeline parallel stages, and embedding metadata, enabling the runtime to accurately interpret model capabilities and memory requirements.

cpp/metadata · high confidence

MLC Chat iOS app restructured with updated model configuration

The ios/MLCChat directory has been restructured, removing legacy Swift and Objective-C implementation files (ChatState, ChatView, LLMChat.mm, etc.) and introducing a new README and mlc-package-config.json. The configuration file now defines a model list including Llama-3.2-3B-Instruct, Gemma-2-2B, Phi-3.5-mini, and Qwen3 variants, specifying their Hugging Face sources, estimated VRAM usage, and inference overrides like prefill chunk size and context window size.

ios/MLCChat · high confidence

New CI task scripts for building and testing MLC-LLM

Added shell and batch scripts in the ci/task directory to standardize the CI workflow. build\_clean.sh removes previous build artifacts, build\_lib.sh handles library compilation with support for CUDA, ROCm, Metal, and Vulkan backends using ccache and auditwheel, build\_win.bat manages Windows-specific builds by resolving linker conflicts, test\_model\_compile.sh installs nightly wheels and runs integration tests for various targets including WASM and iOS, and test\_unittest.sh executes unit tests categorized with pytest markers.

ci/task · high confidence

New centralized parameter loader with HuggingFace support and presharding

The model loading subsystem has been reorganized into a new \mlc\_llm.loader\ package that centralizes weight loading logic. This introduces a \HuggingFaceLoader\ capable of reading PyTorch and SafeTensor formats, including support for sharded index files. A new \make\_standard\_huggingface\_loader\ helper automates the mapping of common HuggingFace weight structures (such as QKV and gate/up concatenations) to MLC model parameters. The loader now supports applying presharding after quantization and exposes a centralized \LOADER\ registry for different weight formats.

_python/mlc\llm/loader · high confidence

Refactor documentation build and deployment scripts

The documentation build process has been restructured to separate concerns: new scripts \build\_mlc\_for\_docs.sh\ and \build\_site.sh\ now handle the CMake/Make build of MLC components and the Sphinx/Jekyll site generation respectively, replacing the previous monolithic approach. The \prep\_deps.sh\ script has been removed, and \gh\_deploy\_site.sh\ and \local\_deploy\_site.sh\ have been updated to invoke these new build scripts. Additionally, a new \check\_url\_validity.py\ script was added to validate links in documentation files.

scripts · high confidence

Restructure iOS app with new state management and model configuration

The iOS app has been restructured to introduce a new \AppState\ and \ModelState\ for managing model downloads and chat sessions, replacing the previous flat \ChatState\ as the primary entry point. This change adds support for loading app and model configurations from JSON files (\mlc-app-config.json\, \mlc-chat-config.json\, \tensor-cache.json\) and enables downloading models from remote URLs. The UI now presents a \StartView\ for managing multiple models, each with its own download progress and deletion options, rather than immediately launching a chat. Additionally, the app entitlements have been updated to include \com.apple.developer.kernel.extended-virtual-addressing\ to support larger memory usage.

ios/MLCChat/MLCChat · high confidence

Restructure iOS build system and add Mac Catalyst support

The iOS build process has been reorganized to support Mac Catalyst (macabi) and iOS Simulator architectures (x86\_64, arm64) via a revamped \prepare\_libs.sh\ script. The project structure was updated to include a new \MLCEngineExample\ Xcode project and reorganized \MLCChat\ workspace files, while the legacy \prepare\_params.sh\ script was removed.

ios · high confidence

Website rebranding to llm.mlc.ai with updated homepage and navigation

The project website has been migrated from mlc.ai/mlc-llm to the new domain llm.mlc.ai, updating the site configuration, CNAME, and base paths. The homepage content has been significantly simplified to focus on an overview and direct links to the new documentation site, replacing the previous detailed local installation instructions. A new privacy policy page has been added, and the main navigation menu now includes a link to the Docs section.

site · high confidence

iOS app project restructured with new UI components and dependencies

The iOS project has been reorganized, moving the project file to a nested directory and updating the Xcode workspace version. This change introduces several new Swift source files including StartView, ModelView, AppState, ModelConfig, ModelState, ImageProcessing, ParamsConfig, AppConfig, and Constants, while removing older files like ThreadWorker and LLMChat.mm. The app now embeds the MLCSwift and MarkdownUI frameworks, enabling markdown text display in the chat interface, and updates the copy files phase to include a 'bundle' directory instead of the previous 'dist' folder.

ios/MLCChat/MLCChat.xcodeproj · high confidence

Test coverage

Add server integration tests for embedding, function calling, and image capabilities; Added Python tests for JSON FFI Engine capabilities; Added comprehensive test suite for MLC LLM serving engine; Added correctness tests for new and updated ML operations; Added integration test for model compilation across multiple targets and quantizations; Added model and quantization unit tests; Added test categorization markers for Python test suite; Added testing utilities and debug tools for MLC LLM; Added tests for AWQ and HuggingFace model loading; Added tests for AWQ and group quantization; Added tests for FasterTransformer dequantize matmul epilogue fusion; Added tests for conversation templates and tokenizer streamers; Added tests for the Python router with multiple endpoint configurations; Added unit tests for JSONFFI conversation template parsing; Added unit tests for MLC LLM support modules and CLI weight conversion; Removal of debug comparison library; Reorganize test infrastructure and remove legacy scripts.

Dependencies

Initial dependency manifests for Android, iOS, Python, Node.js, and Rust components

This change introduces the foundational dependency configuration files for the project's various platforms. For Android, it adds Gradle build files for the MLCChat and MLCEngineExample apps and the mlc4j library, specifying AndroidX, Kotlin, and Compose dependencies. For iOS, it establishes the Swift Package Manager configuration (Package.swift and Package.resolved) for the MLCSwift package, pinning NetworkImage, swift-cmark, and swift-markdown-ui. Python dependencies are defined in pyproject.toml and requirements.txt, listing packages like apache-tvm-ffi, torch, and transformers, while also configuring scikit-build-core for wheel building and ruff for linting. Node.js REST API examples are configured via package.json with dependencies on langchain and openai. Additionally, the Rust tokenizers-cpp dependency is removed from the 3rdparty directory.

(dependencies) · high confidence

Updated 3rd-party submodules and removed SentencePiece-JS

The 3rd-party dependencies have been updated to new submodule commits for TVM, tokenizers-cpp, XGrammar, GoogleTest, and STB, while the sentencepiece-js submodule has been removed from the repository.

3rdparty · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

Baseline

  • First survey — no prior run to compare against. CAI 61.

Lenses

  • Code Health 88
  • Architecture 94
  • Maturity 61
  • Readiness 50
  • Security 70

Changes since last survey

  • 300 commits — 227 feature/other, 73 fixes

By area

  • python/mlc_llm — 157 commits
  • 3rdparty/tvm — 41 commits
  • cpp/serve — 36 commits
  • ci/task — 8 commits
  • (root) — 7 commits
  • android/mlc4j — 7 commits
  • docs/install — 6 commits
  • 3rdparty/tokenizers-cpp — 5 commits
  • android/MLCChat — 5 commits
  • ci/jenkinsfile.groovy — 4 commits
  • docs/compilation — 3 commits
  • cpp/multi_gpu — 2 commits
  • cpp/tokenizers — 2 commits
  • docs/deploy — 2 commits
  • docs/get_started — 2 commits
  • python/setup.py — 2 commits
  • (repo) — 1 commit
  • .github/workflows — 1 commit
  • ci/build-environment.yaml — 1 commit
  • cmake/gen_cmake_config.py — 1 commit

Notable commits

  • fix: Add modules_to_not_convert and fix activation scale name
  • fix: CI FIxes (#3415)
  • fix: Fix discord badge in readme (#2975)
  • fix: Fix follow-up to TVM runtime refactor (#3496)
  • fix: Fix model_task pydantic warning and skip flashinfer on non-linux (#3460)
  • fix: Fix relativ path comment (#3008)
  • fix: Fix the dtype of batch_size to ensure mixtral ir well-formed (#2891)
  • fix: Fix: Use correct tvm.s_tir.Schedule API in compiler passes (#3417)
  • fix: Fixed log prob result generation (#3150)
  • fix: Fixed missing trace_enabled from model constructor, which caused inaccurate profiling reports (#3154)
  • fix: Revert "[Fix] Resolve TVM Compatibility Issue by Removing Unsupported Argument" (#3165)
  • fix: Revert accidental merge (#3402)
  • fix: [Android] Fix header after recent tvm runtime refactor (#3185)
  • fix: [BUG FIXED] fix missing 1/sqrt(d) scaling in attention fallback path (#3426)
  • fix: [Bench] Fix the dataset truncation (#2946)
  • fix: [Bug] Replace alloc_buffer with sblock_alloc_buffer and temporarily bypass CSE (#3454)
  • fix: [C++] Fix DLTensor auto type detection (#3337)
  • fix: [C++] Fix EventTraceRecorder initialization (#3331)
  • fix: [C++] Fix StopStrHandler initialization (#3354)
  • fix: [C++] Fix non-nullable object initialization (#3333)
  • …and 280 more

Architecture

  • 0 containers · 3 bounded contexts · 1 dependency edges (baseline)

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

mlc-ai/mlc-llm was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 19 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit 9fa644f54b04983adea4d0168f49fc6af4a893ba — the exact code this score is about.
  • Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-13a154b7f5d1.