kvcache-ai/ktransformers
40.7
Weak · 19 September 2026
130.7k
lines of production code
Python
with C++
1
measurement over time
What this system is
This system is a modular inference and fine-tuning framework for Large Language Models, centered on the \kt-kernel\ library which provides high-performance, hardware-optimized kernels for Mixture-of-Experts (MoE) architectures. It supports diverse hardware backends including Intel CPUs (via AMX/AVX512), NVIDIA GPUs, and Huawei Ascend NPUs, enabling efficient execution of quantized models in both inference and supervised fine-tuning modes. The project also includes a unified CLI for model management and a Docker-based distribution system, while actively maintaining compatibility with the SGLang runtime.
How it got here
2024 — Architecture consolidation and legacy cleanup
19 changes.
The project restructured its codebase around the kt-kernel and kt-sft modules, replacing the previous monolithic layout with a modernized package shim and installation workflow. This period involved the systematic removal of legacy inference engines, custom CUDA/CPU bindings, and outdated server API implementations to streamline the architecture. Concurrently, documentation was expanded for new hardware support, and third-party dependencies were updated to improve compatibility and stability.
2025–2026 — kt-kernel modularization and CPU/GPU expansion
26 changes.
The project extracted a standalone high-performance inference backend, kt-kernel, to support Mixture-of-Experts models across diverse CPU architectures (AMX, AVX2) and GPU/NPU hardware. This period focused on implementing optimized kernels for various quantization formats, establishing comprehensive benchmarking and testing infrastructure, and introducing a unified CLI and Docker distribution system to streamline deployment.
Features
Add CUDA kernels for GGUF dequantization, GPTQ-Marlin GEMM, and MoE top-k softmax
This change introduces the \kt-kernel/cuda\ module, providing high-performance CUDA implementations for three key operations. First, it adds CUDA kernels to dequantize GGUF weights (q8\_0, q6\_k, q5\_k, q4\_k, q3\_k, q2\_k, and iq4\_xs) into fp32, fp16, or bf16. Second, it integrates a GPTQ-Marlin GEMM kernel for efficient matrix multiplication with 4-bit quantized weights, leveraging tensor cores on compute capability 8.0+ hardware. Third, it includes optimized CUDA kernels for Mixture-of-Experts (MoE) routing, specifically implementing top-k softmax and selection logic. These components are exposed via a Python binding module (\KTransformersOps\) built with PyTorch extensions.
kt-kernel/cuda · high confidence
Initial release of kt-kernel with KVCache operator and performance demo tools
This change introduces the kt-kernel library, starting with the \kt-kernel/operators/kvcache\ module which implements a Key-Value Cache for attention mechanisms. The KVCache supports configurable anchor types (Fixed, Dynamic, QUEST, Block Mean/Max) and retrieval strategies (Layer, KV-Head, Query-Head), and handles data persistence via load/dump functions for both FP16 and Q4\_0 quantized states. To validate and benchmark these capabilities, the \kt-kernel/demo\ directory provides a suite of C++ test programs and Python visualization scripts. These demos exercise the underlying BLAS kernels (via \kblas\ and \blis\) to measure throughput in TFLOPS for int8, bfloat16, and float16 matrix operations, and include specific benchmarks for AOCL GEMM reorder bandwidth.
kt-kernel/demo, kt-kernel/operators/kvcache · high confidence
Introduce CPU backend infrastructure with multi-vendor hardware support
The \kt-kernel/cpu\_backend\ directory now contains the core CPU inference backend, including a task queue, a NUMA-aware worker pool, and shared memory buffer management. This backend provides a unified interface for multiple hardware vendors through adapter headers: CUDA, HIP (ROCm), MUSA, MACA, and Ascend NPU. For Ascend NPU specifically, a dedicated callback worker is implemented to handle asynchronous stream callbacks via the ACL runtime, ensuring proper dispatch of host functions. The implementation includes tests for the task queue and shared memory buffer lifetime to verify correctness.
_kt-kernel/cpu\backend · high confidence
Introduce KVC2 disk-backed KV cache with async I/O and NPU support
The KVC2 module in the balance serve layer now provides a disk-backed key-value cache implementation. This change adds a new C++ library structure (CMakeLists.txt) and source files that implement an asynchronous IO dealer using the Photon framework (supporting io\_uring or libaio) to handle persistent storage of KV cache blocks. It introduces a GPU/NPU page cache layer that manages memory allocation and eviction across multiple devices, with explicit support for both NVIDIA CUDA and Huawei Ascend NPU backends via conditional compilation. The module also includes a metrics subsystem exposing Prometheus counters and histograms for cache operations (insert, lookup, prefix match) and integrates with the existing build system to link against libraries like prometheus-cpp, xxHash, and TBB.
_archive/csrc/balance\serve/kvc2, archive/ktransformers/website · high confidence
Introduce automated Docker image build and distribution system
Added a new Docker packaging infrastructure in the \docker/\ directory, including a comprehensive \Dockerfile\ and supporting shell scripts (\build-docker-tar.sh\, \push-to-dockerhub.sh\, \docker-utils.sh\) along with documentation. This system enables automated building of KTransformers images with standardized naming conventions that embed version information for SGLang, KTransformers, and LLaMA-Factory, as well as CPU variant (AMX/AVX512/AVX2) and CUDA version details. Users can now easily export images to local tar files for offline distribution or push them directly to DockerHub with both full descriptive tags and simplified version tags.
docker · high confidence
Introduce kt-kernel SFT submodule for MoE fine-tuning
The new \kt-kernel/python/sft\ package adds supervised fine-tuning (SFT) capabilities for Mixture-of-Experts (MoE) models, extending the inference-only base with forward/backward passes, LoRA support, and distributed training utilities. It provides a unified autograd function (\KTMoEFunction\) and wrapper classes (e.g., \AMXSFTMoEWrapper\) that offload expert computation to the CPU, supporting BF16, INT8, FP8, and RAWINT4 quantized backends. The module includes architecture detection for models like Kimi K2/K2.5, DeepSeek V2/V3, Qwen MoE, and Mixtral, along with artifact management for loading adapters, saving checkpoints, and handling non-expert caches.
kt-kernel/python/sft · high confidence
Introduce kt-kernel as a standalone high-performance CPU/GPU MoE inference backend
kt-kernel is now available as a separate package providing optimized Mixture-of-Experts (MoE) inference kernels for both CPU and GPU. The package includes a C++ backend with Python bindings (pybind11) and supports multiple hardware backends: AVX2, AVX512, AMX, SYCL (Intel iGPU), Ascend NPU, and CUDA. It features automatic CPU instruction set detection, NUMA-aware execution, and integration with SGLang for CPU-GPU hybrid inference. Installation is supported via PyPI or from source using CMake, with pre-built wheels available for Linux x86-64.
kt-kernel · high confidence
Introduce kt-kernel for CPU-optimized MoE inference and SFT
Adds the kt-kernel Python package, providing high-performance kernel operations for KTransformers. The package automatically detects CPU capabilities (AMX, AVX512, AVX2) and loads the optimal variant at runtime, supporting both inference and supervised fine-tuning (SFT) modes. It exposes the KTMoEWrapper factory interface for CPU-based Mixture-of-Experts operations, including support for various quantization methods (INT4, INT8, FP8, BF16) and LoRA adapters.
kt-kernel/python · high confidence
Introduce unified Python utility layer for MoE backends and weight loading
The \kt-kernel/python/utils\ package now provides a centralized set of Python utilities that expose the underlying C++ MoE kernels and weight loaders to the rest of the system. This includes the \amx.py\ module, which registers and dispatches a wide range of CPU inference backends (such as AMX, AVX2, AVX-VNNI, and SYCL) for various quantization formats including INT4, INT8, FP8, BF16, MXFP4, and MXFP8. It also introduces \llamafile.py\ to support GGUF-based quantized weights via the Llamafile backend, and \loader.py\ to handle SafeTensor and GGUF file parsing with specific support for GGML quantization types. These utilities are exported via \\_\init\\_.py\ to make wrappers like \AMXMoEWrapper\, \LlamafileMoEWrapper\, and loaders like \SafeTensorLoader\ and \GGUFLoader\ readily available for model initialization.
kt-kernel/python/utils · high confidence
Introduces kt-cli command-line interface with model, chat, and configuration management
Adds a new Python-based CLI (\kt-cli\) providing commands to manage models (download, list, path management), run inference servers, interact via chat, quantize weights, benchmark performance, and diagnose environment issues. The package includes shell completion scripts for Bash, Zsh, and Fish to improve usability.
kt-kernel/python/cli/commands · high confidence
Introduction of the KTransformers unified CLI (kt-cli)
Users now have access to a new, unified command-line interface for KTransformers. This tool provides a structured set of commands for managing the environment and models, including installation assistance, system diagnostics (doctor), model downloading and quantization, configuration management, and launching inference servers or interactive chats. It also supports internationalization with English and Chinese interfaces.
kt-kernel/python/cli · high confidence
New AMX-accelerated GEMM backend with BF16 support and AVX512 fallbacks
This change introduces a new low-level linear algebra backend in \kt-kernel/operators/amx/la\ that leverages Intel AMX (Advanced Matrix Extensions) for accelerated matrix multiplication. The implementation includes native BF16 GEMM kernels, optimized buffer management for quantized data (Q4/Q8), and a unified activation function (\act\_fn\) supporting SwiGLU and silu variants with configurable clamping. To ensure compatibility on hardware lacking full AMX instruction sets, the backend provides AVX512F+BW emulation fallbacks for BF16 and FP8 operations, allowing the system to maintain performance across different CPU architectures.
kt-kernel/operators/amx/la · high confidence
New AMX-based MoE operator implementations for multiple quantization formats
The \kt-kernel/operators/amx\ directory now includes a comprehensive set of new header files implementing Mixture-of-Experts (MoE) operators optimized for Intel AMX and AVX512 instruction sets. These additions introduce native support for several quantization schemes: AWQ Int4 with KGroup quantization and zero-point support (\awq-moe.hpp\), native BF16 inference (\bf16-moe.hpp\), MXFP4 (FP4 E2M1) with BF16 activations (\fp4-moe.hpp\), FP8 with 128x128 block-wise scales (\fp8-moe.hpp\), FP8 with per-channel quantization (\fp8-perchannel-moe.hpp\), and Kimi-K2 Int4 with KGroup scales (\k2-moe.hpp\). A staging helper (\fp8\_tp\_staging.hpp\) is also added to support tensor-parallel FP8 loading. Each operator uses a CRTP pattern to provide specialized GEMM dispatch and weight loading logic tailored to its specific quantization format, enabling efficient CPU inference for these model architectures.
kt-kernel/operators/amx · high confidence
New AVX2 CPU backends for BF16, FP8, and INT4 MoE inference
Added a suite of AVX2-optimized kernels for Mixture-of-Experts (MoE) models, enabling high-performance inference on CPUs that support AVX2. The new backends include native BF16 GEMM and MoE operators, FP8 E4M3 MoE with LUT-based dequantization, and GPTQ INT4 MoE with both standard and AVX-VNNI-256 acceleration. The AVX-VNNI-256 INT4 backend also features a packed weight variant that keeps int4 weights resident in memory, significantly reducing VRAM/CPU memory usage compared to pre-unpacked formats. These additions provide users with efficient, quantization-aware execution paths for modern MoE architectures on compatible Intel/AMD hardware.
kt-kernel/operators/avx2 · high confidence
New Ascend NPU deployment scripts and MXFP4 kernels for DeepSeek-V4-Flash
Added a new toolset under \kt-kernel/tools/ascend\_dsv4\ to enable single-NPU deployment of DeepSeek-V4-Flash on Huawei Ascend hardware. This includes \setup.sh\ for building dependencies (kt-kernel, sgl-kernel, CANN ops) and \serve.sh\ for launching the SGLang server with Ascend-specific configurations (e.g., \--attention-backend ascend\). The package also introduces \verify.sh\ for acceptance testing and a Python-based chat client. Additionally, new AscendC kernels (\mxfp4\_fused\_kernel.cpp\) and Python wrappers (\mxfp4\_fused\_op.py\) are provided to efficiently convert MXFP4 weights to W8A8 format on-device, alongside conversion tools (\convert\_mxfp4\_gguf.py\) to prepare per-layer GGUF files for the model.
kt-kernel/tools · high confidence
New CPU kernel operator library for MoE, MLA, and core transformations
The \kt-kernel/operators\ directory now contains a new set of C++ header files implementing the core computational kernels for the CPU backend. This includes tensor-parallel wrappers for Mixture-of-Experts (MoE) and Multi-Head Latent Attention (MLA) layers, as well as standalone implementations for essential operations like RMS normalization, RoPE (Rotary Positional Embeddings), and Softmax. The code introduces a structured profiling system (\sft\_profile.hpp\) to track forward and backward pass stages, and provides common utilities for memory offsetting and task counting, establishing the foundational operator layer for the \kt-kernel\ component.
kt-kernel/operators · high confidence
New documentation for AMX optimization, DeepSeek-V4-Flash, and Intel XPU support
Added comprehensive English documentation for new hardware and model capabilities: \AMX.md\ details Intel Advanced Matrix Extensions optimizations for Qwen 3 MoE models, including memory layout and kernel usage; \DeepSeek-V4-Flash.md\ and \DeepSeek-V4-Flash\_tutorial\_for\_Ascend\_NPU.md\ provide deployment guides for the new model on NVIDIA GPUs and Ascend NPUs respectively; \Docker\_xpu.md\ introduces an Intel GPU Docker guide; and \FAQ.md\ clarifies expected SGLang warnings. These files establish the user-facing reference for these specific features.
doc/en · high confidence
New kt-cli utility modules for model management and system diagnostics
The kt-cli now includes a suite of new utility modules in \kt-kernel/python/cli/utils\ to enhance model discovery, configuration, and user interaction. \model\_scanner.py\ and \model\_discovery.py\ provide logic to scan local directories for model files (safetensors/gguf) and register them, including automatic detection of MoE architectures via \analyze\_moe\_model.py\. \model\_registry.py\ adds built-in support for specific models like DeepSeek-V4-Flash, MiniMax-M2/M3, and Kimi-K2-Thinking with optimized default parameters. System diagnostics are improved by \environment.py\, which detects GPU, CPU, and virtual environment details, while \console.py\ and \input\_validators.py\ offer standardized Rich-based UI components and robust input handling. Additional utilities include \download\_helper.py\ for HuggingFace/ModelScope interactions, \kv\_cache\_calculator.py\ for memory estimation, and \debug\_configs.py\ for inspecting saved run configurations.
kt-kernel/python/cli/utils · high confidence
New kt-kernel benchmark suite for CPU inference kernels
The kt-kernel/bench directory now includes a comprehensive set of Python-based performance benchmarks for the kt-kernel CPU inference library. These scripts allow users to measure throughput and latency for various kernel implementations, including attention mechanisms (bench\_attention.py), linear layers with diverse quantization modes like FP32, FP16, BF16, and GGML formats (bench\_linear.py), and Mixture-of-Experts (MoE) operators. The MoE benchmarks specifically cover native BF16 (bench\_bf16\_moe.py), MXFP4 (bench\_fp4\_moe.py), FP8 block-wise (bench\_fp8\_moe.py), FP8 per-channel (bench\_fp8\_perchannel\_moe.py), and K2 int4 (bench\_k2\_moe\_amx.py) paths. Additionally, benchmarks for Multi-Latent Attention (bench\_mla.py) and write-buffer performance (bench\_k2\_write\_buffer.py) are provided, along with a Makefile for streamlined execution and a .gitignore to manage generated artifacts.
kt-kernel/bench · high confidence
New kt-kernel examples for DeepSeek-V3, AMX MoE, and LLAMAFILE integration
The kt-kernel/examples directory now includes a suite of new scripts and configuration files to support advanced model architectures and optimization techniques. This adds a DeepSeek-V3 model configuration and PyTorch modeling implementation, alongside benchmarking and validation scripts for AMX-accelerated MoE operations (including INT8, BF16, and AWQ quantization). Additionally, a repro harness for LLAMAFILE-based inference is provided to test CPU-GPU expert scheduling and CUDA stream integration, while utility scripts for RoPE and attention validation are also included.
kt-kernel/examples · high confidence
New weight conversion and CPU feature detection scripts for CPU-GPU hybrid inference
The kt-kernel/scripts directory now includes a suite of tools to support heterogeneous expert placement in CPU-GPU hybrid inference. Users can convert CPU-resident 'cold' expert weights to INT4/INT8 with AMX optimization using convert\_cpu\_weights.py (with low-memory and resume modes), and apply GPTQ/RTN quantization to GPU-resident 'hot' experts via convert\_gpu\_weights.py. A new convert\_kt\_to\_sglang\_adapter.py script merges KT SFT fused expert LoRA checkpoints into SGLang-compatible adapter directories. Additional utilities include check\_cpu\_features.py for validating CPU instruction set support (AMX, AVX512, AVX2), compare\_weights.py for validating quantization accuracy across safetensor and .kt formats, and model-specific converters for Kimi-K2 (FP8 to BF16) and general MoE models. These scripts enable users to prepare and validate model weights for efficient hybrid deployment.
kt-kernel/scripts · high confidence
Removals
Removal of Ollama API completion endpoints
The Ollama compatibility layer in the server API has been removed, specifically deleting the \completions.py\ file that previously exposed \/api/generate\, \/api/chat\, \/api/tags\, and \/api/show\ endpoints. Users relying on this local server to emulate the Ollama API for text generation or model information retrieval will no longer have access to these interfaces.
ktransformers/server/api/ollama · high confidence
Removal of custom StaticCache implementation
The custom \StaticCache\ class in \ktransformers/models/custom\_cache.py\ has been removed. This class previously provided a static cache implementation adapted from Hugging Face Transformers, featuring specific optimizations for DeepSeek V2 models and compatibility with \torch.compile\. Its removal indicates a shift away from this custom caching mechanism, likely to rely on standard library implementations or a different caching strategy.
ktransformers/models · high confidence
Removal of legacy CPUInfer Python bindings
The \ext\_bindings.cpp\ file containing the \cpuinfer\_ext\ Python module has been deleted. This removes the previous Python bindings for \Linear\, \MLP\, and \MOE\ operators that relied on the \CPUInfer\ submit/sync pattern, effectively cleaning up the extension interface in this location.
_ktransformers/ktransformers\ext · high confidence
Removal of legacy CUDA extension build and bindings
The legacy CUDA extension module (KTransformersOps) and its associated build configuration have been removed. This eliminates the Python bindings for specific dequantization functions (q8\_0, q6\_k, q4\_k) and the GPTQ Marlin GEMM operation, along with the setup script used to compile these C++/CUDA sources.
_ktransformers/ktransformers\ext/cuda · high confidence
Removal of legacy Marlin and llamafile operator implementations
The custom Marlin quantization pipeline (including GPTQ quantization, Marlin linear methods, and weight repacking) and the llamafile-based CPU inference operators (Linear, MLP, and MoE) have been removed from the operators extension. Users relying on these specific quantization formats or CPU execution backends will no longer have access to these implementations in this module.
_ktransformers/ktransformers\ext/operators · high confidence
Removal of legacy chat completion schema definitions
The file defining the legacy chat completion request and response models (including Message, ChatCompletionCreate, Choice, and streaming chunk structures) has been deleted from the server schemas. This removes the local implementation of these Pydantic models, indicating that the chat completion API handling has been refactored or replaced by a different implementation elsewhere in the codebase.
ktransformers/server/schemas · high confidence
Removal of legacy custom CUDA GGUF dequantization bindings
The legacy custom CUDA implementation for GGUF dequantization has been removed from the \ktransformers\_ext/cuda/custom\_gguf\ module. This change deletes the Python/C++ binding file (\binding.cpp\), the custom header (\custom\_ggml.h\), and the CUDA kernel source (\dequant.cu\) that previously exposed \dequantize\_q8\_0\, \dequantize\_q6\_k\, and \dequantize\_q4\_k\ functions. Users relying on these specific custom GPU dequantization paths in this location will no longer have access to them via this module.
_ktransformers/ktransformers\_ext/cuda/custom\gguf · high confidence
Removal of legacy inference utility module
The \ktransformers/util/utils.py\ file has been deleted, removing the legacy \prefill\_and\_generate\ function and associated helper utilities (such as \set\_module\, \load\_weights\, and \InferenceState\) that previously handled model loading and token generation via CUDAGraphs. Users relying on this specific module for inference workflows will need to migrate to the current supported inference interfaces.
ktransformers/util · high confidence
Removal of legacy ktransformers OpenAI chat completion endpoint
The file implementing the /chat/completions endpoint for the ktransformers backend has been deleted. This removes the specific implementation that handled streaming and non-streaming chat completions by directly invoking the ktransformers inference interface, indicating a shift away from this legacy integration path.
ktransformers/server/api/openai/endpoints · high confidence
Removal of legacy ktransformers and transformers backend interfaces
The legacy \ktransformers.py\ and \transformers.py\ backend interface implementations have been removed from the server. This deletion eliminates the previous static cache and CUDA graph runner integration, as well as the custom text streaming and message formatting logic contained in these files, indicating a shift away from this specific backend architecture.
ktransformers/server/backend/interfaces · high confidence
Removal of legacy local chat and version initialization files
The \ktransformers/\_\init\\_.py\ file, which previously exposed the package version as 0.1.0, and the \ktransformers/local\_chat.py\ script, which provided a standalone CLI for running local chat with DeepSeek and Qwen2-MoE models, have been removed. This cleanup eliminates the older 0.1.0 version marker and the direct local inference entry point, indicating a shift away from this specific legacy interface in favor of newer architectural patterns or server-based interactions.
ktransformers · high confidence
Removal of legacy operator implementations
The \RoPE\, \attention\, \experts\, and \linear\ operator modules in \ktransformers/operators\ have been removed. This deletion eliminates the previous implementations of Rotary Position Embedding, DeepSeek V2 attention, CPU-based Mixture of Experts (MoE), and Marlin quantized linear layers, indicating a shift in the underlying inference engine or operator architecture.
ktransformers/operators · high confidence
Removal of legacy server argument configuration
The \args.py\ file containing the \ConfigArgs\ Pydantic model and its default settings has been deleted from the server backend. This removes the previous configuration structure for parameters such as context size, batch size, sampling options, and device settings, indicating a shift in how the server handles its runtime arguments.
ktransformers/server/backend · high confidence
Removal of legacy server configuration module
The \ktransformers/server/config/config.py\ file has been deleted. This module previously provided a singleton-based configuration loader that parsed \config.yaml\ to initialize server settings (IP, port), database connections, model paths, and web interface options. Its removal indicates that this specific configuration handling logic is no longer used or has been replaced by a different mechanism elsewhere in the system.
ktransformers/server/config · high confidence
Removal of legacy text completion endpoint
The legacy \/completions\ endpoint for text generation has been removed from the OpenAI-compatible API server. This change eliminates the specific handler that previously processed raw prompt inputs and returned either streamed or non-streamed text completions, effectively narrowing the supported API surface to other endpoints (such as chat completions) that remain in the codebase.
ktransformers/server/api/openai/legacy · high confidence
Architecture
Repository restructured around kt-kernel with new top-level package shim
The project has been reorganized to focus on the kt-kernel and kt-sft modules, replacing the previous monolithic layout. A new top-level Python package (ktransformers.py) now serves as a lightweight shim that exposes version information and an SFT availability check, while the actual runtime kernels reside in kt-kernel. The installation workflow has been modernized with a new install.sh script supporting subcommands for installing sglang and kt-kernel independently or together, and git submodules for custom flashinfer and sglang have been added to manage dependencies.
(repo-wide) · high confidence
Behavioural changes
Archive legacy KTransformers framework code
The original integrated KTransformers framework has been moved to the \archive/\ directory for reference, as the project now focuses on the modular \kt-kernel\ and \kt-sft\ components. This change includes the preservation of legacy documentation, build configurations (Dockerfiles, CMakeLists), and source code to maintain historical context while streamlining the active repository structure.
archive · high confidence
Enforce conventional commit messages and auto-format code on commit
The kt-kernel area now installs git hooks that enforce stricter development workflows. The commit-msg hook requires all commit subjects to follow the Conventional Commits format (e.g., \[feat\]: description), blocking non-conforming messages. The pre-commit hook automatically formats staged C/C++ files using clang-format and Python files using Black; if formatting changes are applied, the commit is aborted so the user can review the updates before re-committing.
kt-kernel/.githooks · high confidence
Improved Windows compilation support and expanded quantization handling in llamafile
The llamafile third-party submodule now compiles correctly on Windows (MSVC) by updating preprocessor checks to recognize the \_M\_X64 target across multiple source files (sgemm.cpp, tinyblas\cpu\.cpp, iqk\_mul\mat\.cpp). Additionally, the matrix multiplication logic in tinyblas\_cpu\_sgemm.inc has been extended to accept a wider range of B-type quantizations (including Q8\_0 and Q8\_1) and natively supports MXFP4 GGUF weights via a GGML vec-dot fallback, while ARM builds now utilize a dedicated iqk\_mul\_mat\_arm.inc implementation.
_third\party/llamafile · high confidence
Test coverage
Add per-commit test suite for KT-Kernel; Added AMX kernel unit and latency tests; New CI test infrastructure and kernel validation tests.
Dependencies
Archiving legacy packages and restructuring dependency manifests
The \kt-sft\ package and the main \ktransformers\ website assets have been moved into the \archive/\ directory, effectively removing them from the active build. The main \pyproject.toml\ has been refactored to use dynamic dependency resolution and now requires Python 3.11+, while the \kt-kernel\ package has been introduced with a pinned \torch==2.9.1\ baseline and specific platform constraints for \triton\. Additionally, the \ktransformers/server\ requirements have been updated to relax the \transformers\ version constraint (now \\>= 4.51.3\), remove the \fire\ dependency, and add \openai\, \zmq\, and \psutil\.
(dependencies) · high confidence
Updated SGLang submodule and added custom FlashInfer submodule
The SGLang third-party submodule has been updated to commit 541ddc37, incorporating recent fixes and performance improvements. Additionally, a new custom FlashInfer submodule (commit fd94393) has been introduced to provide specialized inference kernel support.
_third\party · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Baseline
- First survey — no prior run to compare against. CAI 41.
Lenses
- Code Health 67
- Architecture 99
- Maturity 72
- Readiness 38
- Security 56
- Accessibility 27
Changes since last survey
- 300 commits — 224 feature/other, 76 fixes
By area
- doc/en — 60 commits
- (root) — 40 commits
- kt-kernel/operators — 33 commits
- kt-kernel/python — 32 commits
- .github/workflows — 17 commits
- third_party/sglang — 17 commits
- (repo) — 14 commits
- kt-kernel/scripts — 10 commits
- kt-kernel/README.md — 8 commits
- .github/release — 7 commits
- archive/ktransformers — 7 commits
- doc/zh — 6 commits
- kt-kernel/cpu_backend — 6 commits
- kt-kernel/CMakeLists.txt — 5 commits
- kt-kernel/ext_bindings.cpp — 5 commits
- kt-kernel/test — 5 commits
- .github/ISSUE_TEMPLATE — 3 commits
- KT-SFT/ktransformers — 2 commits
- archive/kt-sft — 2 commits
- doc/assets — 2 commits
Notable commits
- fix: Fix CPU Instruction Set and Installation (#1729)
- fix: Fix K2 MoE decode bug in buffer management (#1686)
- fix: Fix Qwen3.5 FP8 load for VL detection (#1857)
- fix: Fix TaskQueue worker thread 100% CPU spin when idle (#1899)
- fix: Fix download link for Kimi-K2-Thinking weights
- fix: Fix duplicate BF16 loader definition (#1984)
- fix: Fix git submodule (#1586)
- fix: Fix image reference in README.md (#1584)
- fix: Fix kt-kernel compile issue (#1595)
- fix: Fix kt-kernel for new wrapper (#1588)
- fix: Fix moe bug. (#1783)
- fix: Fix worker pool idle CPU usage (#1902)
- fix: Fix/sglang kt detection (#1875)
- fix: Reduce CPU memory usage during large chunk prefill (Fixes #1676) (#1683)
- fix: Revert "[doc]: add Kimi-K2.5 deploy&sft guide (#1810)" (#1811)
- fix: Revert "[doc]: update kimi_k2.5 doc (#1823)" (#1825)
- fix: Revert "kt-kernel: enable CPUInfer stream bridge for ROCm (#1918)" (#1925)
- fix: chore: bump sglang submodule for FP16 FE8M0 Marlin fix (#2082)
- fix: [chore]: sync sglang-kt packaging fix
- fix: [doc] fix kt parameters (#1629)
- …and 280 more
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
kvcache-ai/ktransformers was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 19 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit f7607c0f66c643220ff43b822b24483c98c41392 — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-13a154b7f5d1.