qualcomm/GenieX
67.2
Adequate · 29 September 2026
27.2k
lines of production code
Rust
with Go, C++, C, Python
2
measurements over time
What this system is
GenieX is an inference SDK that enables Large Language and Vision Model execution on Qualcomm hardware, supporting CPU, GPU, and NPU backends via llama.cpp and QAIRT plugins. It provides a unified C API with language bindings for Python, Go, and Android, allowing developers to manage model lifecycles, perform streaming inference, and handle multimodal inputs. The system includes a CLI and server interface for local deployment, featuring an OpenAI-compatible API, robust model downloading from multiple hubs, and comprehensive benchmarking tools for performance validation.
Features
Add bench\_pull example for model download performance tuning
A new \bench\_pull\ example has been added to the \model-manager-core\ crate to help users measure and tune download performance. This utility times a single Hugging Face model pull, printing wall-clock duration and throughput (MiB/s) to the console. It is designed for ad-hoc testing of network configurations, proxy settings, and executor tuning knobs (\GENIEX\DL\\*\ environment variables) by allowing users to specify a model repository and optional quantization hint.
sdk/model-manager/crates/core/examples · high confidence
Android SDK introduces Kotlin model management and multimodal data types
The Android SDK binding now provides a high-level Kotlin API for model lifecycle management and supports multimodal and tool-use capabilities. ModelManagerWrapper exposes idiomatic Kotlin/Coroutines methods (init, pullFlow, list, getPaths, remove, clean, getType, resolveAlias) for downloading and managing models, while new data classes (ChatMessage, VlmChatMessage, ToolCall, VlmCapabilities, ProfilingData) enable structured handling of tool calls, vision/audio content, and detailed generation profiling metrics.
app · high confidence
GenieX SDK C API and plugin infrastructure introduced
The SDK now exposes a unified C API (geniex\\) for LLM and VLM inference, device resolution, and model management, replacing the legacy ml\_ prefix. Users can resolve device aliases (cpu, gpu, npu, hybrid) or specify explicit device lists, with the QAIRT plugin restricted to NPU-only inference. The SDK includes a plugin registry that scans for backend shared libraries, handles UTF-8 token streaming safely, and provides error codes with human-readable messages. A Rust-based model manager is linked via FFI, exposing model pull, remove, list, and detailed lookup capabilities. The build system (CMake) handles cross-platform linking, including Windows-specific symbol exports and Linux version scripts to control the public ABI. CPU feature checks prevent illegal instruction crashes on older ARM boards, and logging is standardized with ANSI color support and configurable callbacks.
sdk/src · high confidence
Initial Android JNI Bridge and SDK Binding
The Android binding is introduced, providing a Kotlin API for the GenieX SDK via a four-layer JNI bridge (Kotlin wrappers, Kotlin JNI interface, C++ bridge, and the core C library). This includes the public \LlmWrapper\ and \VlmWrapper\ for coroutine-based generation, a \ModelManager\ for hub model listing and downloads, and a \GenieXSdk\ entry point that routes SDK logs to Android logcat and supports dynamic plugin registration (e.g., for NPU acceleration). The binding also exposes the QAIRT runtime path, handles device resolution, and ensures correct UTF-8/emoji handling in streamed tokens.
bindings/android · high confidence
Introduce AI Hub source with automatic chipset detection and remote ZIP parsing
The model-manager now includes a new AI Hub source implementation that automatically detects the host chipset (Windows via Adreno GPU/CPU, Linux via device tree, Android via SoC model) to select the correct QAIRT assets, and fetches large remote ZIP archives efficiently by parsing the central directory via HTTP Range requests without downloading the full payload.
_sdk/model-manager/crates/core/src/source/ai\hub · high confidence
Introduce CLI configuration and data directory management
The CLI now manages its own persistent configuration and data directory structure. A new \store\ package provides a singleton \Store\ that resolves the data directory (defaulting to \\~/.cache/geniex\ or a custom path via \--data-dir\/\GENIEX\_DATADIR\) and handles atomic read/write operations for a \config.json\ file. This enables the CLI to persist user settings, such as the \chipset\ key, and resolve the host chipset automatically when not explicitly configured, while ensuring the data directory exists with appropriate permissions.
cli/internal/store · high confidence
Introduce CLI rendering package for progress, spinners, and theming
The new \cli/internal/render\ package provides the visual feedback layer for the CLI, introducing a \ProgressBar\ for download/upload status, a \Spinner\ for loading states (which respects the \NO\_COLOR\ environment variable), and a \Theme\ system that applies colored styling to various output types (info, success, error, commands, flags, etc.) based on terminal capabilities. This includes a Bazel build configuration and unit tests for the theme logic.
cli/internal/render · high confidence
Introduce Go SDK bindings for Qualcomm GenieX
This change adds the initial Go bindings for the GenieX SDK, providing a Go package to interact with the underlying C library. The bindings expose core capabilities including model initialization, device and compute unit resolution, and LLM/VLM generation with support for tool calls and streaming callbacks. It also includes a model manager for pulling models from various hubs (Hugging Face, ModelScope, Docker Hub) and resolving aliases, along with utilities for power mode configuration and crash handling.
bindings/go · high confidence
Introduce Python FFI bindings for the GenieX SDK
The Python bindings now include a new FFI layer (bindings/python/geniex/\_ffi) that exposes the underlying C SDK to Python via ctypes. This adds support for initializing and deinitializing the runtime, managing device and plugin discovery, and performing LLM and VLM inference (including chat templates, generation, and KV-cache operations). It also exposes configuration structures for generation, sampling, and model settings, as well as logging integration and error handling. Users can now interact with the native GenieX library directly from Python with structured access to model info, profiling data, and tool calls.
_bindings/python/geniex/\ffi · high confidence
Introduce Python generation API with streaming support
New Python bindings are added under \bindings/python/geniex/generation\, exposing a \GenerationConfig\ dataclass for model parameters (including a new \sliding\_window\ option for context eviction), a \GenerateOutput\ type, and a \TextIteratorStreamer\ class that enables token-by-token streaming via a background thread and allows generation to be cancelled mid-stream.
bindings/python/geniex/generation · high confidence
Introduce QAIRT plugin for LLM and VLM inference
Adds the QAIRT plugin to the SDK, providing new QairtLlm and QairtVlm classes that implement the ILlm and IVlm interfaces. The plugin supports tool calls, sampler configuration, power modes, and system prompts, and includes utilities for handling QAIRT runtime paths, QNN libraries, and Windows-specific file path encoding.
sdk/plugins/qairt/include · high confidence
Introduce QAIRT plugin for Qualcomm NPU inference
Adds the \geniex\qairt\ plugin (built from \sdk/plugins/qairt/src\) to enable LLM and VLM inference on Qualcomm NPUs via the QNN runtime. The plugin registers a single NPU device, loads model bundles by resolving context-binary shards from \ctx-bins\ (avoiding incorrect \\.bin\ globs), and supports vision encoders, system prompts, and power-mode configuration. It also includes a log-filtering sink to surface QNN backend logs through the SDK logging system.
sdk/plugins/qairt/src · high confidence
Introduce QDC benchmark infrastructure with shared primitives and model catalog
The QDC benchmark location now provides a complete, self-contained suite for running performance and accuracy benchmarks on Qualcomm Device Cloud targets. A new shared module (\_qdc.py) centralizes QDC client creation, job submission with quota-aware retry logic, status polling, and log retrieval, ensuring consistent behavior across the benchmark runner and pytest harness. The benchmark runner (run\_qdc\_jobs.py) orchestrates artifact uploads, job execution across Linux, Windows, and Android platforms, and generates markdown reports, including a new accuracy grading mode that evaluates generated text against a committed prompt set. A new model catalog (bench-models.json) defines supported models, specifying plugins, target devices, download URLs, and configuration for VLM and speculative decoding scenarios.
sdk/benchmark/qdc · high confidence
Introduce Rust-based C FFI for the GenieX model manager
The model-manager FFI layer has been rewritten in Rust to provide a stable C interface for the GenieX SDK. This change introduces new APIs for initializing the model store (\geniex\_model\_init\), detecting host chipsets (\geniex\_model\_detect\_chipset\), and listing supported chipsets (\geniex\_model\_list\_chipsets\). It also adds capabilities to list and query models from AI Hub (\geniex\_model\_list\_hub\, \geniex\_model\_query\), pull models with progress callbacks (\geniex\_model\_pull\), and manage local model paths and details (\geniex\_model\_get\_paths\, \geniex\_model\_get\_detailed\). The FFI now handles hub routing logic (including Docker Hub and ModelScope) and provides structured error codes for better debugging.
sdk/model-manager/crates/ffi · high confidence
Introduce Rust-based model-manager core with multi-hub support and structured error handling
The model-manager core is now implemented in Rust, providing a new \Store\ that manages model caching, downloading, and manifest resolution. Users can now pull models from HuggingFace, ModelScope, AI Hub (Qualcomm), and Docker Hub, with automatic routing based on model name prefixes (e.g., \qualcomm/\, \ai-hub-models/\, \docker.io/\). The system supports chunk-granular resume for large downloads, honors environment variables for endpoints (\HF\_ENDPOINT\, \MODELSCOPE\_ENDPOINT\, \GENIEX\_AIHUBBASEURL\) and tokens (\GENIEX\_HFTOKEN\), and uses a structured error type for better recovery. Manifest inference now handles GGUF quantization detection via header probing and config.json analysis, with a defined priority order for quant selection. The store layout uses \geniex.json\ manifests and \.inflight\ directories for atomic updates, ensuring consistency across processes via file locking.
sdk/model-manager/crates/core/src · high confidence
Introduce geniex-bench C inference benchmark tool
A new C-based benchmark binary, geniex-bench, is added to the SDK to measure inference performance (TTFT, prefill/decode throughput) and accuracy. It supports both LLM and VLM workloads across the llama\_cpp and qairt plugins, with flag naming aligned to llama-bench for familiarity. Key capabilities include resolving models via the model-manager (supporting local paths and hub IDs), running in single-cell or matrix modes, and providing an accuracy mode that applies the bundle's chat template to prompt files. It also features a prefill-only logits mode for on-target accuracy metrics and allows overriding the QAIRT runtime via --qairt-lib.
sdk/benchmark · high confidence
Introduce internal readline and audio recording libraries for the CLI
This change adds the \cli/internal/readline\ and \cli/internal/record\ packages to the codebase. The readline library provides a custom terminal input handler supporting cursor movement, history navigation, and a dim placeholder text when the input is empty, with platform-specific terminal handling for Linux and Windows. The record library introduces functionality to capture audio using the \sox\ command-line tool, offering both file-based recording and a stream-based recorder that outputs raw 16-bit/32-bit float samples at 16kHz for further processing.
cli/internal/readline · high confidence
Introduce model management, SDK fetcher, and profile output APIs
The Python bindings now include a high-level model manager (geniex.model\_manager) for pulling, querying, and caching models from various hubs, an install-time SDK fetcher (\_sdk\_fetch.py) that downloads prebuilt SDK binaries with a progress bar and supports CPU-only variants via GENIEX\_SDK\_VARIANT, and a ProfileData class in geniex.generation.output that exposes detailed generation metrics (time, speed, tokens) with human-readable units.
python · high confidence
Introduce server-side keep-alive model caching and parameter resolution
The server now caches loaded models to avoid reloading them for every request, significantly reducing latency for repeated interactions. This change introduces a keep-alive service that manages a single-model cache, automatically resetting the model's internal state (such as the KV cache) when a new conversation session begins, ensuring conversations remain isolated. It also centralizes the resolution of model parameters—such as compute device, context length, and power mode—into a dedicated service layer that validates inputs and handles runtime-specific constraints (e.g., zeroing context length for non-llama\_cpp runtimes).
cli/server/service · high confidence
Introduce the GenieX Python SDK with unified model loading and CLI
The \bindings/python/geniex\ directory now contains the complete GenieX Python SDK, providing a public API for loading and running LLMs and VLMs. Users can instantiate models via \AutoModelForCausalLM\ and \AutoModelForVision2Seq\, which automatically detect model types and resolve device mappings (e.g., \cpu\, \gpu\, \npu\, \hybrid\) to the appropriate backend plugins. The SDK includes a \model\_manager\ for handling model downloads with \tqdm\ progress bars, a CLI entry point (\geniex-py\) for chat and version reporting, and support for multimodal inputs (images/audio) in VLM chat. It also exposes runtime details like plugin versions, compute units, and QAIRT runtime paths, while ensuring robust handling of device aliases and power modes.
bindings/python/geniex · high confidence
Introduces server middleware for CORS handling and GIL-based request serialization
The server now includes a new middleware package that enforces Cross-Origin Resource Sharing (CORS) policies for browser clients and serializes API requests using a Global Interpreter Lock (GIL). The CORS middleware sets standard access-control headers and handles preflight OPTIONS requests, while the GIL middleware ensures that only one request is processed at a time, preventing model resources from being freed while a generation is in flight. This change supports the CLI server's ability to handle concurrent browser-based interactions safely.
cli/server/middleware · high confidence
New C API for model management via Rust model-manager
The SDK now exposes a new C API (geniex\_model.h) backed by a Rust model-manager implementation, enabling users to initialize the model cache, resolve local file paths for downloaded models, list cached models with detailed metadata (including quantization levels), look up individual model details, and remove or clean cached models. This API is linked into libgeniex.so when the GENIEX\_MODEL\MANAGER CMake option is enabled, with a linker script ensuring only the public geniex\\* symbols are exported to maintain ABI stability.
sdk/model-manager · high confidence
New CLI HTTP downloader with secure redirect handling and proxy support
The CLI now includes a new HTTP downloader component that handles chunked downloads with automatic retries on any network error, not just timeouts. It securely strips authentication tokens when following redirects to different hosts to prevent credential leakage to third-party servers, while preserving them for same-host redirects. The downloader also respects HTTP\_PROXY environment variables for traffic routing through corporate or custom proxies.
cli/internal/downloader · high confidence
New CLI server utility package for tool-call parsing and image handling
A new \cli/server/utils\ package has been added to handle common CLI and server-side utilities. It introduces a \ToolCallScanner\ that parses streaming model responses for tool calls across multiple model-specific formats (Gemma 4, GPT-OSS, LFM2, MiniCPM5, Qwen 3/3.5, and standard JSON), ensuring correct extraction of function names and arguments. Additionally, it provides a \SaveURIToTempFile\ utility that downloads or reads local/data URIs and automatically converts WebP images to PNG for SDK compatibility, along with session hashing logic for conversation continuity.
cli/server/utils · high confidence
New Linux release infrastructure with pre-flight checks and CPU-only support
This change introduces the build and release assets for the GenieX CLI on Linux ARM64. It adds a Bazel build definition (BUILD.bazel) that constructs a Debian 13-based Docker image, including an entrypoint script that dynamically resolves and symlinks Qualcomm host libraries at runtime. A new install.sh script allows users to download and install the CLI from S3, with support for a CPU-only build variant via the --cpu-only flag. Additionally, a check.sh pre-flight script is provided to verify system prerequisites, including required shared libraries and minimum Qualcomm driver versions for GPU and NPU hardware.
cli/release/linux · high confidence
New Windows ARM64 onboarding notebook for GenieX
A new Jupyter notebook (examples/python/windows.ipynb) has been added to guide users through setting up the GenieX Python library on Windows ARM64 devices. It provides step-by-step instructions for verifying the Python architecture, installing ARM64 Python if needed, configuring the virtual environment and Jupyter kernel, and running example LLM and VLM inference code.
examples/python · high confidence
New chat and completion server handlers with streaming, tool calls, and reasoning support
The server now includes a dedicated handler package that implements the OpenAI-compatible chat completions and legacy completions endpoints. This adds support for streaming responses, structured tool calls (including round-trip handling and parallel calls), and separated reasoning content (e.g., deepseek-style thinking blocks). It also introduces per-request prefill/decode timing metrics in responses, configurable compute parameters (nctx, ngl, compute, vit\_compute, power\_mode), and VLM support with base64 audio decoding. The implementation is backed by comprehensive unit tests covering message building, streaming deltas, tool call parsing, and reasoning separation.
cli/server/handler · high confidence
New chunked download executor with resume support and concurrency controls
The model-manager now uses a new Rust-based executor in the core SDK to handle file downloads. This executor supports parallel chunked downloads with a \.progress\ bitmap for resuming interrupted pulls, ensuring compatibility with the existing Go CLI. It introduces configurable concurrency via environment variables (\GENIEX\_DL\_FILE\_CONCURRENCY\, \GENIEX\_DL\_CHUNK\_CONCURRENCY\) and handles various sources including HTTP, local files, and deflated archives. The implementation includes robust error handling, such as keeping sibling chunks alive if one fails, and provides detailed progress reporting to the UI.
sdk/model-manager/crates/core/src/executor · high confidence
New model sources: Docker Hub, ModelScope, and enhanced local filesystem support
The model manager now supports pulling models from Docker Hub (specifically the \ai/\*\ namespace using the Docker Registry HTTP API V2) and ModelScope, in addition to the existing HuggingFace and AI Hub sources. The local filesystem source has been expanded to automatically detect and handle AI Hub extracted directories and local AI Hub \.zip\ archives, alongside the existing HuggingFace GGUF directory layout.
sdk/model-manager/crates/core/src/source · high confidence
Repository initialization with Bazel build system and open-source governance files
The repository has been initialized with a Bazel-based build system (version 9.0.2), including configuration files (.bazelrc, MODULE.bazel) and a .clang-format standard. It also introduces standard open-source governance and developer experience files: a BSD 3-Clause license, a Contributor Covenant Code of Conduct, a SECURITY.md for vulnerability reporting, a CONTRIBUTING.md guide, a CODEOWNERS file, and a repolint.json for automated structure checks. Submodules for third-party dependencies (llama.cpp, geniex-qairt) are configured, and .gitignore is updated to exclude local build artifacts and IDE files.
(repo-wide) · high confidence
Windows installer release target and silent-upgrade support
A new Windows release build target has been added to the CLI, producing an InnoSetup-based installer (geniex-cli-setup.exe) via Bazel. The installer now detects existing installations and automatically uninstalls them before proceeding, preventing the silent-install hang that occurred when upgrading. It also version-stamps the installed icon filename to avoid stale icon caching issues in the Windows shell, and registers the CLI in the user's PATH and App Paths registry keys.
cli/release/windows · high confidence
Removals
Removal of GenieX CLI and SDK build infrastructure
The GenieX CLI tool and its associated SDK build definitions have been removed from the repository. This change deletes the Go-based CLI implementation (including the main command, configuration module, and Bazel build rules) and removes the Bazel build configurations for the GenieX SDK (including llama.cpp and OpenCL build files). Users will no longer be able to build or run the GenieX CLI or SDK using the previous Bazel-based build system.
geniex-cli, geniex-sdk · high confidence
Removal of Go SDK binding and associated build files
The Go language binding for the GenieX SDK has been removed. This change deletes the \geniex-sdk-bindings/go\ directory, including the \ml.go\ source file (which previously exposed a \Version\ function) and its \BUILD.bazel\ build configuration. The parent \geniex-sdk-bindings\ build file and the Android and Python binding build files have also been deleted, indicating a restructuring or cleanup of the SDK bindings module structure.
geniex-sdk-bindings · high confidence
Behavioural changes
Automated synchronization of Python sibling package distributions
A new build helper script, \bindings/sync\_siblings.py\, has been added to automatically mirror the canonical Python source files and generate tailored \pyproject.toml\ configurations for the \geniex-llama-cpp\ and \geniex-qairt\ sibling distributions. This ensures that the specific backend-focused packages remain consistent with the main \geniex\ source while maintaining their distinct package names and descriptions, streamlining the creation of separate sdist archives for each backend variant.
bindings · high confidence
CLI server startup hints network exposure and fixes root redirect
The CLI server now warns users when it binds to a loopback address (e.g., localhost or 127.0.0.1), suggesting they restart with --host 0.0.0.0 to expose the service on the network. Additionally, accessing the root URL now redirects to /docs/ui/ to ensure the Swagger UI and favicon load correctly.
cli/server · high confidence
Centralized CLI error handling and improved user feedback
The \cli/cmd/geniex/common\ package now centralizes error reporting and user interaction logic. Users will see friendly, actionable hints instead of raw error messages for common issues like model load failures, context length exceeded, and hub connectivity problems. The CLI also automatically resets the conversation when the context window is exceeded, and on Windows, the console now correctly supports UTF-8 output.
cli/cmd/geniex/common · high confidence
GenieX CLI restructured with Bazel build and enhanced model management
The GenieX CLI has been rebuilt using Bazel, introducing a new \BUILD.bazel\ configuration and a \go\_cgo\_test\ wrapper for Windows CGO testing. The command-line interface now features a dedicated \config\ command for managing settings like chipset selection, and the \list\ command supports machine-readable \--format csv\ output. Model management capabilities have been expanded: the \pull\ command now supports multiple hubs (HuggingFace, ModelScope, Docker Hub, localfs) and displays model size and precision details, while \remove\ and \clean\ commands now require user confirmation. The \run\ command has been refactored to use the OpenAI Go SDK for server communication, and the \update\ mechanism now fetches release assets from an S3 index instead of GitHub releases.
cli/cmd/geniex · high confidence
Introduce HTP session management and platform-specific tuning for llama.cpp
The llama.cpp plugin now manages Qualcomm HTP (Hexagon Tensor Processor) sessions to prevent crashes during model handoffs, using a reference-counted guard to release and reacquire DSP sessions when the backend is present. It also applies platform- and device-specific defaults for threadpool pinning, polling, and batch sizes (Linux, Windows, Android), and routes internal logs through the GenieX logging system with adjusted severity levels.
_sdk/plugins/llama\cpp/src · high confidence
Introduces chipset-aware defaults and Hugging Face token resolution in CLI configuration
The CLI now automatically adjusts compute and batch-size defaults based on the detected hardware: on RB3 Gen 2 (Qualcomm QCS6490) devices, the compute unit defaults to CPU and the GPU ubatch is capped at 256 MiB to prevent allocation overruns. Additionally, the configuration layer now resolves Hugging Face tokens with a clear precedence chain, preferring the GENIEX\_HFTOKEN environment variable, falling back to the HF\_TOKEN environment variable, and finally reading from the local \~/.cache/huggingface/token file if neither is set.
cli/internal/config · high confidence
Python bindings split into three installable distributions with install-time SDK download
The Python bindings are now distributed as three separate source distributions: \geniex\ (includes both llama.cpp and QAIRT plugins), \geniex-llama-cpp\ (llama.cpp only), and \geniex-qairt\ (QAIRT only). These packages are pure Python and do not ship prebuilt native libraries; instead, they download the required platform-specific SDK binaries at install time via HTTP Range requests from an S3 mirror, falling back to GitHub Releases. The console script has been renamed to \geniex-py\ to avoid conflicts with the Go binary, and the build system now uses a PEP 517 wrapper to support stripped Python 3.12 environments that lack \tomllib\.
bindings/python · high confidence
QDC benchmark runner now uses model-manager for model downloads on Windows
The Windows QDC benchmark entry script (run\_windows.ps1) has been updated to resolve models via the local model-manager C API instead of using Invoke-WebRequest. This change prevents out-of-memory crashes when downloading large GGUF files (over \~7 GB) on Windows devices, ensuring the benchmark can run reliably on hardware like the Snapdragon X2 Elite. The script also cleans up stale result files from previous runs to prevent data pollution.
sdk/benchmark/qdc/windows · high confidence
Split Python bindings into separate packages for llama-cpp and QAIrt backends
The Python bindings have been split into distinct distribution packages: \geniex-llama-cpp\ and \geniex-qairt\. Each package now has its own \setup.py\ driver that configures the build for its specific backend (\llama-cpp\ or \qairt\) while sharing common setup logic. This separation allows users to install only the backend they need, rather than a monolithic package containing all backend implementations.
bindings/python-llama-cpp, bindings/python-qairt · high confidence
Fixes
1 commit (1 fix) fixing cli/internal/thinkfsm
A fix in cli/internal/thinkfsm — 1 commit (1 fix), 3 files.
cli/internal/thinkfsm · medium confidence · unverified
Hexagon power mode control and session lifecycle fixes; Adreno OpenCL kernel and memory-split workarounds
The SDK patches now allow controlling the Hexagon HTP power profile via the GGML\_HEXAGON\_POWER\_MODE environment variable, with session lifecycle managed by new release/reacquire hooks to support plugin handoffs. Adreno OpenCL kernels are patched to avoid constant-register limits on devices like the 643, and the xmem GEMM path is disabled by default. Additionally, tensor splitting no longer crashes when devices report zero free memory.
sdk/patches · high confidence
Test coverage
Add QDC device test harness for Linux, Windows, and Android platforms; Added CaptureOutput helper for CLI tests; Added Python binding contract and integration tests; Added Windows ARM64 QDC pytest harness; Added comprehensive test coverage for the Rust model manager core; Added on-device benchmark tests for QDC Android phones; Added unit tests for Python bindings error handling, model caching, and SDK variant selection; New SDK end-to-end test suite with device-specific matrices.
Dependencies
Add geniex-qairt and llama.cpp as git submodules
The repository now tracks the geniex-qairt and llama.cpp projects as git submodules, initializing them at specific commits (43ac75e and 9425611 respectively). This change integrates these third-party libraries into the build system, enabling the use of their functionality within the main project.
third-party · high confidence
New Android AAR binding and updated Go CLI dependencies
This release introduces a new Android binding packaged as an AAR (Android Archive), including build configuration for compile SDK 35 and NDK 29, and automatically packaging prebuilt native libraries (libgeniex.so, QNN runtime, and QAIRT plugins) into the APK. The Go CLI has been restructured under the \github.com/qualcomm/GenieX/cli\ module (Go 1.25.9) with a comprehensive set of dependencies including \gin\, \cobra\, \sonic\, and \openai-go/v3\, while the previous \geniex-cli\ and \geniex-sdk-bindings/go\ modules are removed. Additionally, the Rust model-manager now uses \rustls\ with platform-specific TLS root verification (OS store on non-Android, bundled webpki-roots on Android) and the Python package is updated to use \setuptools\>=77\ with a shim for \tomllib\ on Python 3.12.
(dependencies) · high confidence
SDK includes updated with build configuration and fmt library headers
The SDK now includes a new \build\_config.h\ header that defines plugin identifiers (\llama\_cpp\, \qairt\) and exposes the bridge version. Additionally, the SDK bundles the fmt library (version 11.2.1) and the Howard Hinnant date library as vendored headers under \sdk/include/external/\, providing formatting, date/time, and OS-specific utilities to the SDK consumers.
sdk/include · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Score
- CAI 65 → 67 (+2.2)
- Rubric changed (rubric-2026.09.9 → rubric-2026.09.17) — scores are not directly comparable.
Lenses
- Code Health 87 → 88 (+0.2)
- Architecture 98 → 97 (-1.1)
- Maturity 76 → 76 (+0.5)
- Readiness 73 → 60 (-13.1)
- Security 49 → 61 (+12.6)
- Performance 100 (new)
Resolved (5)
- Change coupling: GenieXSdk.kt ↔ _api.py (bindings/android/app/src/main/java/com/geniex/sdk/GenieXSdk.kt)
- Off-boarding risk: anonymized user #1
- Off-boarding risk: anonymized user #2
- Off-boarding risk: anonymized user #3
- utils.listEnd (cyclomatic 16) (cli/server/utils/toolcall_lfm2.go)
New (50)
- Duplicated block (10 lines × 2) (cli/server/handler/chat_response.go)
- Duplicated block (10–11 lines × 2) (cli/server/handler/chat_messages.go)
- Duplicated block (12 lines × 2) (bindings/go/llm.go)
- Duplicated block (13 lines × 2) (cli/server/handler/chat_messages.go)
- Duplicated block (14–16 lines × 2) (bindings/android/app/src/main/java/com/geniex/sdk/LlmWrapper.kt)
- Duplicated block (16–17 lines × 2) (bindings/go/model_manager.go)
- Duplicated block (21 lines × 2) (cli/cmd/geniex/infer.go)
- Duplicated block (31–34 lines × 2) (bindings/python/geniex/auto.py)
- Duplicated block (5 lines × 2) (cli/server/handler/chat.go)
- Duplicated block (6 lines × 2) (bindings/python/geniex/modeling.py)
- Duplicated block (6 lines × 2) (cli/server/handler/chat_messages.go)
- Duplicated block (6 lines × 2) (cli/server/handler/completion.go)
- Duplicated block (7 lines × 2) (cli/cmd/geniex/common/process.go)
- Duplicated block (8 lines × 2) (bindings/go/llm.go)
- Duplicated block (8 lines × 2) (bindings/python/geniex/modeling.py)
- Duplicated block (8 lines × 2) (cli/server/handler/chat_messages.go)
- Duplicated block (8 lines × 2) (cli/server/handler/chat_stream.go)
- Duplicated block (8–9 lines × 2) (cli/server/handler/chat.go)
- Duplicated block (9 lines × 2) (bindings/go/llm.go)
- Duplicated block (9 lines × 2) (cli/cmd/geniex/infer.go)
- …and 30 more
Changes since last survey
- 65 commits — 30 feature/other, 35 fixes
By area
- (repo) — 20 commits
- cli/server — 10 commits
- sdk/plugins — 10 commits
- sdk/benchmark — 7 commits
- sdk/patches — 5 commits
- (root) — 3 commits
- sdk/model-manager — 2 commits
- .github/actions — 1 commit
- bindings/android — 1 commit
- bindings/python — 1 commit
- cli/cmd — 1 commit
- cli/internal — 1 commit
- examples/sdk — 1 commit
- third-party/geniex-qairt — 1 commit
- third-party/llama.cpp — 1 commit
Notable commits
- fix: Merge pull request #1342 from mansiverma897993/fix/completions-host-stop
- fix: Merge pull request #1458 from mansiverma897993/fix/input-audio-base64
- fix: Merge pull request #1459 from qualcomm/fix/update-chunk-download-retry
- fix: Merge pull request #1460 from qualcomm/fix/vlm-ttft-params
- fix: Merge pull request #1461 from qualcomm/fix/paul/geniex-bench-top-p-default
- fix: Merge pull request #1462 from mansiverma897993/fix/minicpm5-xml-tool-calls
- fix: Merge pull request #1467 from qualcomm/fix/hf-pull-unsupported-format-message
- fix: Merge pull request #1468 from qualcomm/fix/download-progress-cancel
- fix: Merge pull request #1470 from qualcomm/fix/qnn-log-level
- fix: Merge pull request #1472 from qualcomm/fix/llama-cpp-special-tokens-default
- fix: Merge pull request #1480 from qualcomm/fix/vlm-encoder-fp16-default
- fix: Merge pull request #1484 from qualcomm/fix/adreno643-q4k-constreg
- fix: Merge pull request #1492 from qualcomm/fix/setup-vcvars-clang-cl-discovery
- fix: chore(sdk): bump geniex-qairt to the perf_profile precedence fix
- fix: fix(ci): discover clang-cl.exe recursively for VS2026 Llvm layout
- fix: fix(cli): print download cancel hint before progress
- fix: fix(cli): require both control-token markers for LFM2 tool calls
- fix: fix(go): retry chunk download on any error, not just timeout/EOF
- fix: fix(qairt): drop explanatory comments on the log-level wiring
- fix: fix(qairt): surface QNN backend logs through --log/GENIEX_LOG
- …and 45 more
Architecture
- Containers 0 added · 0 removed · contexts 1 added · 1 removed · edges 0 added · 0 removed
Added bounded contexts (1)
- app
Removed bounded contexts (1)
- geniex-android
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
qualcomm/GenieX was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 29 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 7a0ff49c4e6ce119a069128ef66d6d0ffbe33a95 — the exact code this score is about.
- Scored under rubric-2026.09.17 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-705631bb727e.