Skip to content
CAI
Software that uses CAICheck a score

autodeployai/ai-serving

59.2

Adequate · 20 September 2026

5.5k

lines of production code

Scala

primary language

1

measurement over time

CAI band scale
CAI lens gauges

What this system is

This system is an AI model serving platform that supports deploying and inferring machine learning models in PMML and ONNX formats. It provides HTTP and gRPC endpoints compliant with the v2 inference protocol, featuring automatic request batching to optimize throughput. The service is containerized for both CPU and GPU environments, leveraging ONNX Runtime for acceleration, and includes example notebooks for training and deployment workflows.

Features

Add Docker support for CPU and CUDA-based AI Serving

Users can now build and run AI Serving containers for both CPU and GPU environments. The new Dockerfiles use Ubuntu 22.04 and OpenJDK 17, with the CUDA variant leveraging NVIDIA CUDA 12.3.2 and cuDNN 9. The CPU image builds a standard assembly, while the CUDA image enables the ONNX Runtime backend via the \--add-exports\ and \-Donnxruntime.backend=cuda\ flags to support GPU acceleration.

dockerfiles · high confidence

Added XGBoost Iris model example in PMML format

The examples/models directory now includes a new \xgb-iris.pmml\ file, providing a pre-exported PMML representation of an XGBoost classification model trained on the Iris dataset. This file, generated by Nyoka version 5.5.0, contains the complete model definition including data dictionary, standard scaler transformations, and the tree-based mining model structure, allowing users to load and infer using this specific example without needing to train the model locally.

examples/models · high confidence

New tutorials for PMML and ONNX model serving

Added example notebooks demonstrating how to train, convert, and serve machine learning models using AI-Serving. The \IrisXGBoost.ipynb\ notebook shows how to train an XGBoost classifier and convert it to PMML format. The \AIServingIrisXGBoostPMMLModel.ipynb\ notebook provides a tutorial on deploying and making predictions with that PMML model via HTTP. Additionally, the \AIServingMnistOnnxModel.ipynb\ notebook demonstrates deploying and serving a pre-trained MNIST CNN model in ONNX format. These examples also include the generated \ai\_serving\_pb2.py\ Python file required for interacting with the AI-Serving gRPC/HTTP endpoints.

examples · high confidence

Support for v2 inference APIs and automatic batching

The serving layer now supports the v2 inference protocol for both HTTP and gRPC endpoints, exposing standard health, metadata, and inference routes (e.g., /v2/health/live, /v2/models/{name}/infer) alongside the existing v1 APIs. Additionally, the system introduces automatic batching for supported models (such as ONNX and PMML), which aggregates incoming requests to improve throughput based on configurable batch size and delay parameters.

src/main/scala/com · high confidence

Removals

Removal of legacy model serving components

The model serving infrastructure in the \ai.autodeploy.serving\ package has been removed. This includes the deletion of the \ModelCache\ and \ModelManager\ classes that handled model versioning and deployment, the removal of specific model runtime implementations for ONNX (\OnnxModel\) and PMML (\PmmlModel\), and the elimination of the gRPC \DeploymentServiceImpl\ along with its associated protobuf conversion utilities (\protobuf/package.scala\) and JSON support (\JsonSupport\, \JsonUtils\).

src/main/scala/ai · high confidence

Behavioural changes

AI-Serving API updates: new deployment config, v2 inference protocol, and type changes

The AI-Serving gRPC API now supports automatic batching inference via a new DeployConfig message (allowing configuration of request timeout, max batch size/delay, and warmup settings) in the DeploymentService. The service also introduces the v2 inference protocol through a new grpc\_predict\_v2.proto defining standard server/model metadata and inference RPCs. Additionally, model version identifiers have been changed from integers to strings across ModelInfo, ModelMetadata, and ModelSpec, and the internal tensor representation has been updated to a custom TensorProto message replacing the previous ONNX dependency.

src/main/protobuf · high confidence

Updated AI dispatcher configuration and enhanced logging capabilities

The application's AI dispatcher has been renamed to 'ai-dispatcher' and its thread pool size is now dynamically set to use all available CPU cores (-1) instead of a fixed size of 16. New configuration sections have been added to support ONNX Runtime backend options (including CPU, CUDA, DNNL, TensorRT, and DirectML) and to control request timing logs. Additionally, the logging format has been updated to include method and line number details, the root logger package has been changed from 'ai.autodeploy.serving' to 'com.autodeployai.serving', and explicit logging levels have been configured for Akka and Netty components.

src/main/resources · high confidence

Test coverage

Added test coverage for v2 inference APIs and batch processing; Added test resources for v2 API and PMML model validation; Removed BaseSpec test helper.

Dependencies

Major dependency upgrades and GPU support in build configuration

The build manifest has been updated to version 2.2.0, significantly upgrading core libraries including Akka HTTP (10.1.11 to 10.5.3), Akka Streams (2.6.4 to 2.7.0), PMML4s (0.9.5 to 1.5.8), and ONNX Runtime (1.6.0/1.7.0 history to 1.22.0). The project now supports GPU acceleration for ONNX models via a configurable 'gpu' build property that switches the dependency to 'onnxruntime\_gpu'. Additionally, Logback was upgraded to 1.5.13, and the assembly merge strategy was adjusted to handle Java 9+ module info files, while test execution was configured to fork with specific Java exports.

(dependencies) · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

Baseline

  • First survey — no prior run to compare against. CAI 59.

Lenses

  • Code Health 91
  • Architecture 95
  • Maturity 57
  • Readiness 50
  • Security 66

Changes since last survey

  • 73 commits — 67 feature/other, 6 fixes

By area

  • (root) — 26 commits
  • src/main — 15 commits
  • src/test — 11 commits
  • dockerfiles/scripts — 6 commits
  • .github/workflows — 5 commits
  • (repo) — 3 commits
  • examples/AIServingMnistOnnxModel.ipynb — 3 commits
  • dockerfiles/Dockerfile — 2 commits
  • examples/AIServingIrisXGBoostPMMLModel.ipynb — 2 commits

Notable commits

  • fix: Fix failed UTs for various platforms
  • fix: Fix the build error
  • fix: Fix the memory leak after deploying
  • fix: Fixed warmup failure
  • fix: fix build error
  • fix: fix build error
  • change: Add conf.json
  • change: Add more details of Predict Protocol V2 in README.md
  • change: Add run options to cancel running requests
  • change: Add test cases for v2 APIs
  • change: Build image when a tag pushed only
  • change: Bump pmml4s to 1.5.8
  • change: Correct the link
  • change: Correct the outdated link
  • change: Create cuda-docker-image.yml
  • change: Create docker-image.yml
  • change: Handle errors from onnxruntime
  • change: Merge branch 'master' of https://github.com/autodeployai/ai-serving
  • change: Merge pull request #17 from autodeployai/api-v2
  • change: Merge pull request #19 from autodeployai/api-v2-enhance
  • …and 53 more

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

autodeployai/ai-serving was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 20 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit ab4eb8e569ffd28a56bc54c17afd920287710af6 — the exact code this score is about.
  • Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-b51f968c9b10.