Skip to content
CAI
Software that uses CAICheck a score

huggingface/text-embeddings-inference

56.7

Adequate · 30 September 2026

24.9k

lines of production code

Rust

with Python

2

measurements over time

CAI band scale
CAI trend line
CAI lens gauges

What this system is

This system is a high-performance inference server for text embeddings, classification, and reranking tasks, designed to serve neural network models via HTTP and gRPC interfaces. It supports a wide variety of hardware backends—including CUDA, ROCm, Intel XPU/HPU, and CPU via ONNX Runtime—while providing configurable pooling strategies and optimized attention mechanisms for modern architectures. The service handles tokenization, batching, and prediction, exposing metrics and tracing for observability and supporting diverse model types such as BERT, Mistral, and Jina.

Features

Add ONNX Runtime backend for improved CPU performance

Users can now run models using the ONNX Runtime backend, which provides better CPU performance compared to previous options. This new backend supports float32 precision and automatically detects model requirements for token type IDs, position IDs, and past key values. It also handles padding configuration via tokenizer settings and supports various pooling strategies for embedding and classification models.

backends/ort · high confidence

Add Predict RPC and Score response types to EmbeddingService

The EmbeddingService now exposes a new Predict RPC endpoint that accepts an EmbedRequest and returns a PredictResponse containing a list of Score objects. This change enables clients to retrieve prediction scores alongside or instead of standard embeddings, extending the service's capabilities beyond simple embedding generation.

backends/proto · high confidence

Add reranker support and configurable pooling to Python backend

The Python backend now supports reranker models via a new Predict endpoint, allowing users to obtain similarity scores in addition to standard embeddings. Additionally, the pooling strategy is now configurable (defaulting to 'cls') via a new CLI argument, and the OpenTelemetry service name can be explicitly set. These changes enable more flexible model usage and better observability for the Python-based inference server.

_backends/python/server/text\_embeddings\server · high confidence

Added reranker prediction capability to the gRPC client

The gRPC client in the backends/grpc-client module now supports a new \predict\ method, allowing users to send token-level inputs (input\_ids, token\_type\_ids, position\_ids, cu\_seq\_lengths, max\_length) and receive relevance scores. This addition aligns the client with the Text Embeddings Inference backend, as indicated by updated documentation and a corrected proto file path in the build configuration.

backends/grpc-client · high confidence

Expanded hardware support and configurable observability

The Python backend now supports Intel (XPU via IPEX) and AMD ROCm GPUs in addition to standard CUDA and Habana Gaudi (HPU) devices, with optimized attention kernels for each platform. A new device detection utility ensures correct framework initialization, while ROCm builds can now use pre-downloaded Triton layer-norm kernels for improved startup reliability. Additionally, the OpenTelemetry tracing setup now accepts a configurable service name via the \OTLP\_SERVICE\_NAME\ environment variable, allowing users to customize how traces are identified in their observability backends.

_backends/python/server/text\_embeddings\server/utils · high confidence

Expanded model support and device compatibility in Python backend

The Python backend now supports a wider range of model architectures and hardware devices. New model types include classification models (ClassificationModel), masked language models with Splade pooling (MaskedLanguageModel), and optimized flash implementations for Mistral, Qwen3, and JinaBert models. The backend also extends device support to include Intel XPU, HPU, and AMD ROCm GPUs, replacing the previous CUDA-only constraint for flash attention. Additionally, a configurable pooling strategy (e.g., 'splade') is now available via the \pool\ parameter, and the \TRUST\_REMOTE\_CODE\ environment variable allows loading models with custom code.

_backends/python/server/text\_embeddings\server/models · high confidence

Initial HTTP server implementation for Text Embeddings Inference

This change introduces the core HTTP server module (\router/src/http\) for the Text Embeddings Inference router, establishing the entry point for all HTTP-based interactions. It defines the request and response types (\router/src/http/types.rs\) and the server logic (\router/src/http/server.rs\), including endpoints for model info, health checks, and predictions. The implementation supports flexible input parsing for single strings, string pairs, and batches, and integrates with the inference backend to handle tasks like prediction, embedding, and tokenization while exposing metrics and OpenTelemetry tracing.

router/src/http · high confidence

Initial release of the TEI gRPC API specification

This change introduces the \tei.proto\ file, defining the complete gRPC service contract for the Text Embeddings Inference (TEI) system. It establishes four main services—Info, Embed, Predict, and Rerank—along with a Tokenize service for text processing. The protocol buffers define the request and response structures for embedding generation (including dense, sparse, and all-token embeddings), text classification/prediction, and reranking, as well as metadata and configuration details like truncation directions and model types.

proto · high confidence

Introduce gRPC API for text embeddings and tokenization

The router now exposes a gRPC interface alongside the existing HTTP endpoints, allowing clients to interact with the service using the Protocol Buffers definition. This new server implementation handles requests for embedding (pooled and sparse), tokenization, and model information, providing a binary-protocol alternative for integration.

router/src/grpc · high confidence

Major infrastructure overhaul: multi-arch builds, sccache, and new Dockerfiles

The build and deployment infrastructure has been significantly restructured to support multiple architectures and hardware backends. New Dockerfiles have been introduced for ARM64 (\Dockerfile-arm64\), Intel CPU/XPU/HPU (\Dockerfile-intel\), and AMD ROCm (\Dockerfile-rocm\), alongside a new multi-architecture CUDA image (\Dockerfile-cuda-all\) that bundles binaries for compute capabilities 75, 80, 90, 100, and 120. The primary \Dockerfile\ and \Dockerfile-cuda\ have been updated to use Rust 1.92, CUDA 12.9, and Ubuntu 24.04, and now integrate \sccache\ (v0.10.0) for faster builds. Entrypoint scripts (\cuda-entrypoint.sh\, \cuda-all-entrypoint.sh\) have been added to handle dynamic library paths and select the appropriate binary based on the host's CUDA compute capability. Additionally, Nix flake support (\flake.nix\, \flake.lock\) and pre-commit hooks have been added to the repository root.

(repo-wide) · high confidence

New gRPC load tests and updated HTTP test configuration

Added new k6 load testing scripts for gRPC (\load\_grpc.js\ and \load\_grpc\_stream.js\) to benchmark the TEI service's unary and streaming Embed endpoints, including metrics for total, tokenization, queue, and inference times. Updated the existing HTTP load test (\load.js\) to target the root path (\/\) instead of \/embed\, increased the pre-allocated virtual users from 2000 to 5000, and reduced the request arrival rate from 500 to 50 per second.

_load\tests · high confidence

New model architectures and Flash Attention optimizations

The Candle backend now supports several new model architectures, including DeBERTa-v2, DistilBERT, GTE, Jina (including Jina Code), Mistral, ModernBERT, and NomicBert. To improve performance on CUDA devices, Flash Attention implementations have been added for DistilBERT, GTE, Jina, Jina Code, Mistral, ModernBERT, and NomicBert. Additionally, a generic Dense layer implementation has been introduced to support classification heads that use dense projections with specific activation functions.

backends/candle/src/models · high confidence

New optimized CUDA layer implementations for linear, normalization, and rotary embeddings

The \backends/candle/src/layers\ module now includes new Rust source files that provide optimized CUDA-accelerated implementations for core neural network layers. Specifically, \cublaslt.rs\ introduces fused matrix multiplication with activation support via cuBLASLt, \layer\_norm.rs\ and \rms\_norm.rs\ add fused normalization kernels (including residual addition) using \candle\_layer\_norm\, \linear.rs\ integrates these CUDA paths for linear layers, \index\_select.rs\ adds a faster CUDA index select kernel, and \rotary.rs\ implements rotary embedding logic with support for Llama3 and NTK scaling strategies. These changes improve inference performance on CUDA devices by offloading specific operations to specialized libraries.

backends/candle/src/layers · high confidence

Behavioural changes

Improved CUDA compute capability detection and Flash Attention support for newer GPUs

The Candle backend now detects the GPU's CUDA compute capability at runtime using the CUDA driver API instead of relying on the \nvidia-smi\ command-line tool, which improves reliability in containerized environments. This change also expands support for Flash Attention to include newer architectures (compute capabilities 90, 100, and 120/121) and adds support for ALiBi slopes and attention windowing in Flash Attention v2, enabling better performance and compatibility with models like ModernBert and Jina on Blackwell and Hopper GPUs.

backends/candle/src · high confidence

Introduction of configurable data types and refined HPU warmup logic

The backend now supports a configurable \DType\ enum (Float16, Float32, Bfloat16) with feature-gated availability and specific defaults per backend (e.g., Bfloat16 for Python, Float32 for ORT/Accelerate). Additionally, the HPU warmup process has been adjusted to use dummy inputs with shapes closer to real scenarios, employing exponential growth for sequence lengths and configurable bucket sizes via environment variables like \SEQ\_LEN\_EXPONENT\_BASE\ and \PAD\_SEQUENCE\_TO\_MULTIPLE\_OF\.

backends/src · high confidence

Python backend now supports classification models and configurable pooling strategies

The Python backend has been upgraded to support classification models in addition to embeddings, exposing a new \predict\ method alongside the existing \embed\ method. Users can now configure the pooling strategy (CLS, Mean, LastToken, or Splade) via the \--pool\ argument, and the backend correctly applies CLS pooling for classifier model types. Additionally, an \otlp-service-name\ argument has been added to allow explicit configuration of the OpenTelemetry service name.

backends/python · high confidence

Refactored core inference, tokenization, and queueing for better performance and flexibility

The core inference engine has been restructured to support advanced model features and improved performance. Tokenization now runs in dedicated background threads with configurable workers, supporting default prompts, custom truncation directions, and explicit empty-input validation. The inference pipeline introduces separate tasks for batching and backend communication, enabling batch prefetching and serialization off the blocking thread to reduce latency. The request queue has been moved to a blocking thread, switched from \flume\ to \mpsc\ channels, and now supports padded models and configurable capacity limits. Additionally, model artifact downloading now handles legacy Sentence Transformers configuration files and optional pooling configs to ensure compatibility with various model formats.

core · high confidence

Removal of gRPC metadata extraction capability

The \MetadataExtractor\ struct and its implementation of the \Extractor\ trait have been removed from the \backends/grpc-metadata\ library. While the \MetadataInjector\ remains to handle outgoing context propagation, the ability to extract trace context from incoming gRPC request metadata is no longer available in this component.

backends/grpc-metadata · high confidence

Reworked embedding output types and added pooling strategies

The core library now distinguishes between pooled and raw embeddings via a new \Embedding\ enum (supporting \Pooled\ and \All\ variants) and updated the \Batch\ struct to track \pooled\_indices\ and \raw\_indices\. A new \Pool\ enum introduces configurable pooling strategies including CLS, Mean, SPLADE, and LastToken, allowing users to select how embeddings are derived from model outputs. The \Backend\ trait has been updated to return \Embeddings\ instead of simple vectors and includes a new \predict\ method for classification tasks, while \ModelType\ now explicitly supports classifiers alongside embeddings.

backends/core · high confidence

Router refactored with modular logging, metrics, and new CLI arguments

The router's internal structure has been reorganized into dedicated modules for logging, Prometheus metrics, and graceful shutdown, replacing the previous monolithic server implementation. This change introduces new command-line arguments including \--served-model-name\ for OpenAI compatibility, \--pooling\ to override pooling configuration, \--auto-truncate\ (defaulting to true), and \--default-prompt\/\--default-prompt-name\ for standard sentence-transformers prompt handling. It also adds support for reading the Hugging Face Hub token from the local cache via \--hf-token\, configures OpenTelemetry OTLP endpoints for distributed tracing, and exposes a configurable Prometheus port for monitoring.

router/src · high confidence

Fixes

Fix build script panic on unexpected nvidia-smi output

The build script for the Candle backend now handles unexpected output from \nvidia-smi\ gracefully instead of panicking. Previously, the script assumed the first line of output would always be \compute\_cap\ and used an assertion that would crash the build if this assumption failed. The change replaces the assertion with a conditional check that returns a descriptive error message, improving build stability when the CUDA environment is misconfigured or returns unexpected data.

backends/candle · high confidence

Test coverage

Added integration tests for Gaudi (HPU) platform; Added integration tests for MRL embedding generation; Added integration tests for the Candle backend; Snapshot tests added for BERT, Flash BERT, DeBERTaV2, and Dense embedding models.

Dependencies

Add ONNX Runtime backend and update Python server dependencies

This change introduces a new ONNX Runtime backend (text-embeddings-backend-ort) for CPU inference, adding the ort crate (version 2.0.0-rc.10) and its sys bindings to the Rust workspace. It also updates the Python server dependencies, pinning transformers to 4.51.3, tokenizers to 0.19.1, and upgrading opentelemetry components to 1.15.0, while introducing separate requirements files for AMD, HPU, and Intel hardware targets to manage platform-specific dependencies.

(dependencies) · high confidence

Python server build tooling and dependency updates

The Python backend server's build process has been updated to use newer versions of gRPC tooling (grpcio-tools 1.51.1 to 1.62.2 and mypy-protobuf 3.4.0 to 3.6.0) and relaxed protobuf type constraints. The installation target now uses the --no-deps flag when installing requirements to prevent unintended dependency resolution changes. Additionally, a new kernels.lock file was added to pin the triton-layer-norm kernel variant, and minor whitespace fixes were applied to several Makefiles and the README.

backends/python/server · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

Score

  • CAI 58 → 57 (-1.5)
  • Rubric changed (rubric-2026.09.11 → rubric-2026.09.18) — scores are not directly comparable.

Lenses

  • Code Health 73 → 73 (-0.7)
  • Architecture 100 → 97 (-3.3)
  • Maturity 59 → 44 (-15.4)
  • Readiness 66 → 74 (+8.6)
  • Security 47 → 55 (+8.3)
  • Performance 100 (new)

Resolved (24)

  • Documentation: no installation or build instructions (README.md)
  • Documentation: no installation or build instructions (docs/source/en/index.md)
  • Documentation: no usage examples (README.md)
  • Documentation: no usage examples (docs/source/en/cli_arguments.md)
  • High CVE: [GHSA redacted] (Cargo.lock)
  • High CVE: [GHSA redacted] (Cargo.lock)
  • High CVE: [GHSA redacted] (Cargo.lock)
  • High vulnerability: [GHSA redacted] (Cargo.lock)
  • Medium CVE: [GHSA redacted] (Cargo.lock)
  • Medium CVE: [GHSA redacted] (Cargo.lock)
  • Medium CVE: RUSTSEC-2025-0055 (Cargo.lock)
  • Medium advisory (unmaintained): RUSTSEC-2025-0134 (Cargo.lock)
  • Medium advisory (unmaintained): RUSTSEC-2026-0173 (Cargo.lock)
  • Medium advisory (unsound): RUSTSEC-2026-0097 (Cargo.lock)
  • Medium advisory (unsound): RUSTSEC-2026-0097 (Cargo.lock)
  • Medium advisory (unsound): RUSTSEC-2026-0186 (Cargo.lock)
  • Medium advisory (unsound): RUSTSEC-2026-0190 (Cargo.lock)
  • Medium advisory (unsound): RUSTSEC-2026-0221 (Cargo.lock)
  • Medium vulnerability: [GHSA redacted] (Cargo.lock)
  • Medium vulnerability: RUSTSEC-2026-0204 (Cargo.lock)
  • …and 4 more

New (31)

  • Ambiguous semantic distinction between 'embed' and 'predict' at the Backend interface level. While 'predict' might imply classification/logits and 'embed' implies vectorization, the signatures are identical (accepting the same Batch type and returning Result). Without distinct return types or separate interfaces, it is unclear if these are truly distinct operations or just naming variants for different model heads.
  • Documentation: no architecture or design documentation (docs/source/en/private_models.md)
  • Documentation: no project overview (README.md)
  • Duplicate tokenization logic exposed at two levels. Infer.tokenize delegates to Tokenization.tokenize with nearly identical signatures. This exposes internal implementation details (Tokenization) through the higher-level Infer API, creating redundancy.
  • Duplicated block (10 lines × 2) (backends/python/server/text_embeddings_server/models/default_model.py)
  • Duplicated block (14 lines × 2) (backends/python/server/text_embeddings_server/models/flash_bert.py)
  • Duplicated block (15 lines × 2) (backends/python/server/text_embeddings_server/models/flash_mistral.py)
  • Duplicated block (18–19 lines × 3) (backends/python/server/text_embeddings_server/models/classification_model.py)
  • Duplicated block (19–20 lines × 2) (backends/python/server/text_embeddings_server/models/flash_mistral.py)
  • Duplicated block (20 lines × 2) (backends/python/server/text_embeddings_server/models/flash_mistral.py)
  • Duplicated block (22 lines × 2) (backends/python/server/text_embeddings_server/models/flash_mistral.py)
  • Duplicated block (23 lines × 2) (backends/python/server/text_embeddings_server/models/flash_mistral.py)
  • Duplicated block (24 lines × 2) (backends/python/server/text_embeddings_server/models/flash_mistral.py)
  • Duplicated block (24 lines × 3) (backends/python/server/text_embeddings_server/models/flash_bert.py)
  • Duplicated block (33 lines × 2) (backends/python/server/text_embeddings_server/models/flash_mistral.py)
  • Duplicated block (5 lines × 2) (backends/python/server/text_embeddings_server/models/flash_mistral.py)
  • Duplicated block (5 lines × 2) (backends/python/server/text_embeddings_server/models/flash_mistral.py)
  • Duplicated block (6 lines × 2) (backends/python/server/text_embeddings_server/models/flash_mistral.py)
  • Duplicated block (6 lines × 2) (backends/python/server/text_embeddings_server/models/flash_mistral.py)
  • Duplicated block (8 lines × 2) (backends/python/server/text_embeddings_server/models/flash_mistral.py)
  • …and 11 more

Changes since last survey

  • 4 commits — 4 feature/other, 0 fixes

By area

  • (repo) — 1 commit
  • .github/dependabot.yml — 1 commit
  • .github/workflows — 1 commit
  • backends/python — 1 commit

Notable commits

  • change: Scope GITHUB_TOKEN permissions per job (#934)
  • change: Update version to 1.9.4 & run cargo update (#928)
  • change: chore: enable Dependabot weekly GitHub Actions bumps (#868)
  • change: feat: ROCm flash-attn varlen, triton layer norm, and AMD Dockerfile (#860)

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

huggingface/text-embeddings-inference was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 30 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit 98b7ea2ddb928eccbfde41d96e9576f876d045f4 — the exact code this score is about.
  • Scored under rubric-2026.09.18 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-cb25ca4feafa.