Skip to content
CAI
Software that uses CAICheck a score

SYSTRAN/faster-whisper

58.6

Adequate · 18 September 2026

5.8k

lines of production code

Python

primary language

1

measurement over time

CAI band scale
CAI lens gauges

What this system is

This system is a high-performance audio transcription library that reimplements OpenAI's Whisper model using CTranslate2 to enable faster inference with reduced memory consumption. It provides batched processing capabilities, supports distil-whisper models for efficiency, and includes integrated modules for voice activity detection, tokenization, and audio decoding. The project also offers comprehensive benchmarking tools and Dockerized deployment examples to facilitate performance evaluation and usage.

Features

Add Docker example for inference

A new Dockerfile and infer.py script have been added to the docker directory, providing a ready-to-run example for audio transcription using the faster-whisper library. The container is based on Ubuntu 22.04 with CUDA 12.3.2 and installs the necessary Python dependencies to execute inference on a provided audio file (jfk.flac).

docker · high confidence

Add dedicated benchmarks for WER, speed, and memory usage

The benchmark directory now includes specific scripts to evaluate Word Error Rate (using the \jiwer\ library), inference speed, and memory consumption. The WER benchmark (\wer\_benchmark.py\ and \evaluate\_yt\_commons.py\) utilizes the \BatchedInferencePipeline\ for evaluation on Librispeech and YouTube-Commons datasets, while \speed\_benchmark.py\ and \memory\_benchmark.py\ provide tools to measure performance characteristics of the \WhisperModel\.

benchmark · high confidence

Initial release of faster-whisper with batched inference and distil-whisper support

This entry marks the initial release of the faster-whisper library, a reimplementation of OpenAI's Whisper model using CTranslate2 for faster inference with lower memory usage. The release introduces batched transcription capabilities via the new BatchedInferencePipeline, which significantly improves throughput for processing audio files. It also adds support for running inference with distil-whisper models, specifically the distil-large-v3 checkpoint, allowing for faster transcription with maintained accuracy. The package now requires Python 3.9 or greater and includes comprehensive documentation with benchmark comparisons against other implementations, installation instructions for GPU dependencies (cuBLAS/cuDNN), and development tooling for contributors.

(repo-wide) · high confidence

Behavioural changes

Assets directory converted to a Python package

The faster\whisper/assets directory now includes an \\init\\_.py file, making it a valid Python package. This change enables the assets to be properly distributed and imported as part of the package structure.

_faster\whisper/assets · high confidence

faster\_whisper v1.2.1: New Tokenizer, VAD, and Audio Decoding Modules

This release introduces new, standalone modules for tokenization, voice activity detection (VAD), and audio decoding, alongside a version bump to 1.2.1. The \tokenizer.py\ module adds a dedicated \Tokenizer\ class that wraps \tokenizers.Tokenizer\ to handle language codes, tasks, and special token logic (e.g., \no\_speech\, \no\_timestamps\) explicitly. The \vad.py\ module introduces a \VadOptions\ dataclass and \get\_speech\_timestamps\ function, providing configurable speech segmentation based on Silero VAD probabilities. The \audio.py\ module refactors \decode\_audio\ to use the PyAV library directly, adding support for stereo channel splitting (\split\_stereo\) and improved error handling for invalid audio frames. Additionally, \utils.py\ exposes \available\_models\ and \download\_model\ functions, mapping model aliases (like \turbo\ and \large\) to their Hugging Face repository IDs.

_faster\whisper · high confidence

Test coverage

Initial test suite for transcription, tokenizer, and utility functions

Added a comprehensive test suite covering core library functionality. Tests verify the \available\_models\ and \download\_model\ utilities, tokenizer behavior including suppressed tokens and Unicode splitting, and transcription features such as language detection, batched inference, VAD filtering, stereo channel separation, and multilingual support.

tests · high confidence

Dependencies

Update core dependencies and add benchmark requirements

The project updates several core dependencies: CTranslate2 is upgraded to version 4.x (previously 3.5+), the Hugging Face Hub client is added (\>=0.23), tokenizers is relaxed to allow versions up to 1.x, ONNX Runtime is added (\>=1.14), and PyAV is upgraded to version 11+. Additionally, a new benchmarking environment is introduced via \benchmark/requirements.txt\, specifying tools like \transformers\, \jiwer\, \datasets\, \memory\_profiler\, \py3nvml\, and \pytubefix\ for performance and quality evaluation.

(dependencies) · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

Baseline

  • First survey — no prior run to compare against. CAI 59.

Lenses

  • Code Health 81
  • Architecture 100
  • Maturity 52
  • Readiness 50
  • Security 75

Changes since last survey

  • 263 commits — 237 feature/other, 26 fixes

By area

  • faster_whisper/transcribe.py — 106 commits
  • (root) — 94 commits
  • faster_whisper/version.py — 13 commits
  • faster_whisper/audio.py — 9 commits
  • faster_whisper/tokenizer.py — 9 commits
  • faster_whisper/utils.py — 9 commits
  • faster_whisper/vad.py — 8 commits
  • faster_whisper/assets — 5 commits
  • docker/Dockerfile — 3 commits
  • faster_whisper/init.py — 2 commits
  • .github/workflows — 1 commit
  • benchmark/benchmark.m4a — 1 commit
  • benchmark/evaluate_yt_commons.py — 1 commit
  • faster_whisper/feature_extractor.py — 1 commit
  • tests/data — 1 commit

Notable commits

  • fix: Bugfix: Illogical "Avoid computing higher temperatures on no_speech" (#652)
  • fix: Bugfix: code breaks if audio is empty (#768)
  • fix: Fix #839 incorrect clip_timestamps being used in model (#842)
  • fix: Fix VAD index error when a predicted timestamps is too large (#107)
  • fix: Fix list index out of range in word timestamps (#1157)
  • fix: Fix all_tokens handling
  • fix: Fix broken prompt_reset_on_temperature (#604)
  • fix: Fix error in decode_audio for long audio inputs
  • fix: Fix incorrect attribute access
  • fix: Fix language detection with non-speech audio (#895)
  • fix: Fix neg_threshold (#1191)
  • fix: Fix occasional IndexError on empty segments (#227)
  • fix: Fix typing for device_index argument
  • fix: Fix typing of local_files_only
  • fix: Fix typing of words attribute
  • fix: Fix unset attribute when using English-only models
  • fix: Fix variable name reference (#77)
  • fix: Fix window end heuristic for hallucination_silence_threshold (#706)
  • fix: Fix: add <|nocaptions|> to suppressed tokens (#1338)
  • fix: Remove the usage of transformers.pipeline from BatchedInferencePipeline and fix word timestamps for batched inference (#921)
  • …and 243 more

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

SYSTRAN/faster-whisper was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 18 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit ed9a06cd89a93e47838f564998a6c09b655d7f43 — the exact code this score is about.
  • Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-5d04157a340d.