Skip to content
CAI
Software that uses CAICheck a score

huggingface/datasets

61.4

Adequate · 19 September 2026

85.9k

lines of production code

Python

primary language

1

measurement over time

CAI band scale
CAI lens gauges

What this system is

This system is the Hugging Face Datasets library, a Python package designed for loading, manipulating, and saving structured data at scale. It provides a unified API for accessing diverse data formats—including tabular, audio, image, video, and specialized scientific types like biological sequences and 3D meshes—via local files, remote storage, and streaming. The library supports modular I/O operations, parallel processing, and integration with major data frameworks like PySpark and Polars, while offering built-in benchmarking and validation tools for performance tracking.

How it got here

2020 — Legacy cleanup and library initialization

17 changes.

This period focused on initializing the 'datasets' library by removing the legacy 'nlp' package, its associated CLI tools, and outdated dataset builders. The work established a new project structure with modern tooling, introduced core data loading APIs, and implemented comprehensive test coverage and benchmarking infrastructure.

2021 — modular I/O and formatting architecture

14 changes.

This period focused on restructuring the datasets library into a modular I/O and formatting system, introducing dedicated readers and writers for various formats alongside support for JAX and Polars. Significant enhancements included expanding packaged module support for diverse data types, adding streaming capabilities for compressed files, and implementing new feature types for biological and 3D data.

2022–2024 — Packaged module expansion and infrastructure modernization

14 changes.

This period focused on significantly expanding the library's built-in dataset builders, introducing support for diverse formats such as images, audio, PDB structures, Spark DataFrames, and WebDatasets. Concurrently, the project modernized its underlying infrastructure by refactoring the download subsystem for better streaming support and migrating configuration to pyproject.toml with the ruff linter.

2025–2026 — specialized file format support

9 changes.

This period focused on expanding the datasets library with dedicated builders for a wide range of specialized file formats, including scientific, biological, and time-series data. New modules were added to support HDF5, Lance, CoNLL, Apache Iceberg, TsFile, FASTA, FASTQ, GenBank, and mmCIF formats. These additions enable direct loading and efficient processing of complex structured data without requiring external conversion steps.

Features

Add Apache Iceberg dataset format support

Users can now load datasets directly from Apache Iceberg tables using the new \Iceberg\ builder. This feature allows reading from a specified catalog and table, with support for filtering rows, selecting specific columns, and performing time-travel queries via snapshot IDs. The implementation leverages \pyiceberg\ for efficient scanning and integrates with the existing datasets library's parallel processing capabilities.

_src/datasets/packaged\modules/iceberg · high confidence

Add AudioFolder packaged module for loading audio from directories

A new packaged module, AudioFolder, is introduced to allow users to load audio datasets directly from folder structures. This module extends the existing FolderBasedBuilder and supports a wide range of audio and video container formats (including .wav, .mp3, .flac, .ogg, .mp4, and .mkv) by leveraging TorchCodec for decoding, replacing the previous reliance on Soundfile.

_src/datasets/packaged\modules/audiofolder · high confidence

Add CoNLL dataset format loader

Users can now load CoNLL-style datasets (including CoNLL-2000, CoNLL-2003, and CoNLL-U formats) directly using the \conll\ builder. This new module parses whitespace-separated token files, handling sentence boundaries, optional comment prefixes, and document start markers, while allowing custom column configurations for NER, chunking, or dependency parsing tasks.

_src/datasets/packaged\modules/conll · high confidence

Add GenBank file format support for biological sequence data

Users can now load GenBank files, a standard text-based format for nucleotide and protein sequences with annotations, using the new GenBank dataset builder. This feature introduces a pure Python state machine parser that handles sequence data, header fields, and complex multi-line feature annotations (including location parsing and qualifier handling) without requiring external dependencies. The implementation supports configurable batching to manage memory usage for large sequences and exposes structured columns such as locus name, accession, organism, sequence, and detailed feature annotations.

_src/datasets/packaged\modules/genbank · high confidence

Add HDF5 dataset support

Users can now load HDF5 files (\.h5\, \.hdf5\) directly into the datasets library. This new builder automatically infers schema from the HDF5 structure, supporting complex data types (converted to real/imaginary float components), compound dtypes, and variable-length arrays, while converting the data into Arrow tables for efficient processing.

_src/datasets/packaged\modules/hdf5 · high confidence

Add Spark dataset builder for PySpark DataFrames

Introduces a new \Spark\ dataset builder and \SparkConfig\ in the \packaged\_modules\ directory, enabling users to create Hugging Face Datasets directly from PySpark DataFrames via \Dataset.from\_spark\. The implementation includes a \SparkExamplesIterable\ to efficiently stream data from Spark partitions, supports sharding and shuffling of data sources, and handles cache directory validation for multi-node clusters.

_src/datasets/packaged\modules/spark · high confidence

Add TsFile (Apache IoTDB) packaged builder with per-device wide format

Users can now load Apache IoTDB TsFile datasets directly using the \tsfile\ builder. This new capability reads time-series data in a per-device wide format, where each output row represents a single device identified by its TAG values, with the \time\ column and all FIELD columns stored as Arrow list columns. The builder supports filtering by table name, specific field columns, and time ranges, and automatically merges data across multiple files for the same device while sorting by time.

_src/datasets/packaged\modules/tsfile · high confidence

Add XML dataset builder

Users can now load datasets from XML files using the new \xml\ builder. This feature allows specifying the file encoding (defaulting to UTF-8) and error handling strategies. The builder reads the entire content of each XML file as a single string value, supporting both automatic schema inference and explicit feature definitions.

_src/datasets/packaged\modules/xml · high confidence

Add support for loading FASTA biological sequence files

Users can now load FASTA files (e.g., .fa, .fasta, .fna) directly into datasets. This new module parses nucleotide and peptide sequences into columns for ID, description, and sequence data, with an optional raw record column. It supports configuration for batch sizes and byte limits to handle large genomic files efficiently, and integrates with Biopython for advanced sequence handling when available.

_src/datasets/packaged\modules/fasta · high confidence

Add support for loading FASTQ sequencing data files

Users can now load FASTQ files (with .fq or .fastq extensions) directly into datasets. This new packaged module provides a lightweight, pure-Python parser that extracts sequence identifiers, descriptions, nucleotide sequences, and quality scores. The loader supports configurable batching to manage memory usage and allows users to select specific columns, including an optional raw record column for advanced use cases.

_src/datasets/packaged\modules/fastq · high confidence

Add support for loading mmCIF protein structure files

Users can now load macromolecular structures stored in mmCIF format (.cif, .mmcif) using the new MmcifFolder builder. This feature follows the existing folder-based pattern, allowing datasets to be constructed from directories where folder names serve as labels or where a metadata CSV provides additional columns. Each row in the resulting dataset contains one complete structure file, decoded into a Bio.PDB.Structure.Structure object via Biopython.

_src/datasets/packaged\modules/mmcif · high confidence

Add support for reading Lance datasets

Users can now load datasets stored in the Lance format. This new module handles both multi-file Lance datasets (resolving URIs and handling storage options) and single-file Lance formats, including automatic inference of binary blob types (Audio, Image, Video) via magic bytes and support for Hugging Face Hub authentication.

_src/datasets/packaged\modules/lance · high confidence

Added default DVC plot configurations for benchmarking

The .dvc directory now includes default Vega-Lite plot templates (confusion, default, scatter, and smooth) and a .gitignore to exclude local config and cache files. This enables users to visualize experiment metrics and benchmarks directly within DVC using standardized chart types like line plots, scatter plots, and confusion matrices without manually defining visualization schemas.

.dvc · high confidence

Expanded packaged dataset module support and extension mapping

The \src/datasets/packaged\modules/\\init\\_.py\ file has been significantly expanded to include loaders for a wide variety of new data formats, enabling users to load datasets directly from files in these formats without custom scripts. New supported formats include Lance, CoNLL/CoNLL-U, Inspect AI eval logs, Parquet, NDJSON, GenBank, NIfTI, PDB, mmCIF, FASTA, FASTQ, Apache Iceberg, Vortex, TsFile, 3D Mesh, PDF, SQL databases, HDF5, Webdataset, and XML. The module also registers these new loaders in the \\_PACKAGED\_DATASETS\_MODULES\ dictionary for caching and updates the \\_EXTENSION\_TO\_MODULE\ mapping to automatically infer the correct loader based on file extensions (e.g., \.lance\, \.conll\, \.parquet\, \.fa\, \.pdb\, \.cif\, \.tsfile\, \.vortex\, \.mesh\, \.pdf\, \.hdf5\, \.xml\, \.sql\, etc.). Additionally, it handles case-insensitive extension matching for folder-based loaders (Image, Audio, Video, Mesh, PDF, NIfTI) and updates metadata file/extension mappings for Lance and other folder types.

_src/datasets/packaged\modules · high confidence

Initial release of the HuggingFace Datasets library (v5.0.2.dev0)

This change introduces the \datasets\ library as a new package, establishing the core API for loading, manipulating, and saving datasets. The \src/datasets/\_\init\\_.py\ file exposes key classes such as \Dataset\, \IterableDataset\, \DatasetDict\, and \IterableDatasetDict\, along with loading functions like \load\_dataset\ and \load\_from\_disk\. It also includes utilities for dataset inspection, splitting, and fingerprinting, providing the foundational interface for users to interact with structured data.

src/datasets · high confidence

Introduce Cache dataset builder for reloading cached datasets

A new \Cache\ dataset builder has been added to \src/datasets/packaged\_modules/cache\, enabling users to reload previously cached datasets directly. This builder supports automatic cache discovery via \hash='auto'\ and \version='auto'\, allowing the system to locate the most recent cached version of a dataset by scanning the cache directory structure. It handles config-specific lookups, raises clear errors when multiple configurations exist without a specified config name, and facilitates streaming from cached Arrow files by reconstructing split generators and generating tables from the stored data.

_src/datasets/packaged\modules/cache · high confidence

Introduce Generator packaged module for creating datasets from callables

A new Generator packaged module has been added to the library, enabling users to construct datasets directly from Python generator functions. This module includes a GeneratorConfig dataclass that validates the input, ensuring the provided generator is a callable and handling optional keyword arguments, and a Generator builder class that manages splitting and sharding of the generated examples. This provides a standardized way to ingest data produced by custom generator logic without writing a full builder implementation.

_src/datasets/packaged\modules/generator · high confidence

Introduce ImageFolder packaged dataset builder

Adds a new ImageFolder builder that loads image datasets from local directories, supporting a comprehensive list of image extensions (such as .jpg, .png, .gif, .tiff) while explicitly excluding formats like .h5 that may contain multiple images. The builder includes configuration options to drop labels or metadata and relies on the underlying folder-based builder infrastructure.

_src/datasets/packaged\modules/imagefolder · high confidence

Introduce Parquet dataset builder with streaming and filtering support

Users can now load Parquet files directly using the \parquet\ builder, which supports streaming mode, column selection, and predicate pushdown filters for efficient data access. The implementation includes configuration options to handle bad files (error, warn, or skip) and allows customizing fragment scan options for caching and prefetching behavior.

_src/datasets/packaged\modules/parquet · high confidence

Introduce WebDataset packaged module with multi-modal support

Users can now load datasets stored in the WebDataset format (TAR archives) using the new \WebDataset\ builder. This module automatically infers schema from the first few examples and supports decoding images, audio, video, and 3D meshes based on file extensions. It also handles compressed files and ensures special keys like \\_\key\\_\ are positioned last in the output examples.

_src/datasets/packaged\modules/webdataset · high confidence

Introduces metadata resource files for dataset card validation

Adds a new \src/datasets/utils/resources\ package containing JSON and YAML configuration files that define the structure and allowed values for dataset metadata. This includes \languages.json\ for ISO language codes, \creators.json\ for data creation types, \size\_categories.json\ for dataset size buckets, \multilingualities.json\ for language distribution types, and \readme\_structure.yaml\ which defines the required subsections and validation rules for dataset README cards.

src/datasets/utils/resources · high confidence

Introduces parallel processing backend and refactors utility modules

This change introduces a new \parallel\ module that allows users to configure a parallel backend (such as joblib or Spark) for dataset operations via the \parallel\_backend\ context manager and \parallel\_map\ function. It also reorganizes the \src/datasets/utils\ package by introducing dedicated modules for file locking (\\_filelock\), extraction (\extract\), and deprecation handling (\deprecation\_utils\), while updating the pickling logic in \\_dill\ to support additional types like \torch.Generator\ and \tiktoken.Encoding\ and ensuring consistent hashing for Arrow tables.

src/datasets/utils · high confidence

New JSON packaged module with agent trace support and streaming

The JSON loader has been rewritten as a packaged module (\src/datasets/packaged\_modules/json\) to support streaming, parallelization, and on-the-fly extraction of compressed archives. It now includes native support for parsing agent traces (including Hermes format) when the \teich\ library is available, and introduces new configuration options for encoding, error handling, and field selection. The implementation leverages pandas with PyArrow backend for improved performance and handles complex feature types, including nested structures and ClassLabel encoding.

_src/datasets/packaged\modules/json · high confidence

New PDB protein structure loader and folder-based builder base class

This change introduces a new \PdbFolder\ packaged module for loading Protein Data Bank (PDB) files, which leverages a newly extracted \FolderBasedBuilder\ base class. The \PdbFolder\ loader supports \.pdb\ and \.ent\ files, automatically inferring labels from directory structure or using optional metadata files (CSV, JSONL, Parquet), and stores decoded structures in a \BioStructure\ column. The underlying \FolderBasedBuilder\ provides the shared logic for these folder-based datasets, including on-the-fly extraction, metadata handling, and label inference.

_src/datasets/packaged\_modules/folder\_based\builder · high confidence

New benchmark suite for dataset performance

Added a comprehensive set of benchmark scripts to the \benchmarks\ directory to measure and track dataset performance. The suite includes \benchmark\_array\_xd.py\ for Array2D read/write speeds, \benchmark\_getitem\_100B.py\ for large-scale indexing, \benchmark\_indices\_mapping.py\ for operations like select and shuffle, \benchmark\_iterating.py\ for iteration and formatting overhead, and \benchmark\_map\_filter.py\ for map/filter operations with various output formats. Utility functions in \utils.py\ support data generation and timing, while \format.py\ provides a tool to convert JSON benchmark results into Markdown reports.

benchmarks · high confidence

New decodable feature types for bio, 3D, medical, and document data

The \src/datasets/features\ module now exports several new feature classes that allow datasets to natively store and decode complex data types: \BioSequence\ and \BioStructure\ for biological sequence and macromolecular structure files (using Biopython), \Mesh\ for 3D mesh files like GLB and PLY (using trimesh), \Nifti\ for neuroimaging NIfTI files (using nibabel), and \Pdf\ for PDF documents. These features follow the same pattern as existing \Audio\ and \Image\ features, accepting file paths or bytes and decoding them into their respective library objects (e.g., \SeqRecord\, \Structure\, \Trimesh\, \Nifti1Image\, or PDF text/structure) when accessed, while supporting optional decoding disablement to return raw path/byte dictionaries.

src/datasets/features · high confidence

New modular formatting system with JAX and Polars support

The dataset formatting logic has been restructured into a modular system within the \src/datasets/formatting\ package, introducing a central registry in \\_\init\\_.py\ that dynamically loads formatters based on available libraries. This change adds native support for JAX arrays via \JaxFormatter\ (including device selection) and Polars DataFrames via \PolarsFormatter\, alongside existing formatters for NumPy, PyTorch, TensorFlow, Pandas, and Python. The new architecture also includes a \BaseArrowExtractor\ and \TensorFormatter\ base classes to standardize data extraction and type conversion, ensuring consistent behavior across different output formats.

src/datasets/formatting · high confidence

New release utility script for managing version updates

A new \utils/release.py\ script has been added to automate the release process. This tool allows maintainers to update version strings in \src/datasets/\_\init\\_.py\ and \setup.py\ for pre-release and post-release workflows, supporting both minor version bumps and patch releases via command-line arguments.

utils · high confidence

Repository initialization and project restructuring

The repository has been initialized with a new structure, introducing standard project files such as a Code of Conduct, Security policy, and Zenodo metadata. The legacy 'nlp-cli' tool has been removed, and the project now uses 'ruff' for code quality checks via a new Makefile and pre-commit configuration. Documentation has been updated to reflect the 'datasets' library name, and the README now highlights key features like streaming mode and multi-modal support. The setup.py has been updated to define current dependencies, including pyarrow\>=24.0.0 and huggingface-hub\>=1.31.0, and introduces optional extras for audio and vision support.

(repo-wide) · high confidence

Support for loading and streaming Arrow datasets

Users can now load and stream datasets stored in the Apache Arrow IPC format (both stream and file binary formats) using the new Arrow packaged module. The builder automatically infers dataset features from the Arrow schema if not explicitly provided, validates record batches during reading, and handles both stream and file-based Arrow files transparently.

_src/datasets/packaged\modules/arrow · high confidence

Support for streaming compressed files via new filesystem protocols

The \src/datasets/filesystems\ module now includes built-in support for reading compressed files (BZ2, GZIP, LZ4, XZ, and Zstandard) as if they were remote filesystems. This allows users to stream data directly from compressed files (e.g., \gzip://file.txt::http://example.com/data.gz\) without downloading and decompressing them locally first, leveraging fsspec's registry to handle these protocols transparently.

src/datasets/filesystems · high confidence

Text dataset builder now supports loading by paragraph or document

The text dataset builder has been extended to allow users to load data by line (the default), paragraph, or entire document. This is controlled via the new \sample\_by\ configuration parameter, which splits input text into rows based on double newlines for paragraphs or reads the whole file as a single row for documents. Additionally, a new \keep\_linebreaks\ boolean parameter allows users to preserve newline characters in the output, and the builder now supports complex feature types for data casting.

_src/datasets/packaged\modules/text · high confidence

Removals

Removal of legacy NLP CLI commands

The \src/nlp/commands\ directory has been completely removed, eliminating the legacy command-line interface tools. This deletes the \convert\ command (used to transform TensorFlow Datasets to HuggingFace NLP format), the \download\ command, the \env\ (environment info) command, and the \user\ command suite (including login, whoami, logout, S3 management, and file upload). Users relying on these specific CLI subcommands will no longer have access to them through this module.

src/nlp/commands · high confidence

Removal of legacy NLP dataset builders

The \datasets\ module has removed a large set of legacy NLP dataset builders that were previously implemented using the deprecated \nlp\ and \tensorflow\_datasets\ APIs. Specifically, the files for GLUE, AESLC, Amazon US Reviews, BigPatent, BillSum, BLiMP, C4, CFQ, and Civil Comments have been deleted. This cleanup removes outdated data loading implementations, likely as part of a migration to a new, unified dataset loading system.

datasets · high confidence

Removal of legacy NLP download module

The \src/nlp/download\ package, which previously provided the \DownloadManager\, \DownloadConfig\, and associated utilities for downloading and extracting dataset files, has been completely removed. This change eliminates the legacy download and extraction logic, including checksum verification, archive handling, and Kaggle integration, from the library.

src/nlp/download · high confidence

Removal of legacy NLP utility modules

The \src/nlp/utils\ package has been removed, deleting the \\_\init\\_.py\, \file\_utils.py\, \py\_utils.py\, \tqdm\_utils.py\, and \version.py\ files. This eliminates the legacy dataset caching, file downloading, Python helper functions, progress bar wrappers, and versioning classes that were previously exported from this location.

src/nlp/utils · high confidence

Removal of the legacy \`nlp\` package source code

The source files for the \nlp\ package (version 0.0.1) have been deleted from the repository. This includes the core module entry point (\\_\init\\_.py\), the \Dataset\ class implementation (\arrow\_dataset.py\), the \DatasetBuilder\ base class (\builder.py\), and the feature connectors (\features/\). Users relying on this specific legacy package version will no longer have access to these components in this location.

src/nlp · high confidence

Behavioural changes

CSV builder now supports pandas 2.0+ and 2.2+ parameters

The CSV dataset builder has been updated to handle version-specific parameters in pandas' \read\_csv\. It now supports the \date\_format\ parameter introduced in pandas 2.0 and correctly handles the deprecation of the \verbose\ parameter in pandas 2.2. The builder dynamically filters these arguments based on the installed pandas version to ensure compatibility and prevent errors when loading CSV files.

_src/datasets/packaged\modules/csv · high confidence

Deprecation of the Pandas builder

The Pandas builder is now deprecated and will be removed in the next major version of the datasets library. Users relying on this builder will see a FutureWarning when loading data, indicating that they should migrate to alternative data loading methods.

_src/datasets/packaged\modules/pandas · high confidence

Introduce structured datasets-cli with env, test, and delete commands

The datasets CLI is now structured around a base command class and three specific subcommands: \env\ (prints system and library versions for debugging), \test\ (validates dataset loading with options for caching, verification, and info saving), and \delete\_from\_hub\ (removes a dataset config from the Hub). This refactoring replaces the previous monolithic CLI entry point with a modular architecture, improving maintainability and allowing users to run targeted dataset validation or environment checks directly from the command line.

src/datasets/commands · high confidence

New modular I/O subsystem for datasets

The \src/datasets/io\ package introduces a new modular architecture for reading and writing datasets, replacing the previous monolithic implementation. This change adds dedicated reader and writer classes for CSV, JSON, Parquet, SQL, Spark, and text formats, as well as a generator stream. Users benefit from a consistent interface across all I/O operations, including support for multi-process parallelism in CSV, JSON, and SQL writers, fsspec integration for remote storage access, and configurable batch sizes and compression options. The new structure also standardizes streaming and non-streaming dataset creation across all supported formats.

src/datasets/io · high confidence

Refactored download module with new configuration and streaming support

The download subsystem has been reorganized into a dedicated \src/datasets/download\ package, introducing a structured \DownloadConfig\ dataclass that centralizes settings such as cache directories, parallel processing limits (\num\_proc\), authentication tokens, and storage options. This change separates the logic into \DownloadManager\ for standard cached downloads and \StreamingDownloadManager\ for lazy, on-the-fly data access, providing a clearer API for users managing data retrieval and streaming workflows.

src/datasets/download · high confidence

Remove legacy notebooks and add documentation index

The notebooks directory has been cleaned up by removing the outdated 'HuggingFace Datasets library demo' notebook (and its checkpoint) and the 'playground.py' script, which relied on the deprecated \nlp\ library. A new README.md has been added to serve as an index for official and community notebooks, including a link to the Quickstart guide.

notebooks · high confidence

Test coverage

Added comprehensive test coverage for new and existing feature types; Added test coverage for CSV, JSON, Parquet, SQL, and Text I/O readers and writers; Added test fixtures for Protein Data Bank (PDB) and CIF formats; Added tests for packaged module loaders; Added tests for the CLI test command and dataset info YAML generation; Initial test suite structure and fixtures; New test fixtures for file formats, fsspec, and Hub integration; Removal of download subsystem test suite; Removed obsolete test suite files.

Dependencies

Migrate configuration to pyproject.toml and configure ruff linter

Project configuration has moved from setup.cfg to pyproject.toml, introducing the ruff linter with specific rule selections (C, E, F, I, W) and ignores for line length, undefined names in type annotations, and complex functions. The file also configures import sorting via isort and sets up pytest to treat FutureWarnings from huggingface\_hub as errors during CI runs.

(dependencies) · high confidence

Housekeeping

Published benchmark results for dataset operations

Added JSON files in the benchmarks/results directory containing performance metrics for various dataset operations, including array I/O, indexing, iteration, mapping, filtering, and indices mapping.

benchmarks/results · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

Baseline

  • First survey — no prior run to compare against. CAI 61.

Lenses

  • Code Health 79
  • Architecture 99
  • Maturity 57
  • Readiness 74
  • Security 54

Changes since last survey

  • 300 commits — 202 feature/other, 98 fixes

By area

  • src/datasets — 199 commits
  • (root) — 56 commits
  • docs/source — 22 commits
  • .github/workflows — 7 commits
  • .github/conda — 2 commits
  • tests/features — 2 commits
  • tests/test_hub.py — 2 commits
  • .github/dependabot.yml — 1 commit
  • tests/fixtures — 1 commit
  • tests/io — 1 commit
  • tests/packaged_modules — 1 commit
  • tests/test_arrow_dataset.py — 1 commit
  • tests/test_filesystem.py — 1 commit
  • tests/test_inspect.py — 1 commit
  • tests/test_load.py — 1 commit
  • tests/test_tqdm.py — 1 commit
  • tests/utils.py — 1 commit

Notable commits

  • fix: Bug Fix: Resuming Twice Resets the Dataloader (#8295)
  • fix: Bump fsspec upper bound to 2026.2.0 (fixes #7994) (#7995)
  • fix: Datasets streaming iterable packaged fixes (#8068)
  • fix: Fix 7846: add_column and add_item erroneously(?) require new_fingerprint parameter (#7884)
  • fix: Fix CI: commit operation equality (hfh 1.20.0) and pytest parametrize collection error (#8283)
  • fix: Fix CSV loader dropping on_bad_lines/encoding_errors on pandas 2.0-2.2 (#8358)
  • fix: Fix DatasetDict.push_to_hub leaving removed splits in the dataset card (#8367)
  • fix: Fix JSON loader schema inference for files starting with a UTF-8 BOM (#8241) (#8243)
  • fix: Fix Parquet streaming hangs at the end of script (#8176)
  • fix: Fix Pdf.decode_example raising TypeError on remote non-Hub paths (#8566)
  • fix: Fix deepcopy bottleneck in state_dict() (#8387)
  • fix: Fix archive path check allowing sibling-directory escapes (#8490)
  • fix: Fix async map checkpoint accounting across repeated resumes (#8575)
  • fix: Fix audio bytes (#7670)
  • fix: Fix base_path in integration tests (#8155)
  • fix: Fix batch(by_column=...) crashing after shard/shuffle/split (#8259)
  • fix: Fix batch_size default description in to_polars docstrings (#7824)
  • fix: Fix batching for table-formatted datasets (#8126)
  • fix: Fix bucket dataset card handling and push metadata accounting (#8354)
  • fix: Fix casting a nullable LargeList to a different inner type (#8346)
  • …and 280 more

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

huggingface/datasets was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 19 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit 3e2c1a6c33b883fd9710f55f7b6fa4ab64d0ee07 — the exact code this score is about.
  • Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-13a154b7f5d1.