Skip to content
CAI
Software that uses CAICheck a score

apache/carbondata

52.2

Adequate · 27 September 2026

242.1k

lines of production code

Java

with Scala

3

measurements over time

CAI band scale
CAI trend line
CAI lens gauges

What this system is

This release delivers a comprehensive overhaul of the core datastore and query execution engine, introducing a new V3 data format with optimized columnar storage, adaptive encoding, and vectorized scanning. It significantly expands the query engine with advanced filter resolution, lazy loading, and distributed index pruning capabilities. Additionally, the update introduces materialized view support, geospatial indexing, and enhanced integration with Flink, Hive, and Hadoop ecosystems.

Features

Add C++ SDK wrapper for CarbonData operations

Introduced a new C++ SDK (CSDK) in the sdk/sdk directory, providing C++ wrappers for core CarbonData operations. The implementation includes CarbonReader for data retrieval, CarbonWriter for data insertion, CarbonRow for row manipulation, CarbonSchemaReader for schema introspection, and Configuration/CarbonProperties for managing settings. These components bridge C++ applications to the Java-based CarbonData engine via JNI, enabling high-performance data processing in C++ environments.

sdk/sdk · high confidence

Add CarbonHiveMetastoreListener to sync schema between Spark and Hive

A new listener class, CarbonHiveMetastoreListener, has been added to the Hive integration module. This component implements the Hive Metastore pre-event listener interface to automatically synchronize table schemas and storage configurations when Spark creates or alters tables using the CarbonData provider. This ensures that Hive metadata accurately reflects the schema and storage handler settings defined by Spark, resolving issues where schema changes in Spark were not reflected back in Hive.

integration/hive/src/main/scala · high confidence

Add CarbonRowReadSupport for date/timestamp conversion

A new read support implementation, CarbonRowReadSupport, has been added to handle the conversion of date and timestamp data types during row reading. This change introduces logic to transform integer-based date values into java.sql.Date objects and convert timestamp values to java.sql.Timestamp, ensuring correct data type handling for Hadoop-based reads.

hadoop/src/main/java/org/apache/carbondata/hadoop/readsupport/impl · high confidence

Add EXPLAIN command output for query pruning and rewrite information

Users can now see detailed information about query rewriting and block pruning in the output of the EXPLAIN command. This new functionality, enabled by the \enable.query.statistics\ configuration flag, displays which indexes (such as CG and FG indexes) contributed to pruning data blocks and blocklets, showing the number of skipped blocks and blocklets for each stage of the query execution plan.

core/src/main/java/org/apache/carbondata/core/profiler · high confidence

Add EscapeSequences enum for common escape characters

A new enum, EscapeSequences, has been added to the core module. It defines constants for common escape sequences including newline, backspace, tab, and carriage return, providing a centralized way to access their string and character representations.

core/src/main/java/org/apache/carbondata/core/enums · high confidence

A new Flink example has been added to demonstrate how to read CarbonData files using Apache Flink. The example shows how to write a CarbonData file using Spark and then read specific columns from that file using Flink's Hadoop file reader.

examples/flink · high confidence

The Flink integration now supports writing streaming data to Carbon via two new writer implementations: a local file-based writer (CarbonLocalWriter) and an S3-based writer (CarbonS3Writer). These are backed by new factory classes (CarbonLocalWriterFactory, CarbonS3WriterFactory) and their respective builders, which are registered via the Java ServiceLoader mechanism (META-INF/services). This enables users to ingest data into Carbon tables using either local storage or Amazon S3, with configurable commit thresholds and temporary write paths.

integration/flink/src/main · high confidence

The integration/flink-proxy module introduces a set of new classes (ProxyFileSystem, ProxyRecoverableWriter, ProxyFileWriter, and related factories and serializers) that act as a bridge between Flink's StreamingFileSink and the CarbonData SDK. This enables users to write Flink streaming data directly to Carbon tables, supporting partitioned table ingestion as indicated by the commit messages.

integration/flink-proxy · high confidence

Add QueryExecutionException for query execution errors

A new exception class, QueryExecutionException, has been added to the core module to handle query execution errors. This change introduces a dedicated exception type for query execution issues, improving error handling and clarity in the CarbonData core library.

core/src/main/java/org/apache/carbondata/core/scan/executor/exception · high confidence

Add direct compression codec for column pages

A new \DirectCompressCodec\ class has been introduced to handle the direct compression and decompression of column pages. This codec supports both standard and varchar-specific encoding types, and provides optimized decoding paths that can fill vectors directly from compressed data, improving performance for direct scan queries.

core/src/main/java/org/apache/carbondata/core/datastore/page/encoding/compress · high confidence

Add time-series pre-aggregation support

Introduces new core components to support time-series data in pre-aggregated tables. This includes a DaysOfWeekEnum for configuring the first day of the week, a TimeSeriesFunctionEnum defining supported granularities (second, minute, hour, day, week, month, year, and 5/10/15/30-minute intervals), and a TimeSeriesUDF that applies these time-based transformations to timestamp data.

core/src/main/java/org/apache/carbondata/core/preagg · high confidence

Added Boolean data type conversion utilities

A new BooleanConvert class was added to the core datastore page encoding module, providing utility methods to convert between boolean, byte, and string representations. This enables support for the Boolean data type within the system.

core/src/main/java/org/apache/carbondata/core/datastore/page/encoding/bool · high confidence

Added ColumnUniqueIdGenerator for generating unique column identifiers

Users can now generate unique identifiers for columns via the new ColumnUniqueIdGenerator class, which implements the ColumnUniqueIdService interface to produce random UUIDs for each column schema.

core/src/main/java/org/apache/carbondata/core/service/impl · high confidence

Added DictionaryByteArrayWrapper for dictionary cache

A new class, DictionaryByteArrayWrapper, was added to the core cache dictionary package. This class wraps a byte array and provides custom equals and hashCode implementations using XXHash32, enabling efficient object comparison and hashing for dictionary data in the cache.

core/src/main/java/org/apache/carbondata/core/cache/dictionary · high confidence

Added DictionaryThresholdReachedException for local dictionary overflow

A new exception class, DictionaryThresholdReachedException, has been introduced in the core module to handle cases where the local dictionary reaches its capacity threshold. This allows the application to properly catch and manage dictionary overflow conditions during data loading operations.

core/src/main/java/org/apache/carbondata/core/localdictionary/exception · high confidence

Added Hadoop MapReduce integration for CarbonData

Introduced new Hadoop MapReduce components to enable direct reading and writing of CarbonData files via the Hadoop ecosystem. This includes the \CarbonFileInputFormat\ and \CarbonTableInputFormat\ for reading data, along with \CarbonTableOutputFormat\ and \CarbonOutputCommitter\ for writing data. These classes provide the necessary interfaces for Hadoop jobs to interact with CarbonData tables and segments, supporting both transactional and non-transactional table scenarios.

hadoop/src/main/java/org/apache/carbondata/hadoop/api · high confidence

Added KeyGeneratorFactory to centralize key generation logic

A new \KeyGeneratorFactory\ class has been introduced in the core module to manage the creation of \KeyGenerator\ instances. This factory determines whether to generate fully filled or variable-length multi-dimensional keys based on the \IS\_FULLY\_Filled\_BITS\ configuration, abstracting the selection logic away from direct instantiation.

core/src/main/java/org/apache/carbondata/core/keygenerator/factory · high confidence

Added Min-Max and Lucene index implementations for query optimization

Added new index implementations to improve query performance. This includes a Min-Max index example that tracks min/max values per blocklet to enable blocklet pruning, and a Lucene-based fine-grain index that supports text matching and blocklet-wise storage. The Lucene index allows for efficient text search and supports configuration for cache flushing and blocklet splitting. These changes provide new capabilities for optimizing data scanning and filtering.

index/lucene · high confidence

Added SQL parser and execution support for Spark 2.3/2.4

The integration layer now includes a new SQL parser grammar (CarbonSqlBase.g4) and a suite of adapter and execution classes (e.g., CarbonDataSourceScanHelper, SparkVersionAdapter, CarbonAnalyzer) specifically for Spark 2.3 and 2.4. This enables the system to parse and execute SQL commands, handle data source scans, and manage query plans for these specific Spark versions.

integration/spark/src/main · high confidence

Added Thrift schema conversion for table metadata

A new \SchemaConverter\ interface and its \ThriftWrapperSchemaConverterImpl\ implementation were added to the core metadata converter package. This introduces a dedicated mechanism to convert internal CarbonData schema objects (such as \TableSchema\, \ColumnSchema\, and \SchemaEvolution\) to and from the external Thrift format. This enables the system to serialize and deserialize table metadata structures for external storage or communication, supporting features like schema evolution and metadata persistence.

core/src/main/java/org/apache/carbondata/core/metadata/converter · high confidence

Added V3 format support for column chunks

The core datastore module introduces a new V3 data format for handling column chunks. This includes a new \AbstractRawColumnChunk\ class to manage uncompressed blocklets and min/max value metadata, alongside a \DimensionColumnPage\ interface that defines methods for data filling, vector conversion, and memory management. These changes enable the system to read and write data using the V3 format, which is designed to optimize scan performance and support features like adaptive encoding and page-level reading.

core/src/main/java/org/apache/carbondata/core/datastore/chunk · high confidence

Added V3 measure chunk page reader implementation

Introduced new Java classes, MeasureChunkReaderV3 and MeasureChunkPageReaderV3, to handle reading and decoding measure column data in the V3 file format. These components parse the V3 data structure, managing offsets and decompression for measure columns, enabling the system to read and process this specific data layout.

core/src/main/java/org/apache/carbondata/core/datastore/chunk/reader/measure/v3 · high confidence

Added benchmark and example scripts for Spark

Added new benchmark scripts (ConcurrentQueryBenchmark, SCDType2Benchmark, SimpleQueryBenchmark) and example scripts (AlluxioExample, AlterTableExample, CDCExample, CarbonDataFrameExample, CarbonSessionExample, CarbonSortColumnsExample, CaseClassDataFrameAPIExample, CustomCompactionExample, DataFrameComplexTypeExample) to demonstrate various CarbonData features including concurrent queries, SCD Type 2 handling, CDC, DataFrame APIs, and complex types.

examples/spark/src/main/scala · high confidence

Added build documentation and scripts for CarbonData notebook Docker images

The build directory now includes new documentation and scripts to help users build and run CarbonData notebook Docker images. This includes a Dockerfile that sets up a Jupyter Spark notebook environment with the CarbonData JAR, along with two Markdown guides: one for building the image via a Dockerfile and another for a manual build process. Additionally, shell and batch scripts (carbondata-build-info.sh and .bat) were added to generate build information properties files, and a README.md provides an overview of the build process and links to the new guides.

build · high confidence

Added custom log4j log levels and rolling file appender for audit and statistics logging

The CarbonData common module now includes new logging components to support specialized logging needs. A new \AuditLevel\ and \StatisticLevel\ have been introduced to provide custom log levels for audit trails and statistical data, respectively. Additionally, a new \AuditExtendedRollingFileAppender\ has been added to handle audit-specific log rotation, while \ExtendedRollingFileAppender\ provides a customizable rolling file appender that manages log file size and backup counts. These changes allow users to configure more granular logging for audit and statistics events.

common/src/main/java/org/apache/carbondata/common/logging/impl · high confidence

Added development tooling and style configurations

Added the \carbon\_pr.py\ script and its \carbon-pr-readme.md\ documentation to automate the process of reviewing and merging GitHub pull requests, including squashing commits and updating JIRA tickets. Additionally, introduced configuration files for code quality and style enforcement: \javastyle-config.xml\ and \javastyle-suppressions.xml\ for Checkstyle rules, \java-code-format-template.xml\ for IDE formatting, \findbugs-exclude.xml\ to suppress specific FindBugs warnings, and \java.header\ for license headers.

dev · high confidence

Added direct dictionary key generation for date and timestamp types

Introduced the DirectDictionaryGenerator interface and DirectDictionaryKeyGeneratorFactory to handle surrogate key generation for direct dictionary columns. The factory now supports DATE and TIMESTAMP data types, delegating to specific generators (DateDirectDictionaryGenerator and TimeStampDirectDictionaryGenerator) to manage key generation and value retrieval for these temporal types.

core/src/main/java/org/apache/carbondata/core/keygenerator/directdictionary · high confidence

Added geohash-based spatial index and polygon filtering support

The geo module now includes a new geohash-based spatial index implementation (GeoHashIndex) that generates column values from longitude and latitude, enabling efficient spatial queries. This change introduces polygon expression processors (PolygonExpression, PolygonListExpression, PolygonRangeListExpression, and PolylineListExpression) that allow filtering data based on geometric shapes. A dedicated filter executor (PolygonFilterExecutorImpl) is added to prune blocks and blocklets based on these spatial ranges, improving query performance for geospatial data.

geo · high confidence

Added interface audience and stability annotations

New Java annotation classes, InterfaceAudience and InterfaceStability, have been added to the common module. These annotations allow developers to mark public and developer-facing interfaces as User, Developer, or Internal, and to specify their stability level (Stable, Evolving, or Unstable). This provides clearer guidance on which APIs are safe for external use versus those that are internal or subject to change.

common/src/main/java/org/apache/carbondata/common/annotations · high confidence

Added new wrapper classes for scan data

Added new wrapper classes, specifically ByteArrayWrapper and IntArrayWrapper, to the core scan package. These classes provide structured containers for query scan data, enabling more efficient handling of dictionary keys, no-dictionary keys, and complex types during query execution and aggregation.

core/src/main/java/org/apache/carbondata/core/scan/wrappers · high confidence

Added new writers for delete delta and index file merging

The core module introduces new classes to support delete operations and index management: a \CarbonDeleteDeltaWriter\ interface and its implementation \CarbonDeleteDeltaWriterImpl\ for writing delete delta files, a \CarbonIndexFileWriter\ for serializing index data via Thrift, and a \ThriftWriter\ utility class that handles the low-level Thrift serialization and file I/O. These components enable the system to track and write delete deltas and merge index files more effectively.

core/src/main/java/org/apache/carbondata/core/writer · high confidence

Added safe dimension data chunk stores and vector fillers for direct scan queries

New classes were added to the core datastore module to support safe, direct-scan reading of dimension data. This includes \SafeAbstractDimensionDataChunkStore\ and its implementations (\SafeFixedLengthDimensionDataChunkStore\, \SafeVariableLengthDimensionDataChunkStore\, etc.) which store dimension data and inverted indices. Additionally, \AbstractNonDictionaryVectorFiller\ and its concrete fillers (e.g., \StringVectorFiller\, \BooleanVectorFiller\) were introduced to efficiently fill \CarbonColumnVector\ objects during direct scan queries, enabling optimized vectorized reading of non-dictionary columns.

core/src/main/java/org/apache/carbondata/core/datastore/chunk/store/impl/safe · high confidence

Added sample CSV datasets for Spark examples

Added new CSV resource files to the Spark examples directory, including test data for float datatype support (Test\_Data1.csv), complex data structures (complexdata.csv), and various sample datasets (data.csv, data1.csv, dataSample.csv, dimSample.csv, factSample.csv, sample.csv, streamSample.csv) to support example queries and testing.

examples/spark/src/main/resources · high confidence

Added streaming segment pruning with min/max index support

The core stream package now includes new classes to support streaming segment pruning. ExtendedByteArrayInputStream, ExtendedByteArrayOutputStream, and ExtendedDataInputStream provide extended byte stream handling. StreamFile models segment metadata including min/max index information, while StreamPruner implements logic to filter out segments that do not match filter criteria using min/max index data. This enables more efficient streaming data processing by skipping irrelevant segments.

core/src/main/java/org/apache/carbondata/core/stream · high confidence

Added streaming table support via new Hadoop input format and record reader

The Hadoop streaming module now includes new classes—CarbonStreamInputFormat, CarbonStreamRecordReader, StreamBlockletReader, and CarbonStreamUtils—to enable reading data from streaming tables. This introduces a new input format and record reader that handle blocklet parsing, header reading, and row iteration for stream-based data sources, allowing users to process streaming table data through the Hadoop MapReduce interface.

hadoop/src/main/java/org/apache/carbondata/hadoop/stream · high confidence

Added timestamp direct dictionary generators for date and timestamp types

Introduced new classes in the core module to handle direct dictionary generation for date and timestamp data types. The changes add \AbstractDirectDictionaryGenerator\, \DateDirectDictionaryGenerator\, and \TimeStampDirectDictionaryGenerator\ along with supporting constants and enums. This enables the system to generate surrogate keys for date and timestamp columns, supporting configurable time granularity (seconds, minutes, hours, days) and a customizable cutoff timestamp.

core/src/main/java/org/apache/carbondata/core/keygenerator/directdictionary/timestamp · high confidence

Added utility classes for iteration and string formatting

The common module now includes three new utility classes: CarbonIterator, which provides a base iterator with default remove and lifecycle methods; Maps, which adds a getOrDefault helper to avoid JDK 8 dependencies; and Strings, which provides a mkString utility and a formatSize method for human-readable byte size formatting.

common/src/main/java/org/apache/carbondata/common · high confidence

Bloom filter index implementation for blocklet-level pruning

Added a new Bloom filter index mechanism for blocklet-level pruning. This includes the core writer and builder classes (AbstractBloomIndexWriter, BloomIndexBuilder, BloomIndexWriter) that construct and maintain Bloom filters for indexed columns. A cache layer (BloomIndexCache, BloomCacheKeyValue) is introduced to store and retrieve Bloom filters efficiently. The query path is updated with BloomCoarseGrainIndex and its factory to perform pruning using these filters. Additionally, file storage and merging logic (BloomIndexFileStore) and input split handling (BloomIndexInputSplit) are added to support the new index type.

index/bloom · high confidence

Introduce AI\_carbon module for Agent project archives

The Agent\_module now includes the AI\_carbon module, which provides a unified \.AI\_carbon\ archive format for storing Agent-generated files, their context, and conversation history. This module enables developers to create, inspect, and optimize project archives using a Python API, a local web management interface, and command-line tools. The module also includes a Codex skill (\ai-carbon-sync\) to automatically synchronize generated files and context into these archives during development sessions.

_Agent\module · high confidence

Introduce CarbonCli command-line tool for data inspection and benchmarking

A new standalone CLI tool, CarbonCli, is added to the CarbonData project to inspect and analyze CarbonData files. Users can run commands such as 'summary' to view schema, segment, and table properties, 'benchmark' to measure read performance, and 'sort\_columns' to retrieve sort column information. The tool supports various flags to control output detail, such as printing all information, schema, segment details, blocklet details, and column statistics.

tools/cli · high confidence

Introduce CarbonTablePath utility for centralized path construction

A new \CarbonTablePath\ utility class has been added to the \core/src/main/java/org/apache/carbondata/core/util/path\ package. This class provides static methods to construct and identify paths for various table components, including metadata, schema files, dictionary files, segments, and stage data. This centralizes path logic, replacing scattered string concatenation with a unified API for managing table file locations.

core/src/main/java/org/apache/carbondata/core/util/path · high confidence

Introduce LocalDictDimensionDataChunkStore for local dictionary encoding

A new \LocalDictDimensionDataChunkStore\ class has been added to handle dimension data with local dictionary encoding. This implementation wraps an existing \DimensionDataChunkStore\ and a \CarbonDictionary\ to manage data storage and retrieval using surrogate keys, enabling more efficient processing of dimension data through dictionary-based lookups.

core/src/main/java/org/apache/carbondata/core/datastore/chunk/store/impl · high confidence

Introduce PyCarbon Python SDK for AI framework integration

Added the PyCarbon Python SDK, providing a unified API for reading CarbonData datasets into AI frameworks like TensorFlow and PyTorch. The change introduces the \pycarbon\ package with core modules for dataset handling (\CarbonDataset\), filesystem resolution (HDFS, S3, local), and reader workers (Arrow and PyDict). This enables efficient data loading and filtering for machine learning training pipelines.

python · high confidence

Introduce RLE codec for integral column pages

Added RLECodec and RLEEncoderMeta classes to implement Run-Length Encoding for integral column pages. This new encoding supports boolean, byte, short, int, and long data types, enabling more efficient storage and faster processing for repeated values in columnar data.

core/src/main/java/org/apache/carbondata/core/datastore/page/encoding/rle · high confidence

Introduce V3 data reader interfaces and factory

Added new interfaces (DimensionColumnChunkReader, MeasureColumnChunkReader) and a factory (CarbonDataReaderFactory) to support reading data in the V3 columnar format. The factory routes requests to V3-specific readers (e.g., DimensionChunkPageReaderV3, MeasureChunkPageReaderV3) when the V3 format version is specified, enabling the system to read data using the new V3 structure.

core/src/main/java/org/apache/carbondata/core/datastore/chunk/reader · high confidence

Introduce V3 dimension chunk readers for improved data access

The core module now includes new V3 dimension chunk readers (AbstractDimensionChunkReader, DimensionChunkReaderV3, and DimensionChunkPageReaderV3) that implement a more efficient, page-level reading strategy for dimension columns. This change modifies how the database engine reads and decodes dimension data from storage, which may affect query performance and memory usage for dimension column scans.

core/src/main/java/org/apache/carbondata/core/datastore/chunk/reader/dimension · high confidence

Introduce a new LRU-based in-memory cache framework

Added a new caching infrastructure in the core module, introducing a \Cache\ interface, a \CacheProvider\ for instantiation, and a \CarbonLRUCache\ implementation backed by an \ExpiringMap\. This replaces the previous Guava-based cache with a configurable LRU cache that respects JVM memory limits and supports access-based expiration policies.

core/src/main/java/org/apache/carbondata/core/cache · high confidence

Introduce abstract base class for measure chunk reading

The codebase now includes an abstract \AbstractMeasureChunkReader\ class that serves as the foundation for reading measure column chunks. This new component centralizes the logic for reading raw measure data in groups based on block indexes, providing a reusable template for specific reader implementations.

core/src/main/java/org/apache/carbondata/core/datastore/chunk/reader/measure · high confidence

Introduce adaptive encoding codecs for column pages

The core module now includes a new \adaptive\ package containing \AdaptiveCodec\ and its implementations (\AdaptiveDeltaFloatingCodec\, \AdaptiveDeltaIntegralCodec\, \AdaptiveFloatingCodec\, \AdaptiveIntegralCodec\). These classes implement the \ColumnPageCodec\ interface to handle encoding and decoding of column pages using adaptive strategies, including delta encoding for floating-point and integral types, and type casting to minimize storage size.

core/src/main/java/org/apache/carbondata/core/datastore/page/encoding/adaptive · high confidence

Introduce core block metadata and distribution classes

Added new classes in the \core/src/main/java/org/apache/carbondata/core/datastore/block\ package to manage block-level metadata and task distribution. \AbstractIndex\ provides a base implementation for index structures with access counting and memory size tracking. \Distributable\ is a new interface for retrieving node locations for task distribution based on locality. \SegmentProperties\ and \SegmentPropertiesAndSchemaHolder\ manage the restructuring and column mapping details for each segment, with the holder acting as a singleton cache for segment properties. \TableBlockInfo\ and \TaskBlockInfo\ are new classes to pass block details and task-specific block lists to executors, supporting sorting by block size and serialization for distributed processing.

core/src/main/java/org/apache/carbondata/core/datastore/block · high confidence

Introduce core infrastructure for Materialized Views

Added new classes in the core view package to support materialized views, including MVSchema for storing view definitions and properties, MVManager for managing schemas, MVProvider for persistence, and MVCatalog for in-memory registry. These changes provide the internal foundation for materialized view functionality.

core/src/main/java/org/apache/carbondata/core/view · high confidence

Introduce event-driven architecture for operations

Added a new eventing framework in the core module, introducing the \Event\ base class, \OperationContext\ for carrying operation data, an \OperationEventListener\ interface for subscribers, and an \OperationListenerBus\ singleton to manage listener registration and event firing. This provides a centralized, thread-safe mechanism for decoupling components via publish/subscribe events during operations.

core/src/main/java/org/apache/carbondata/events · high confidence

Introduce index type enumeration and block index info model

Added the IndexType enum to define supported index providers (Lucene, Bloomfilter, and Secondary Index) and their corresponding factory class names, alongside a new BlockIndexInfo class to store block-level index metadata including row counts, file offsets, and blocklet information. This change provides the core data structures and type definitions needed for the new index metadata handling.

core/src/main/java/org/apache/carbondata/core/metadata/index · high confidence

Introduce local dictionary encoding and fallback mechanism for blocklet data

The system now supports local dictionary encoding for column pages within a blocklet. When a column is encoded using a local dictionary, the system tracks the page-level dictionary and manages a fallback mechanism. If pages are not encoded with a local dictionary, the system initiates a fallback process to re-encode those pages, ensuring consistent handling of encoded and non-encoded pages within the same blocklet. This change introduces new classes \BlockletEncodedColumnPage\ and \EncodedBlocklet\ to manage these encoded pages and their fallback states.

core/src/main/java/org/apache/carbondata/core/datastore/blocklet · high confidence

Introduce local dictionary generation interface and implementation

Added the LocalDictionaryGenerator interface and its ColumnLocalDictionaryGenerator implementation to handle per-column dictionary generation. The new generator manages a threshold-based dictionary store, allowing the system to build and retrieve local dictionaries for column data while tracking when the dictionary size threshold is reached.

core/src/main/java/org/apache/carbondata/core/localdictionary/generator · high confidence

Introduce modular plan architecture for materialized views

Added a new modular plan representation for materialized views, including core plan nodes (Select, GroupBy, Union), expression types (ScalarModularSubquery, ModularListQuery, ModularExists), and supporting utilities (DSL, Harmonizer, SignatureGenerator). This refactors the query plan structure to support multi-tenant and incremental data load scenarios for materialized views.

mv/plan · high confidence

Introduce new core classes for handling update and delete delta operations

Added several new classes in the \core/src/main/java/org/apache/carbondata/core/mutate\ package to support update and delete delta functionality. These include \CarbonUpdateUtil\ for utility methods related to tuple IDs and block paths, \CdcVO\ for cache loading in the Index server, \DeleteDeltaBlockDetails\ and \DeleteDeltaBlockletDetails\ to store block details for delete delta files, \DeleteDeltaVo\ to track deleted rows using a BitSet, \FilePathMinMaxVO\ to store file paths and min/max values for each blocklet, \SegmentUpdateDetails\ to manage segment update status, \TupleIdEnum\ to define tuple ID indices, and \UpdateVO\ to store update operation details. These changes provide the foundational data structures and utilities required for processing and managing update and delete delta files within the CarbonData core module.

core/src/main/java/org/apache/carbondata/core/mutate · high confidence

Introduce new data type processing classes and exception handling

The processing module now includes new classes for handling complex data types during data loading. Specifically, it adds \GenericDataType\ as an interface for complex types like \ArrayDataType\ and \StructDataType\, alongside \PrimitiveDataType\ for standard columns. These classes manage the serialization and deserialization of data during the load process. Additionally, new exception classes (\DataLoadingException\, \MultipleMatchingException\) and a log resource file (\CARBON\_PROCESSINGLogResource.properties\) are introduced to support error handling and logging within the data processing pipeline.

processing · high confidence

Introduce new index metadata models for secondary index support

Added two new Java classes, IndexMetadata and IndexTableInfo, to the core metadata schema package. These models store and manage secondary index information, including provider mappings, parent table details, index columns, and status tracking, enabling the system to handle coarse-grain and fine-grain index metadata for features like Presto query integration.

core/src/main/java/org/apache/carbondata/core/metadata/schema/indextable · high confidence

Introduce new memory management classes for off-heap and on-heap allocations

Added a new set of classes in the core memory package to manage memory allocation, including CarbonUnsafe, HeapMemoryAllocator, UnsafeMemoryAllocator, UnsafeMemoryManager, and UnsafeSortMemoryManager. These classes provide mechanisms for allocating and freeing memory blocks, supporting both off-heap and on-heap memory types, and managing memory usage for tasks. The implementation includes utilities for unsafe memory operations, memory block tracking, and memory pool management.

core/src/main/java/org/apache/carbondata/core/memory · high confidence

Introduce new result iterator implementations for query execution

Added new iterator classes in the scan result iterator package to handle query execution and data retrieval. This includes \AbstractDetailQueryResultIterator\ as a base class for detail queries, \DetailQueryResultIterator\ for row-based results, \VectorDetailQueryResultIterator\ for vector batch processing, \RawResultIterator\ for raw row processing with prefetching support, \ChunkRowIterator\ for chunked row iteration, \ColumnDriftRawResultIterator\ to handle column drift scenarios, and \PartitionSplitterRawResultIterator\ for partitioned data splitting. These changes provide the underlying mechanism for iterating over query results in the core scanning engine.

core/src/main/java/org/apache/carbondata/core/scan/result/iterator · high confidence

Introduce new table metadata schema classes

Added new classes to the core metadata schema package, including PartitionType, CarbonTable, CarbonTableBuilder, IndexSchema, RelationIdentifier, TableInfo, TableSchema, TableSchemaBuilder, Writable, and WritableUtil. These classes define the structure for table metadata, index schemas, and partition types, supporting features like non-transactional tables, local dictionaries, and schema evolution.

core/src/main/java/org/apache/carbondata/core/metadata/schema/table · high confidence

Introduce page-level dictionary for blocklet storage

A new PageLevelDictionary class has been added to manage page-level dictionary values for a column. This class generates and stores unique dictionary values for each page, supporting both primitive and complex types, and handles the encoding and compression of the local dictionary chunk for blocklet-level dictionary storage in CarbonData files.

core/src/main/java/org/apache/carbondata/core/localdictionary · high confidence

Introduce pluggable file system abstraction for CarbonData

The core module now provides a new pluggable file system abstraction via the \CarbonFile\ interface and its implementations (\LocalCarbonFile\, \HDFSCarbonFile\, \S3CarbonFile\, \AlluxioCarbonFile\, and \ViewFSCarbonFile\). This change allows CarbonData to interact with various storage backends—including local file systems, HDFS, S3, and Alluxio—through a unified API, enabling support for multiple distributed file systems and cloud storage services.

core/src/main/java/org/apache/carbondata/core/datastore/filesystem · high confidence

Introduce pluggable file system abstraction for data storage

The core module now supports pluggable file operations through a new \FileTypeInterface\ and \DefaultFileTypeProvider\ in the \datastore/impl\ package. This change introduces a \FileFactory\ that routes file access requests to specific implementations like \DFSFileReaderImpl\ for HDFS/Alluxio/S3 or \FileReaderImpl\ for local storage, enabling the system to handle diverse storage backends and custom file providers.

core/src/main/java/org/apache/carbondata/core/datastore/impl · high confidence

Introduce pluggable file-locking framework for concurrent operations

The core module now includes a new \locks\ package that provides a unified, pluggable locking mechanism for coordinating concurrent operations. The \CarbonLockFactory\ dynamically selects the appropriate lock implementation—such as \HdfsFileLock\, \S3FileLock\, \AlluxioFileLock\, \LocalFileLock\, or \ZooKeeperLocking\—based on the underlying file system or configured lock type. This abstraction allows CarbonData to manage metadata, table status, and segment locks consistently across different storage backends, ensuring thread safety for operations like compaction, load, and drop table.

core/src/main/java/org/apache/carbondata/core/locks · high confidence

Introduce vectorized columnar batch processing for scan results

Added new classes in the \core/scan/result/vector\ package to support vectorized data reading. This includes the \CarbonColumnVector\ interface for typed data storage, \CarbonColumnarBatch\ to manage batches of column vectors, \CarbonDictionary\ for dictionary lookups, \ColumnVectorInfo\ to hold metadata about each column's state, and \MeasureDataVectorProcessor\ with fillers for various data types (integral, boolean, short, etc.) to efficiently fill vectors from column pages.

core/src/main/java/org/apache/carbondata/core/scan/result/vector · high confidence

Introduced vectorized column and dictionary implementations

Added new concrete implementations for column vectors and local dictionaries within the scan result package. The \CarbonColumnVectorImpl\ class provides a typed storage mechanism for various data types (including boolean, float, short, int, long, float, double, decimal, string, and complex types) to support efficient vectorized reading. Additionally, \CarbonDictionaryImpl\ was introduced to manage local dictionary data, enabling optimized lookups and data retrieval during scan operations.

core/src/main/java/org/apache/carbondata/core/scan/result/vector/impl · high confidence

Introduces a custom logging service and factory

Added LogService and LogServiceFactory classes to the common logging module. LogService extends log4j's Logger to provide a unified logging interface with support for audit and statistic logging levels, while LogServiceFactory provides a static method to retrieve logger instances by class name.

common/src/main/java/org/apache/carbondata/common/logging · high confidence

Introduces a new ResultCollectorFactory to manage scan result collection

A new ResultCollectorFactory has been added to the core scan module, providing a centralized way to create specific result collector instances (such as DictionaryBased, RawBased, and Vector-based collectors) based on the query's execution context. This factory coordinates with the new ScannedResultCollector interface, which defines methods for collecting results in both row and columnar batch formats, enabling the system to select the appropriate collector implementation for different scan scenarios.

core/src/main/java/org/apache/carbondata/core/scan/collector · high confidence

Introduces atomic file operations for S3 and Alluxio

Added a new \AtomicFileOperations\ interface and implementations to handle file writes more safely. For S3 and Alluxio storage, the system now performs direct overwrites to avoid the brief window where a file might not exist, preventing access failures. For standard file systems, it uses a temporary file that is renamed upon successful completion, ensuring data integrity during data loads, updates, or compactions.

core/src/main/java/org/apache/carbondata/core/fileoperations · high confidence

Introduces new column page encoding and decoding interfaces

The core module now includes new interfaces and classes for column page encoding and decoding, including ColumnPageCodec, ColumnPageEncoder, ColumnPageDecoder, and related metadata and factory classes. This change introduces a new encoding strategy for column pages, allowing for more flexible and efficient data storage and retrieval.

core/src/main/java/org/apache/carbondata/core/datastore/page/encoding · high confidence

Introduces new column page implementations and a decoder-based fallback mechanism for local dictionary encoding

The core data store page module now includes a suite of new classes to handle column data more efficiently. A new \ColumnPage\ hierarchy supports both safe and unsafe memory modes for various data types, including decimal, complex, and variable-length byte arrays. The update adds a \DecoderBasedFallbackEncoder\ that allows the system to fall back to decoding actual data when local dictionary encoding is not optimal, reducing memory footprint and improving query performance. Additionally, a \LazyColumnPage\ decorator is introduced to perform decoding lazily, and a \ComplexColumnPage\ manages nested data structures. These changes support more flexible encoding strategies and better memory management during data loading and querying.

core/src/main/java/org/apache/carbondata/core/datastore/page · high confidence

Introduces new core datastore abstractions for data access and column typing

The core datastore package now includes new classes and interfaces that define how data is read and typed. A new \ColumnType\ enum categorizes columns into global/direct/plain/complex/measures, while \DataRefNode\ provides an interface for iterating over data blocks and reading dimension/measure chunks. A \FileReader\ interface standardizes byte-level file access, and a \ReusableDataBuffer\ class manages memory-efficient buffer reuse. Additionally, \TableSpec\ and \TableSegmentUniqueIdentifier\ handle table metadata and segment identification, supporting the underlying data storage and retrieval mechanisms.

core/src/main/java/org/apache/carbondata/core/datastore · high confidence

Introduces new index store classes for blocklet and extended blocklet handling

The core module adds several new classes to support the index server and distributed index architecture. This includes \ExtendedBlocklet\ and \ExtendedBlockletWrapper\ to carry detailed blocklet information and handle serialization for network transfer, \BlockletIndexStore\ and \BlockletIndexWrapper\ for managing blocklet indexes in a cache, and \AbstractMemoryDMStore\ with \SafeMemoryDMStore\ for in-memory index row storage. Additionally, new types like \BlockMetaInfo\, \BlockletDetailInfo\, and \PartitionSpec\ are introduced to support metadata and partitioning information.

core/src/main/java/org/apache/carbondata/core/indexstore · high confidence

Introduces new query model classes for scan execution

The core scan model is expanded with new classes—ProjectionColumn, ProjectionDimension, ProjectionMeasure, QueryModel, QueryModelBuilder, and QueryProjection—designed to carry query plan details from the driver to the executor. These classes encapsulate projection and filter information, enabling the execution engine to process queries more efficiently by providing raw detailed records and supporting vector-based row pruning push-down.

core/src/main/java/org/apache/carbondata/core/scan/model · high confidence

Introduction of CarbonReadSupport interface for data reading abstraction

A new interface, CarbonReadSupport, has been added to the Hadoop read support module. This interface defines the contract for converting data read via RecordReader into row representations, providing methods for initialization, reading rows, and cleanup. This change introduces a new abstraction layer for data reading operations within the CarbonData Hadoop integration.

hadoop/src/main/java/org/apache/carbondata/hadoop/readsupport · high confidence

Legacy dimension index codecs added to the core datastore

The core module now includes a new \legacy\ subpackage containing four new classes—\ComplexDimensionIndexCodec\, \DirectDictDimensionIndexCodec\, \PlainDimensionIndexCodec\, and their supporting base classes \IndexStorageCodec\ and \IndexStorageEncoder\. These classes implement the \ColumnPageCodec\ interface to handle the encoding of dimension data pages, supporting various encoding types such as DICTIONARY, INVERTED\_INDEX, RLE, and DIRECT\_COMPRESS\_VARCHAR. This change introduces the underlying storage and compression logic for dimension index data.

core/src/main/java/org/apache/carbondata/core/datastore/page/encoding/dimension · high confidence

New BiDictionary interface for data loading

A new BiDictionary interface has been added to the core API, providing methods to manage bidirectional key-value mappings for data loading operations. This includes retrieving or generating keys for values, looking up keys by value, and retrieving values by key, along with a size query.

core/src/main/java/org/apache/carbondata/core/devapi · high confidence

New BlockletFilterScanner and BlockletFullScanner classes for query scanning

The core scan scanner implementation now includes two new classes, BlockletFilterScanner and BlockletFullScanner, which handle the processing of blocklets for filter and full scan queries respectively. BlockletFilterScanner applies filter conditions and supports min/max range checks, while BlockletFullScanner handles full scans without filters. These changes introduce new scanning logic for query execution, affecting how data is read and filtered at the blocklet level.

core/src/main/java/org/apache/carbondata/core/scan/scanner/impl · high confidence

New CLI and server management scripts for CarbonData

Added new shell scripts to the bin directory: a CarbonData-specific Spark SQL CLI launcher (carbon-spark-sql) and scripts to start (start-indexserver.sh) and stop (stop-indexserver.sh) the distributed index server. These scripts provide a convenient way to launch the SQL command-line interface and manage the background index server process.

bin · high confidence

New Hadoop input split and serialization support for distributed reads

Added new Hadoop classes to support distributed reading of CarbonData files. The \CarbonInputSplit\ class now defines the structure for input splits, including fields for segment, bucket, blocklet, and delete delta files, along with serialization logic for Hadoop's MapReduce framework. A new \CarbonInputSplitWrapper\ handles the serialization and deserialization of multiple \CarbonInputSplit\ objects. Additionally, \ObjectArrayWritable\ provides a Hadoop \Writable\ implementation for object arrays, and the \Block\ interface defines the contract for HDFS blocks, enabling the system to manage blocklet-level scanning and full scan requirements during distributed processing.

core/src/main/java/org/apache/carbondata/hadoop · high confidence

New Hadoop integration classes for CarbonData

Added new Hadoop integration classes including AbstractRecordReader, CarbonMultiBlockSplit, CarbonProjection, CarbonRecordReader, and InputMetricsStats. These classes provide the foundational components for reading CarbonData files via Hadoop's MapReduce framework, supporting both single and multi-block splits for concurrent query optimization.

hadoop/src/main/java/org/apache/carbondata/hadoop · high confidence

New Hadoop utility classes for vectorized reading and split management

Added three new utility classes in the Hadoop module: CarbonInputFormatUtil, which provides factory methods to configure and create CarbonTableInputFormat instances; CarbonInputSplitTaskInfo, which manages task information and location balancing for input splits; and CarbonVectorizedRecordReader, which enables direct reading of data into columnar batches for improved query performance.

hadoop/src/main/java/org/apache/carbondata/hadoop/util · high confidence

New Hive integration components for CarbonData

Added new Java classes to the Hive integration module, including a storage handler, SerDe, record reader, input/output formats, and expression converters. These additions enable Hive to read and write CarbonData files directly, supporting schema inference, data type conversion, and query predicate pushdown.

integration/hive/src/main/java/org/apache/carbondata/hive · high confidence

New QueryExecutor interface and factory for row-based scanning

A new QueryExecutor interface and a corresponding QueryExecutorFactory have been added to the core scan executor package. The interface defines the contract for executing queries based on a query model and returning an iterator over results, while the factory selects the appropriate executor implementation (vector or detail) based on the query type. This provides a standardized way to execute row-based CarbonRecordReader queries.

core/src/main/java/org/apache/carbondata/core/scan/executor · high confidence

New SDK and SQL example programs for CarbonData

Added five new Java example programs in the Spark examples directory: CarbonReaderExample demonstrates reading and writing data using the CarbonData SDK; SDKS3Example and SDKS3ReadExample show how to write to and read from S3 storage using the SDK; SDKS3SchemaReadExample illustrates reading schema information from S3; and JavaCarbonSessionExample provides a complete SQL-based workflow using the CarbonSession. These examples cover basic I/O, S3 integration, schema reading, and SQL operations.

examples/spark/src/main/java · high confidence

New SQL exception types for schema, index, and materialized view operations

The CarbonData common module introduces several new exception classes in the \org.apache.carbondata.common.exceptions.sql\ package to handle specific SQL command failures. These include \CarbonSchemaException\ for schema-related errors, \MalformedCarbonCommandException\ as a base for malformed commands, and specific subclasses for index and materialized view operations: \MalformedIndexCommandException\, \MalformedMVCommandException\, \NoSuchIndexException\, and \NoSuchMVException\. These changes provide more granular error handling for SQL parsing and execution, particularly supporting features like materialized views and index management.

common/src/main/java/org/apache/carbondata/common/exceptions/sql · high confidence

New Thrift schema definitions for CarbonData V3 format

The \format\ module now includes new Thrift interface definitions (\carbondata.thrift\, \carbondata\_index.thrift\, \carbondata\_index\_merge.thrift\, \dictionary.thrift\, and \schema.thrift\) that describe the structure of the V3 file format. These changes introduce new data structures such as \DataChunk3\ and \BlockletInfo3\ to support the V3 storage layout, alongside updated schema definitions for column encodings, partitioning, and bucketing.

format · high confidence

New blocklet index metadata classes for min/max and B-Tree storage

Three new Java classes—BlockletBTreeIndex, BlockletIndex, and BlockletMinMaxIndex—have been added to the core metadata package. These classes define the data structures for storing blocklet index information, including B-Tree indices and min/max value tracking for columns within a blocklet, enabling the system to persist and retrieve blocklet-level metadata.

core/src/main/java/org/apache/carbondata/core/metadata/blocklet/index · high confidence

New blocklet-level index implementation for finer-grained caching and query performance

The core module introduces a new blocklet-level index implementation, including \BlockIndex\, \BlockletIndex\, \BlockletDataRefNode\, \BlockletIndexFactory\, and related model classes. This change adds support for caching index data at the blocklet level, which allows the system to cache and retrieve index information more granularly. This is expected to improve query performance by reducing the amount of index data loaded into memory and enabling more efficient filtering and scanning of data blocks.

core/src/main/java/org/apache/carbondata/core/indexstore/blockletindex · high confidence

New column page and raw chunk implementations for dimension and measure data

The core datastore module introduces new classes to handle dimension and measure column data: AbstractDimensionColumnPage, DimensionRawColumnChunk, FixedLengthDimensionColumnPage, MeasureRawColumnChunk, and VariableLengthDimensionColumnPage. These classes provide the underlying storage and decoding logic for column pages, enabling the system to read, decode, and fill vector data from raw column chunks. This change supports the new V3 format and improves query performance by optimizing how column pages are accessed and processed.

core/src/main/java/org/apache/carbondata/core/datastore/chunk/impl · high confidence

New columnar block indexer storage classes for optimized data access

Added a new set of classes in the core datastore columnar package to handle block indexing and data compression. The abstract base class BlockIndexerStorage and its concrete implementations (ByteArrayBlockIndexerStorage, ObjectArrayBlockIndexerStorage, etc.) provide mechanisms for row ID encoding and run-length encoding (RLE) of data pages. This includes a new UnBlockIndexer utility for decompressing index and data structures, supporting more efficient columnar data retrieval and storage.

core/src/main/java/org/apache/carbondata/core/datastore/columnar · high confidence

New conditional expression types for filtering

The core module introduces a new \conditional\ subpackage containing a set of new expression classes that implement conditional and implicit filtering logic. This includes base and concrete implementations for comparison operators (EqualTo, NotEquals, GreaterThan, GreaterThanEqualTo, LessThan, LessThanEqualTo, StartsWith), set membership (In, NotIn), and implicit block scanning (ImplicitExpression, CDCBlockImplicitExpression). These classes provide the evaluation logic for these filter types during scan execution.

core/src/main/java/org/apache/carbondata/core/scan/expression/conditional · high confidence

New configuration templates for CarbonData

Added carbon.properties.template and dataload.properties.template files that define default settings for system, performance, and data loading configurations, including parameters for sorting, compaction, and CSV parsing.

conf · high confidence

New data transfer objects for block and row count details

Added BlockMappingVO and RowCountDetailsVO classes in the core mutate data package. These new value objects encapsulate block-to-segment mappings, row counts, and segment block counts, providing structured data containers for segment and block information.

core/src/main/java/org/apache/carbondata/core/mutate/data · high confidence

New exception classes for data writing and index building

Added CarbonDataWriterException and IndexBuilderException to the core datastore exception package, providing specific error handling for data writing and index building operations.

core/src/main/java/org/apache/carbondata/core/datastore/exception · high confidence

Added new classes to read specific metadata and delta files: CarbonDeleteDeltaFileReader and its implementation for processing delete deltas; CarbonDeleteFilesDataReader for reading multiple delete delta files concurrently; CarbonDictionaryColumnMetaChunk and CarbonDictionaryReader for managing and reading dictionary data; CarbonFooterReader and CarbonFooterReaderV3 for reading file footers (including version 3); CarbonHeaderReader for reading file headers and schemas; CarbonIndexFileReader for reading index files; and a generic ThriftReader to handle Thrift-based file formats. These components enable the system to read and process delete deltas, dictionary information, file headers, footers, and index data.

core/src/main/java/org/apache/carbondata/core/reader · high confidence

New filter execution and resolution classes for scan optimization

Added new classes to the core scan filter package to support optimized filter execution and resolution. This includes \ColumnFilterInfo\ for managing filter data, \FilterExecutorUtil\ for executing measure filters based on data types, \FilterExpressionProcessor\ for resolving filter expression trees, and various resolver implementations like \FalseConditionalResolverImpl\ and \TrueConditionalResolverImpl\ to handle boolean filter conditions.

core/src/main/java/org/apache/carbondata/core/scan/filter · high confidence

New filter executor implementations for query optimization

The core scan filter package introduces a suite of new filter executor classes to handle various filtering scenarios more efficiently. This includes \AndFilterExecutorImpl\ and \OrFilterExecutorImpl\ for combining filter conditions, \IncludeFilterExecutorImpl\ and \ExcludeFilterExecutorImpl\ for standard inclusion/exclusion logic, and \FalseFilterExecutor\ for handling false conditions. Additionally, \RangeValueFilterExecutorImpl\ is added to support range-based filtering, while \CDCBlockImplicitExecutorImpl\ and \ImplicitIncludeFilterExecutorImpl\ enable block and blocklet pruning based on implicit column filters. The update also introduces supporting infrastructure such as the \BitSetUpdaterFactory\ and \FilterBitSetUpdater\ interface to manage bitset operations for filter results.

core/src/main/java/org/apache/carbondata/core/scan/filter/executer · high confidence

New filter interface and enum types for query optimization

The core module introduces a new filter interface package (core/src/main/java/org/apache/carbondata/core/scan/filter/intf) containing the ExpressionType and FilterExecutorType enums, along with the FilterOptimizer interface and RowImpl/RowIntf classes. These additions provide the internal types and abstractions required for the upcoming filter optimization and execution logic, enabling the query engine to process and optimize filter expressions more efficiently.

core/src/main/java/org/apache/carbondata/core/scan/filter/intf · high confidence

New filter resolver implementations for row-level and range-based filtering

The core scan filter resolver package now includes new implementations for handling row-level and range-based filters. Specifically, \ConditionalFilterResolverImpl\ provides the base logic for resolving conditional expressions, while \LogicalFilterResolverImpl\ handles logical (AND/OR) combinations of filters. Additionally, \RowLevelFilterResolverImpl\ and \RowLevelRangeFilterResolverImpl\ introduce specialized resolvers for row-level filtering and range-based filtering respectively, enabling the system to resolve and execute these specific filter types during scan operations.

core/src/main/java/org/apache/carbondata/core/scan/filter/resolver · high confidence

New filter resolver visitors for direct, implicit, and no-dictionary columns

The core module introduces a new visitor-based architecture for resolving filter information, adding \ResolvedFilterInfoVisitorIntf\ and specific implementations: \CustomTypeDictionaryVisitor\, \ImplicitColumnVisitor\, \MeasureColumnVisitor\, \NoDictionaryTypeVisitor\, \RangeDirectDictionaryVisitor\, and \RangeNoDictionaryTypeVisitor\, all selected via \FilterInfoTypeVisitorFactory\. This change enables the query engine to resolve filter conditions directly for columns with direct, implicit, or no dictionary encoding, improving filter evaluation for high-cardinality and measure columns.

core/src/main/java/org/apache/carbondata/core/scan/filter/resolver/resolverinfo/visitor · high confidence

New info classes for query execution metadata

Added new classes BlockExecutionInfo, DeleteDeltaInfo, DimensionInfo, and MeasureInfo in the core scan executor infos package. These classes encapsulate metadata required for query execution, including block indices, measure and dimension details, delete delta file information, and default values for missing columns.

core/src/main/java/org/apache/carbondata/core/scan/executor/infos · high confidence

New key generation infrastructure for multi-dimensional keys

The core module introduces a new key generation framework for handling multi-dimensional keys. This includes a \KeyGenerator\ interface, a \KeyGenException\ class, and an \AbstractKeyGenerator\ base class. The \mdkey\ sub-package provides the concrete \MultiDimKeyVarLengthGenerator\ implementation, which uses a \Bits\ utility class to manage bit-level operations for generating and parsing byte arrays from multiple dimension keys.

core/src/main/java/org/apache/carbondata/core/keygenerator/mdkey · high confidence

New logical expression and exception classes for filter evaluation

The core module introduces new classes to support logical filter evaluation: \AndExpression\, \OrExpression\, \RangeExpression\, \TrueExpression\, and \FalseExpression\, along with the base \BinaryLogicalExpression\ and related exception classes (\FilterIllegalMemberException\, \FilterUnsupportedException\). These additions enable the system to evaluate logical conditions (AND, OR, range, true/false) during data scanning, providing the underlying logic for filter execution.

core/src/main/java/org/apache/carbondata/core/scan/expression/logical · high confidence

Added BlockletInfo and DataFileFooter classes to the core metadata package. These new classes store metadata about blocklets and data files, including row counts, chunk offsets, index information, and schema details. This change introduces new data structures for managing blocklet and file-level metadata, which may affect how metadata is serialized and accessed.

core/src/main/java/org/apache/carbondata/core/metadata/blocklet · high confidence

New metadata model classes for table and column identification

The core metadata package now includes new classes: AbsoluteTableIdentifier, CarbonTableIdentifier, ColumnIdentifier, ColumnarFormatVersion, DatabaseLocationProvider, SegmentFileStore, and ValueEncoderMeta, alongside IndexProperty in the schema/index subpackage. These classes provide the foundational identifiers, versioning, and storage structures for table and column metadata, supporting the updated internal metadata model.

core/src/main/java/org/apache/carbondata/core/metadata · high confidence

New modular compression framework with Gzip and Zstd support

The core datastore compression layer has been refactored into a pluggable architecture. A new \Compressor\ interface and \AbstractCompressor\ base class provide a unified API for compression and decompression. The \CompressorFactory\ now supports plugging in custom compressors via reflection, while natively supporting Snappy, Gzip, and Zstd. Users can now configure the column compression algorithm (e.g., switching from the default to Gzip or Zstd) to optimize for storage or performance.

core/src/main/java/org/apache/carbondata/core/datastore/compression · high confidence

New page statistics collection framework for columnar data

The core library introduces a new \ColumnPageStatsCollector\ interface and several implementations (\PrimitivePageStatsCollector\, \StringStatsCollector\, \KeyPageStatsCollector\, \DummyStatsCollector\) to collect min/max statistics for different data types. A \SimpleStatsResult\ interface and \TablePageStatistics\ class are added to aggregate and manage these statistics for dimensions and measures, enabling more efficient storage and filtering by tracking value ranges per page.

core/src/main/java/org/apache/carbondata/core/datastore/page/statistics · high confidence

New query and restructuring utility classes for the scan executor

Added QueryUtil and RestructureUtil classes in the core scan executor utility package. QueryUtil provides helper methods for query execution, including mapping dimensions and measures to their respective block indexes, identifying sort dimensions, and handling masked keys. RestructureUtil handles restructuring logic, specifically creating dimension and measure information for block execution, matching query dimensions against table block dimensions, and managing complex type children. These utilities support the query scan and restructuring processes within the core module.

core/src/main/java/org/apache/carbondata/core/scan/executor/util · high confidence

New query statistics recording framework for driver and executor phases

The core stats module now includes a new set of classes to record and log query performance metrics. A \QueryStatisticsRecorder\ interface and its implementations (\QueryStatisticsRecorderImpl\ for active recording, \QueryStatisticsRecorderDummy\ as a no-op placeholder) manage the collection of timing and count data. The \TaskStatistics\ class structures this data into a tabular format for logging, while \QueryStatistic\ and \QueryStatisticsConstants\ define the specific metrics tracked, such as SQL parse time, block allocation, and scan times. This provides the foundation for detailed query execution statistics.

core/src/main/java/org/apache/carbondata/core/stats · high confidence

New query type classes for complex data types

Added new classes (ArrayQueryType, MapQueryType, PrimitiveQueryType, StructQueryType, ComplexQueryType) in the core scan package to handle querying complex data types, enabling the system to read and process nested structures like arrays, maps, and structs during query execution.

core/src/main/java/org/apache/carbondata/core/scan/complextypes · high confidence

New read-committed scope implementations for non-transactional and managed tables

The core read-committed scope interface and two new implementations—LatestFilesReadCommittedScope for non-transactional tables and TableStatusReadCommittedScope for managed tables—have been added to the readcommitter package. These classes define how the engine identifies and retrieves committed index files and segment metadata during reads, enabling the system to correctly handle both transactional and non-transactional table scenarios.

core/src/main/java/org/apache/carbondata/core/readcommitter · high confidence

New result container classes for query scanning

Added BlockletScannedResult and RowBatch classes to the core scan result package. BlockletScannedResult serves as the primary container for scanned query results, managing dimension and measure column pages, reusable buffers, and delete deltas. RowBatch provides an iterator-based interface for accessing query result rows, supporting batch retrieval and version tracking.

core/src/main/java/org/apache/carbondata/core/scan/result · high confidence

New result provider classes for query scanning

Added FilterQueryScannedResult and NonFilterQueryScannedResult classes in the scan result implementation package. These new classes handle the retrieval of scanned data for filter and non-filter query scenarios respectively, providing specific logic for populating column vectors and valid row IDs based on the presence or absence of filter conditions.

core/src/main/java/org/apache/carbondata/core/scan/result/impl · high confidence

New scan processor classes for blocklet iteration and raw column chunk handling

The core module introduces three new classes in the scan processor package: BlockletIterator, DataBlockIterator, and RawBlockletColumnChunks. BlockletIterator provides iteration over data blocks, while DataBlockIterator manages the scanning and aggregation of blocklet results, supporting both filter and full scan modes. RawBlockletColumnChunks encapsulates the raw dimension and measure column chunks for a single blocklet, facilitating efficient data access during query execution.

core/src/main/java/org/apache/carbondata/core/scan/processor · high confidence

New schema metadata classes for table structure and evolution

Added new classes to the core metadata schema package: BucketingInfo, ColumnRangeInfo, PartitionInfo, SchemaEvolution, SchemaEvolutionEntry, SchemaReader, and SortColumnRangeInfo. These classes introduce support for tracking schema evolution (adding/removing columns), managing partition information, and handling sort column ranges, enabling the system to read and infer table schemas from storage.

core/src/main/java/org/apache/carbondata/core/metadata/schema · medium confidence

New serializable comparator implementations for sorting

The core module now includes a new \SerializableComparator\ interface and specific comparator classes for Boolean, BigDecimal, byte array, double, float, int, long, short, and String types. These new comparators are registered in the \Comparator\ utility class, enabling consistent, serializable comparison logic for sorting operations across the system.

core/src/main/java/org/apache/carbondata/core/util/comparator · high confidence

New streaming data ingestion components

The streaming module introduces a new set of classes to support streaming data ingestion, including a custom exception class, an output format, a record writer, a blocklet writer, a file index, and parsers for CSV and Spark SQL rows. These components enable the system to write streaming data into CarbonData segments, handling row parsing, compression, and metadata management for real-time or near-real-time data loads.

streaming · high confidence

Added new utility classes in the core module to manage data file footers and blocklet index information. The new \AbstractDataFileFooterConverter\ class provides logic for reading and converting index file data into \DataFileFooter\ objects, supporting both file path and byte array inputs. Additionally, \BitSetGroup\ was introduced to manage groups of bitsets for filter execution, and \BlockletIndexUtil\ was added to handle block metadata information mapping. These changes enhance the core's ability to process index and footer data during query execution.

core/src/main/java/org/apache/carbondata/core/util · high confidence

Register Carbon Data sources and test executors via Spark service files

Added META-INF/services files to register the Carbon Data source classes (CarbonSource, SparkCarbonFileFormat) and the test query executor (SparkTestQueryExecutor) with Spark's service loader mechanism. This ensures that the data source and test utilities are automatically discovered and registered by Spark, simplifying the format name and cleaning up unused code.

integration/spark/src/resources · high confidence

Segment-level min/max metadata caching for improved query pruning

The core module now maintains in-memory segment-level min/max statistics for each column, including sort column flags and drift indicators. By caching these metadata values, the system can perform more effective data pruning during query execution, which reduces the amount of data scanned and lowers driver memory usage for caching.

core/src/main/java/org/apache/carbondata/core/segmentmeta · medium confidence

Support for direct scan queries with inverted index and delete delta

The core module now includes a new set of classes in the directread package that enable direct scanning of columnar data using inverted indexes and delete deltas. This change introduces a factory and wrapper classes that handle the mapping of data from column pages to the actual vectors, accounting for deleted rows and null values. This allows the system to efficiently process queries that involve inverted indexes and delete deltas, improving performance and correctness for such operations.

core/src/main/java/org/apache/carbondata/core/scan/result/vector/impl/directread · high confidence

Support for querying stage files

The status manager now supports querying stage files. New classes including StageInput, StageInputCollector, and FileFormat have been added to the core status manager package. This enables the system to collect and create input splits from stage files, allowing data written to these temporary locations to be included in queries.

core/src/main/java/org/apache/carbondata/core/statusmanager · high confidence

Unsafe memory storage for dimension data chunks

Added new classes in the \core/src/main/java/org/apache/carbondata/core/datastore/chunk/store/impl/unsafe\ package to store dimension data in off-heap memory using the \Unsafe\ API. This includes an abstract base class \UnsafeAbstractDimensionDataChunkStore\ and concrete implementations for fixed-length (\UnsafeFixedLengthDimensionDataChunkStore\) and variable-length (\UnsafeVariableLengthDimensionDataChunkStore\, \UnsafeVariableIntLengthDimensionDataChunkStore\, \UnsafeVariableShortLengthDimensionDataChunkStore\) dimension data, enabling more efficient memory management for query execution.

core/src/main/java/org/apache/carbondata/core/datastore/chunk/store/impl/unsafe · high confidence

Behavioural changes

Added encoding type enumeration and validation for data file reading

A new \Encoding\ enum was added to the core module, defining supported encoding types such as DICTIONARY, DELTA, RLE, and others. This change introduces a \validateEncodingTypes\ method that checks whether the encodings present in data files are supported for reading in the current version, throwing an \UnsupportedOperationException\ if there is a mismatch or unsupported encoding is encountered.

core/src/main/java/org/apache/carbondata/core/metadata/encoder · high confidence

The core module now includes a standard Apache License 2.0 header in the CARBON\_CORELogResource.properties file, ensuring compliance with licensing requirements for all core resources.

core · high confidence

Centralized configuration and load options for data loading and storage

The constants for CarbonData configuration and load options have been consolidated into dedicated classes. \CarbonCommonConstants\ now holds system-level properties such as store location, blocklet size, compressor, and bad record settings. \CarbonLoadOptionConstants\ defines load-time options including bad record handling, date/timestamp formats, sort scope, and binary decoders. Additionally, \SortScopeOptions\ provides the \NO\_SORT\, \LOCAL\_SORT\, and \GLOBAL\_SORT\ enum values, while \CarbonV3DataFormatConstants\ and \CarbonVersionConstants\ manage V3 format specifics and version metadata. This change organizes these settings to make them easier to manage and configure.

core/src/main/java/org/apache/carbondata/core/constants · high confidence

Data type system refactored from enum to class-based hierarchy

The internal representation of data types has been refactored from a simple enum to a class-based hierarchy, with each type (such as String, Int, Decimal, and complex types like Array and Map) now having its own class extending a common \DataType\ base. This change improves type safety and allows for more complex type definitions, such as adding precision and scale to \DecimalType\ or element types to \ArrayType\. Additionally, a \DataTypeAdapter\ and \DataTypeDeserializer\ have been introduced to handle backward compatibility when deserializing table metadata from older versions, ensuring that string-based type representations in older metadata files are correctly converted to the new object-based format.

core/src/main/java/org/apache/carbondata/core/metadata/datatype · high confidence

Improved Hive integration with new utility classes for type conversion and table loading

Added new utility classes, DataTypeUtil and HiveCarbonUtil, to the Hive integration module. DataTypeUtil provides a comprehensive mapping of Hive SQL types to CarbonData internal types, supporting primitives, decimals, arrays, maps, and structs. HiveCarbonUtil introduces methods to construct CarbonLoadModel and CarbonTable instances from Hive metastore properties, handling both transactional and non-transactional table scenarios, including fixes for empty tables and complex type handling.

integration/hive/src/main/java/org/apache/carbondata/hive/util · high confidence

Introduce LoggerAction enum for bad record handling

A new LoggerAction enum is added to define how bad records are processed during data loading. Users can now choose between FORCE (convert to null), REDIRECT (write to raw CSV), IGNORE (skip writing), or FAIL (abort load) when bad records are encountered.

common/src/main/java/org/apache/carbondata/common/constants · medium confidence

Introduce TableOperation enum to define supported table modification actions

A new TableOperation enum has been added to the core features package, explicitly defining supported table modification actions including ALTER\_REN, ALTER\_DROP, ALTER\_ADD\_COLUMN, ALTER\_CHANGE\_DATATYPE, and others. This change provides a centralized, type-safe way to identify and manage table modification capabilities within the system.

core/src/main/java/org/apache/carbondata/core/features · high confidence

Introduce new schema classes for index metadata storage

The index store schema is refactored to use a new \CarbonRowSchema\ hierarchy, introducing \FixedCarbonRowSchema\, \VariableCarbonRowSchema\, and \StructCarbonRowSchema\ to represent fixed, variable-length, and structured data types respectively. A new \SchemaGenerator\ class is added to construct these schemas for blocks, blocklets, and task summaries, explicitly defining the layout for metadata such as min/max values, row counts, file paths, and blocklet counts. This change alters how index metadata is structured and persisted in the core module.

core/src/main/java/org/apache/carbondata/core/indexstore/schema · high confidence

Introduce vectorized read path for Presto integration

The Presto integration now supports a vectorized read path, enabling the engine to read data in columnar batches rather than row-by-row. This change introduces new classes such as CarbonVectorBatch, ColumnarVectorWrapperDirect, and PrestoCarbonVectorizedRecordReader to handle batched data retrieval, which improves query performance by reducing the overhead of individual row processing.

integration/presto/src/main · high confidence

Introduces new row data structures for the write path

The core datastore now uses a new \CarbonRow\ class and a \WriteStepRowUtil\ helper to manage row data during the write phase. \CarbonRow\ stores dictionary-encoded dimensions, no-dictionary/complex column data, and measure columns in a structured array, while \WriteStepRowUtil\ provides static methods to construct and extract these components, replacing previous ad-hoc handling of row data in the write step.

core/src/main/java/org/apache/carbondata/core/datastore/row · high confidence

Lazy loading of blocklets and pages for improved query performance

The scan scanner module now supports lazy loading of blocklets and pages. A new BlockletScanner interface defines the scanning contract, while LazyBlockletLoader and LazyPageLoader classes implement lazy loading logic. This allows the system to defer reading and decompressing column chunks and pages until they are actually accessed by the execution engine, which is particularly beneficial for filter queries with high cardinality columns.

core/src/main/java/org/apache/carbondata/core/scan/scanner · high confidence

Local dictionary storage now enforces configurable size limits

The local dictionary implementation has been refactored to support size-based fallback. A new \DictionaryStore\ interface and \MapBasedDictionaryStore\ implementation track both the count of unique keys and the total memory size. The system now enforces a configurable size threshold (defaulting to 16MB) alongside the existing count threshold, throwing a \DictionaryThresholdReachedException\ when either limit is exceeded, ensuring memory usage for local dictionaries remains bounded.

core/src/main/java/org/apache/carbondata/core/localdictionary/dictionaryholder · high confidence

New exception classes for deprecated features, metadata, streaming, and locking

The \common\ module introduces four new exception classes to handle specific error conditions: \DeprecatedFeatureException\ is thrown when using features deprecated in CarbonData 2.0, specifically global dictionary and custom partition; \MetadataProcessException\ handles failures during metadata processing; \NoSuchStreamException\ is raised when a stream is not found; and \TableStatusLockException\ is thrown when acquiring a table status lock fails or times out.

common/src/main/java/org/apache/carbondata/common/exceptions · high confidence

New exception types for file operations, concurrency, and configuration

Three new exception classes have been added to the core module to improve error handling: CarbonFileException for file-related failures, ConcurrentOperationException for table lock conflicts, and InvalidConfigurationException for invalid settings. These changes provide more specific error reporting for users encountering these conditions.

core/src/main/java/org/apache/carbondata/core/exception · high confidence

Refactored column metadata schema into dedicated classes

The column metadata schema in the core module has been refactored to use a new set of classes: CarbonColumn, CarbonDimension, CarbonMeasure, ColumnSchema, and related utilities. This change reorganizes how column definitions, encodings, and properties are stored and accessed, which may affect how table schemas are interpreted or serialized.

core/src/main/java/org/apache/carbondata/core/metadata/schema/table/column · high confidence

Refactored core expression classes for query scanning

The core module's scan expression package was restructured by introducing new base and concrete expression classes, including the abstract \Expression\ class and its subclasses \BinaryExpression\, \ColumnExpression\, \LiteralExpression\, \MatchExpression\, and \UnknownExpression\, alongside supporting classes like \ExpressionResult\ and \RangeExpressionEvaluator\. This refactoring organizes the query scanning logic by defining a clear hierarchy for handling different types of filter and column expressions during data scans.

core/src/main/java/org/apache/carbondata/core/scan/expression · high confidence

Refactored dimension chunk store architecture with new wrapper and factory classes

The core module introduces a new \ColumnPageWrapper\ class that acts as a wrapper for \ColumnPage\ and \CarbonDictionary\, handling data retrieval and vector filling logic for dimension columns. A new \DimensionChunkStoreFactory\ is added to manage the creation of various \DimensionDataChunkStore\ implementations (safe and unsafe variants) based on column size and configuration. The \DimensionDataChunkStore\ interface is also introduced to standardize operations like \putArray\, \fillVector\, and \getRow\ across different storage implementations. These changes restructure how dimension data is stored and accessed in memory, supporting both safe and unsafe memory access modes.

core/src/main/java/org/apache/carbondata/core/datastore/chunk/store · high confidence

Refactored index core classes to support distributed index pruning

The index module in the core library has been refactored to support distributed index pruning and improved query performance. New classes such as AbstractIndexJob, IndexChooser, IndexFilter, IndexInputFormat, and IndexInputSplit have been introduced to handle index job execution, filter resolution, and split generation. The IndexChooser now selects between coarse-grain (CG) and fine-grain (FG) indexes based on filter expressions, enabling more efficient data retrieval. Additionally, the IndexInputFormat and related classes facilitate the distribution of index pruning tasks across a cluster, improving scalability and performance for complex queries.

core/src/main/java/org/apache/carbondata/core/index · high confidence

Refactored query execution engine with new executor classes

The query execution logic in the core module has been refactored to introduce new executor classes: AbstractQueryExecutor, DetailQueryExecutor, VectorDetailQueryExecutor, and QueryExecutorProperties. This change restructures how queries are executed, separating the base logic into an abstract class and providing specific implementations for detail and vectorized queries. Users will experience this as an internal optimization that may improve query performance and maintainability, though the external API remains unchanged.

core/src/main/java/org/apache/carbondata/core/scan/executor/impl · high confidence

Refactored scan result collectors to support schema changes and restructured blocks

The scan result collection logic has been refactored to handle restructured blocks and newly added columns. A new abstract base class, AbstractScannedResultCollector, was introduced to share common logic for filling measure data. Specific collectors (DictionaryBased, RawBased, RestructureBased) now properly handle scenarios where dimensions or measures are newly added or missing in the current block, filling default values or skipping non-existing columns. This ensures that queries on tables with altered schemas (e.g., after ADD or ALTER operations) return correct results without throwing exceptions or returning nulls for new columns.

core/src/main/java/org/apache/carbondata/core/scan/collector/impl · high confidence

Test coverage

Add CI test suite for Spark examples; Added C++ SDK test harness and build configuration; Added StoreCreator test utility for Hadoop module; Added comprehensive test coverage for CarbonData bucketing; Added embedded Hive server utility for integration testing; Added empty test file for AbstractDictionaryCache; Added empty test file for CarbonDictionaryWriterImpl; Added integration tests for CarbonData Spark module; Added integration tests for Flink streaming and stage file management; Added integration tests for Hive and CarbonData interactions; Added integration tests for Lucene fine-grain and coarse-grain indexes; Added integration tests for Presto-Carbon data types and operations; Added integration tests for PrestoSQL; Added integration tests for Spark command execution and table deletion; Added integration tests for Spark data loading scenarios; Added integration tests for complex data types; Added integration tests for primitive data types and adaptive encoding; Added test cases for CarbonData table restructuring operations; Added test configuration for Carbon Data; Added test configuration for CarbonData logging; Added test coverage for CarbonData table schema validation and truncate operations; Added test coverage for Strings utility class; Added test coverage for materialized view query rewriting and plan modularization; Added test coverage for stream record reading and binary utility functions; Added test data file for Hadoop tests; Added test data files for Hive integration; Added test data for Spark integration; Added test resources for Presto integration; Added test suites for Carbon Data integration; Added test suites for Spark utility functions and Carbon commands; Added test suites for query execution and filtering; Added test suites for set command, numeric bad records, delete with subqueries, and profiler; Added test utility for creating Carbon Data stores in Presto integration tests; Added tests for Bloom filter coarse-grain index functionality; Added tests for Carbon Data table creation and cache management commands; Added tests for CarbonData table registration and DDL operations; Added tests for CarbonData vector reader restructuring operations; Added tests for CarbonSegmentUtil; Added tests for CarbonTablePath utility methods; Added tests for bad record handling, streaming table operations, and table status backup; Added tests for binary data handling and table creation in Spark integration; Added tests for binary data handling and table status recovery; Added tests for compressed file reading; Added tests for empty row and CSV handling edge cases; Added tests for integer and big decimal data types; Added tests for materialized view functionality; Added tests for secondary index data file merge operations; Added tests for secondary index file merge and compaction behavior; Added tests for secondary index operations; Added tests for spatial index validation and geo query operations; Added tests for time-series materialized view creation, loading, and query rollup; Added tests for vector reader functionality; Added unit tests for CarbonData logging implementation classes; Added unit tests for CarbonData version constants; Added unit tests for CarbonFile implementations; Added unit tests for CarbonIndexFileReader; Added unit tests for CarbonMetadata and DatabaseLocationProvider; Added unit tests for CarbonTable metadata schema components; Added unit tests for FixedLengthDimensionDataChunk; Added unit tests for HeapMemoryAllocator; Added unit tests for IndexServer utility functions; Added unit tests for LoadMetadataDetails; Added unit tests for LogServiceFactory; Added unit tests for MD key generator components; Added unit tests for ObjectSerializationUtil; Added unit tests for RLE codec and encoding factory selection; Added unit tests for ThriftWrapperSchemaConverterImpl; Added unit tests for complex type scanning; Added unit tests for conditional expression evaluation; Added unit tests for core datastore block metadata classes; Added unit tests for core metadata and locking components; Added unit tests for core scan expression components; Added unit tests for core utility classes; Added unit tests for filter executors; Added unit tests for filter expression processing and utility functions; Added unit tests for query and restructure utilities; Added unit tests for query statistics recorders; Added unit tests for the LRU cache implementation; Added unit tests for the blocklet index implementation; Added unit tests for the data storage filesystem implementation; Added unit tests for the direct dictionary key generator; Added unit tests for the local dictionary components; Added validation tests for Carbon properties; Expanded SDV cluster test coverage for complex data types, bad records, and table operations.

Dependencies

Introduce new build modules and Python agent configuration

The project now includes a new Python module, Agent\_module, with a pyproject.toml defining the 'apache-carbondata-agent' package (version 1.0.0) for self-contained .AI\_carbon archives. Additionally, a suite of new Maven modules has been added to the build system: carbondata-assembly, carbondata-common, carbondata-core, carbondata-hadoop, carbondata-format, carbondata-geo, carbondata-bloom, carbondata-index-examples, carbondata-lucene, and carbondata-secondary-index. These modules establish the build structure for the core, format, geospatial, and indexing components, while the assembly module configures the final JAR packaging.

(dependencies) · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

Score

  • CAI 41 → 52 (+10.7)
  • Rubric changed (rubric-2026.08.15 → rubric-2026.09.15) — scores are not directly comparable.

Lenses

  • Code Health 82 → 54 (-27.8)
  • Architecture 96 → 99 (+3.8)
  • Maturity 58 → 53 (-4.9)
  • Readiness 21 → 49 (+28.1)
  • Security 48 → 49 (+0.7)

Resolved (219)

  • Coverage not measured — test suite did not build
  • Dimension evaluation failed
  • Duplicated block (10 lines × 2) (core/src/main/java/org/apache/carbondata/core/datastore/chunk/store/impl/safe/AbstractNonDictionaryVectorFiller.java)
  • Duplicated block (10 lines × 2) (core/src/main/java/org/apache/carbondata/core/scan/result/BlockletScannedResult.java)
  • Duplicated block (10 lines × 2) (core/src/main/java/org/apache/carbondata/core/statusmanager/SegmentStatusManager.java)
  • Duplicated block (10 lines × 2) (core/src/main/java/org/apache/carbondata/core/util/ByteUtil.java)
  • Duplicated block (100 lines × 2) (core/src/main/java/org/apache/carbondata/core/datastore/page/LocalDictColumnPage.java)
  • Duplicated block (109 lines × 2) (core/src/main/java/org/apache/carbondata/core/scan/expression/RangeExpressionEvaluator.java)
  • Duplicated block (11 lines × 2) (core/src/main/java/org/apache/carbondata/core/scan/filter/executer/RangeValueFilterExecutorImpl.java)
  • Duplicated block (11 lines × 2) (core/src/main/java/org/apache/carbondata/core/util/CarbonUtil.java)
  • Duplicated block (11 lines × 3) (core/src/main/java/org/apache/carbondata/core/index/dev/expr/OrIndexExprWrapper.java)
  • Duplicated block (11 lines × 3) (core/src/main/java/org/apache/carbondata/core/util/path/CarbonTablePath.java)
  • Duplicated block (11 lines × 4) (core/src/main/java/org/apache/carbondata/core/datastore/compression/SnappyCompressor.java)
  • Duplicated block (111 lines × 2) (core/src/main/java/org/apache/carbondata/core/indexstore/schema/SchemaGenerator.java)
  • Duplicated block (115 lines × 2) (core/src/main/java/org/apache/carbondata/core/indexstore/schema/SchemaGenerator.java)
  • Duplicated block (1158 lines × 3) (core/src/main/java/org/apache/carbondata/core/util/CarbonProperties.java)
  • Duplicated block (13 lines × 2) (core/src/main/java/org/apache/carbondata/core/scan/filter/executer/RowLevelRangeLessThanEqualFilterExecutorImpl.java)
  • Duplicated block (13 lines × 2) (core/src/main/java/org/apache/carbondata/core/scan/filter/executer/RowLevelRangeLessThanFilterExecutorImpl.java)
  • Duplicated block (13 lines × 2) (core/src/main/java/org/apache/carbondata/core/util/ObjectSerializationUtil.java)
  • Duplicated block (130 lines × 2) (core/src/main/java/org/apache/carbondata/core/metadata/blocklet/BlockletInfo.java)
  • …and 199 more

New (1606)

  • 1.call (cognitive 33) (core/src/main/java/org/apache/carbondata/core/index/TableIndex.java)
  • 1.decodeAndFillVector (cognitive 24) (core/src/main/java/org/apache/carbondata/core/datastore/page/encoding/compress/DirectCompressCodec.java)
  • 1.decodeAndFillVector (cognitive 63) (core/src/main/java/org/apache/carbondata/core/datastore/page/encoding/adaptive/AdaptiveDeltaFloatingCodec.java)
  • 1.decodeAndFillVector (cognitive 63) (core/src/main/java/org/apache/carbondata/core/datastore/page/encoding/adaptive/AdaptiveFloatingCodec.java)
  • 1.decodeAndFillVector (cyclomatic 26) (core/src/main/java/org/apache/carbondata/core/datastore/page/encoding/adaptive/AdaptiveDeltaFloatingCodec.java)
  • 1.decodeAndFillVector (cyclomatic 26) (core/src/main/java/org/apache/carbondata/core/datastore/page/encoding/adaptive/AdaptiveFloatingCodec.java)
  • 1.fillPrimitiveType (cognitive 289) (core/src/main/java/org/apache/carbondata/core/datastore/page/encoding/compress/DirectCompressCodec.java)
  • 1.fillPrimitiveType (cyclomatic 71) (core/src/main/java/org/apache/carbondata/core/datastore/page/encoding/compress/DirectCompressCodec.java)
  • 1.fillVector (cognitive 215) (core/src/main/java/org/apache/carbondata/core/datastore/page/encoding/adaptive/AdaptiveIntegralCodec.java)
  • 1.fillVector (cognitive 306) (core/src/main/java/org/apache/carbondata/core/datastore/page/encoding/adaptive/AdaptiveDeltaIntegralCodec.java)
  • 1.fillVector (cyclomatic 57) (core/src/main/java/org/apache/carbondata/core/datastore/page/encoding/adaptive/AdaptiveIntegralCodec.java)
  • 1.fillVector (cyclomatic 68) (core/src/main/java/org/apache/carbondata/core/datastore/page/encoding/adaptive/AdaptiveDeltaIntegralCodec.java)
  • AbstractDataFileFooterConverter.getBlockletIndexForDataFileFooter (cognitive 28) (core/src/main/java/org/apache/carbondata/core/util/AbstractDataFileFooterConverter.java)
  • AbstractFactDataWriter.commitCurrentFile (cognitive 18) (processing/src/main/java/org/apache/carbondata/processing/store/writer/AbstractFactDataWriter.java)
  • AbstractQueryExecutor.getBlockExecutionInfoForBlock (cognitive 20) (core/src/main/java/org/apache/carbondata/core/scan/executor/impl/AbstractQueryExecutor.java)
  • AbstractQueryExecutor.getDataBlocks (cognitive 31) (core/src/main/java/org/apache/carbondata/core/scan/executor/impl/AbstractQueryExecutor.java)
  • AbstractQueryExecutor.updateColumns (cognitive 18) (core/src/main/java/org/apache/carbondata/core/scan/executor/impl/AbstractQueryExecutor.java)
  • AbstractScannedResultCollector.getMeasureData (cognitive 20) (core/src/main/java/org/apache/carbondata/core/scan/collector/impl/AbstractScannedResultCollector.java)
  • AdaptiveCodec.getPageBasedOnDataType (cognitive 43) (core/src/main/java/org/apache/carbondata/core/datastore/page/encoding/adaptive/AdaptiveCodec.java)
  • AdaptiveCodec.getPageBasedOnDataType (cyclomatic 16) (core/src/main/java/org/apache/carbondata/core/datastore/page/encoding/adaptive/AdaptiveCodec.java)
  • …and 1586 more

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

apache/carbondata was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 27 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit f86ac085ddbbd8381b0a5b65658731fe618eff4a — the exact code this score is about.
  • Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-d00c643c3f66.