pingcap/tispark
49.0
Weak · 28 September 2026
35k
lines of production code
Java
with Scala
2
measurements over time
What this system is
TiSpark is a connector that enables Apache Spark to read from and write to TiDB and TiKV clusters. It provides a Spark SQL interface for querying data with features like columnar execution, query pushdown, and authentication, while also supporting batch data ingestion and deletion. The system includes a dedicated Java client for low-level TiKV communication and manages complex integration details such as TLS security, cache consistency, and version compatibility across multiple Spark releases.
How it got here
2017 — TiSpark v2 API and TiKV Java client introduction
20 changes.
This period established the foundational structure of the TiSpark project and introduced a new TiKV Java client library built with Bazel. The core module was restructured to support the Spark v2 data source API and multiple Spark versions, replacing legacy RDD-based components with a new columnar execution engine. Significant effort was also dedicated to implementing cache invalidation mechanisms and comprehensive unit test coverage for the new client and core modules.
2018–2019 — statistics, columnar execution, and test expansion
25 changes.
This period focused on enhancing query optimization through a new statistics module and table size estimation, while introducing columnar execution support via TiColumnVectorAdapter. Significant effort was also dedicated to expanding test coverage for TiKV client codecs, batch write operations, and SQL expression parsing, alongside adding configuration templates for the broader TiDB ecosystem.
2020–2023 — v2 DataSource API and batch write expansion
30 changes.
This period focused on implementing the Spark 3.x DataSource V2 API and a comprehensive TiBatchWrite engine, enabling robust data insertion, updates, and deletion via the standard Spark interface. The work also included extending compatibility to Spark 3.1 and 3.2, introducing TiCatalog integration, and adding critical features like authentication, authorization, and GC safe point management.
Features
Add Spark 3.1 wrapper support
This change introduces a new Spark 3.1 wrapper module, adding the necessary adapter files to support TiSpark on Spark 3.1. The implementation includes \SparkWrapper\ for version-specific utility functions, \TiBasicExpression\ and \TiBasicLogicalPlan\ for translating SQL expressions and logical plans, \TiDBWriteBuilder\ for handling data writes, and \TiStrategy\ along with \TiAggregationProjectionV2\ for query planning and optimization. These files adapt the core TiSpark logic to the specific APIs and structures required by the Spark 3.1 catalyst.
spark-wrapper/spark-3.1 · high confidence
Add Spark 3.2 compatibility wrapper
Introduces a new Spark 3.2 wrapper module containing the core integration components required to run TiSpark on Spark 3.2. This includes the \SparkWrapper\ utility for version-specific API handling, \TiBasicExpression\ for translating Spark SQL expressions to TiKV, \TiBasicLogicalPlan\ for authorization verification, \TiDBWriteBuilder\ for data insertion support, and \TiStrategy\ along with \TiAggregationProjectionV2\ for query planning and aggregation push-down optimizations.
spark-wrapper/spark-3.2, spark-wrapper/spark-3.4 · high confidence
Add SpecialSum aggregate expression with configurable initial value
A new \SpecialSum\ aggregate expression has been introduced in the SQL catalyst layer, allowing for a custom initial value and return type during aggregation. This enables specific optimization paths, such as promoting \LongType\ inputs to \DecimalType\ for push-down scenarios or handling count-based sums with a zero initial value, differing from the standard \Sum\ behavior.
core/src/main/scala/org/apache/spark/sql/catalyst/expressions/aggregate · high confidence
Add columnar execution support for Spark via TiColumnVectorAdapter
This change introduces the \TiColumnVectorAdapter\ and \TiColumnarBatchHelper\ classes in the TiKV core module to enable columnar data execution in Spark. The adapter wraps internal TiKV column vectors to implement the Spark \ColumnVector\ interface, allowing Spark to read data in a columnar format for performance gains. It currently supports primitive types (boolean, byte, short, int, long, float, double), decimals, strings, and binary data, while explicitly throwing \UnsupportedOperationException\ for complex types like arrays, maps, and child vectors.
core/src/main/java/com/pingcap/tikv/columnar · high confidence
Add metadata dump script and build-time version generation
The core/scripts directory now includes a new SQL dump file (DumpHiveMetastore.sql) for exporting Hive Metastore metadata, a shell script (fetch-test-data.sh) to initialize Git submodules for test data, and a new version generation script (version.sh) that dynamically creates a Scala source file containing the TiSpark release version, Git commit hash, branch, and build timestamp.
core/scripts · high confidence
Add statistics information module
Introduces a new statistics module that enables TiSpark to load and cache table statistics (such as row counts, histograms, and CMSketches) from TiDB. This feature allows the query optimizer to leverage statistical data for improved index selection and supports broadcast join optimizations by maintaining a local cache of table metadata and column/index histogram information.
core/src/main/scala/com/pingcap/tispark/statistics · high confidence
Add type mapping for TiKV data types to Spark SQL types
A new TypeMapping class has been added to the TiKV core module to handle the conversion of TiKV internal data types into Apache Spark SQL data types. This mapping ensures that TiDB chunk data formats are correctly interpreted within Spark, supporting columnar execution and insert operations via SQL. Specifically, it handles conversions for dates, timestamps, decimals (with precision limits), strings, floats, doubles, binary data, and integers, including special handling for unsigned long integers to prevent overflow by mapping them to decimal types.
core/src/main/java/com/pingcap/tikv/datatype · high confidence
Added TiDB data source registration and debug logging configuration
The core module now registers the TiDB table provider (com.pingcap.tispark.v2.TiDBTableProvider) via the Spark DataSourceRegister service file, enabling automatic discovery of TiDB as a data source. Additionally, a new log4j properties template has been introduced to configure logging levels, specifically setting the com.pingcap.tispark logger to DEBUG to facilitate local debugging and troubleshooting.
core/src/main/resources · high confidence
Added TiDB-based authentication and authorization checks
The TiSpark connector now supports authentication and authorization via TiDB. A new \TiAuthorization\ component connects to TiDB to fetch and periodically refresh user privileges (global, database, and table levels) and enforces these permissions during data source operations, throwing SQL exceptions if a user lacks the required privileges for a specific command.
core/src/main/scala/com/pingcap/tispark/auth · high confidence
Added TiSpark catalyst rules for authorization and statistics
New logical plan rules have been introduced to handle TiSpark-specific operations within the query optimizer. The TiAuthorizationRule enforces access control by verifying permissions for SELECT, DELETE, and catalog/namespace changes when authentication is enabled, while the TiStatisticsRule automatically loads and applies table statistics (row counts and size) to DataSourceV2 relations to improve query planning accuracy.
core/src/main/scala/org/apache/spark/sql/catalyst/rule · high confidence
Added automated Scala and Java formatting tools and configuration
Developers can now enforce consistent code style in CI by using new formatting scripts and configuration files. A Scalafmt configuration (version 2.5.3) and a \scalafmt\ script are provided for Scala code, while a Google Java Style XML configuration and a \javafmt\ script (using the fmt-maven-plugin) are included for Java. These tools allow for automatic formatting via IDE plugins or command-line execution before committing.
dev · high confidence
Added cache invalidation listeners for TiKV region errors
New \CacheInvalidateListener\ and \PDCacheInvalidateListener\ components have been added to the TiSpark core to handle cache invalidation events from TiKV. The \CacheInvalidateListener\ registers a Spark listener that accumulates invalidation events on executor nodes, while the \PDCacheInvalidateListener\ processes these accumulated events on the driver side when a job ends, forwarding them to a handler to ensure local cache consistency with the TiKV region manager.
core/src/main/scala/com/pingcap/tispark/listener · high confidence
Added usage telemetry collection for TiSpark
TiSpark now collects and reports anonymous usage information to help improve the product. When enabled, the system gathers a unique track ID, hardware details (OS, CPU, memory, disks), TiSpark and TiDB versions, and current configuration settings, then sends this data via HTTP POST to the PingCAP telemetry server. This feature is controlled by the \spark.tispark.telemetry.enable\ configuration option and runs asynchronously at startup.
core/src/main/scala/com/pingcap/tispark/telemetry · high confidence
Initial configuration templates for TiDB, TiKV, PD, and TiFlash
Added default configuration files (templates) for the core TiDB ecosystem components: TiDB, TiKV, Placement Driver (PD), and TiFlash. These include standard configurations as well as specific variants for version 4.0 compatibility, TLS security settings, and daily CI test environments. The TiDB templates enable features such as table locks and alter primary key by default, while the PD and TiKV templates configure replication, storage paths, and logging. Additionally, a Hive metastore configuration template is provided to connect to a TiDB backend.
config · high confidence
Initial release of the TiKV Java client with Bazel build support
This change introduces the TiKV Java client library, providing a Java-based interface to communicate with TiDB and TiKV clusters via gRPC. The client supports data read/write operations, region discovery through the Placement Driver, and coprocessor push-down calculations. The project is now built using Bazel, as evidenced by the new BUILD, WORKSPACE, and Makefile configurations, which also handle dependency shading for Netty to prevent conflicts. The release includes the Apache 2.0 license and documentation outlining usage with TiSpark and direct client integration.
tikv-client · high confidence
Initial repository structure and documentation for TiSpark
This change establishes the foundational structure of the TiSpark project by adding the initial commit of key configuration and documentation files. It introduces the Apache 2.0 license, project ownership (OWNERS), and a comprehensive README that defines TiSpark as a Spark layer for TiDB/TiKV/TiFlash, detailing feature support, limitations (such as lack of view support), and the requirement to manually provide the mysql-connector-java dependency. The commit also adds CI configuration (Travis, .lift.toml), Docker Compose files for local TiDB cluster development (including TLS support), and a detailed changelog documenting releases from v2.1.3 through v3.2.0, covering features like telemetry, TLS, and Spark 3.3 support.
(repo-wide) · high confidence
Introduce TiCatalog for Spark SQL integration
Added a new TiCatalog implementation that enables TiDB tables to be managed and queried via the Spark SQL catalog interface. This change introduces support for namespace listing, table loading, and authorization checks (including visibility filtering) within the Catalyst catalog layer, allowing users to interact with TiDB data sources using standard Spark SQL commands.
core/src/main/scala/org/apache/spark/sql/catalyst/catalog · high confidence
Introduce cache invalidation event handler for TiSpark driver
Added a new \CacheInvalidateEventHandler\ class in the handler package that processes \CacheInvalidateEvent\ messages from the TiKV client. This handler updates the local region and store caches in the Spark driver by delegating to the \RegionManager\ for specific cache types (region/store, leader, and request failures), ensuring that stale cache data is properly invalidated during runtime operations.
core/src/main/scala/com/pingcap/tispark/handler · high confidence
Introduce columnar execution and type conversion for TiSpark
This change introduces new execution nodes and utilities to support columnar data processing and type mapping between Spark SQL and TiDB. The new \CoprocessorRDD\ and \ColumnarRegionTaskExec\ classes enable columnar batch execution, allowing TiSpark to fetch data in a columnar format and handle index scan downgrades more efficiently. Additionally, the \TiConverter\ object provides explicit type mapping logic to convert Spark SQL data types to TiDB types and vice versa, ensuring correct data handling during batch writes and reads.
core/src/main/scala/org/apache/spark/sql/execution · high confidence
Introduce table size estimation for query optimization
Added the TableSizeEstimator trait and its DefaultTableSizeEstimator implementation to calculate estimated row sizes, row counts, and total table sizes. This logic, which relies on StatisticsManager for row counts and TiTableInfo for row width, is used to determine table size in bytes, directly influencing decisions such as whether to broadcast a table during query execution.
core/src/main/scala/com/pingcap/tispark/statistics/estimate · high confidence
Introduction of Service GC Safe Point management
TiSpark now actively manages the GC safe point for its service to prevent premature garbage collection of data required by long-running transactions. A new \ServiceSafePoint\ component registers a service-specific GC safe point with the Placement Driver (PD) and periodically updates it based on the minimum transaction start timestamp. This mechanism ensures that TiKV retains necessary historical versions for the duration of TiSpark operations, configurable via \spark.tispark.gc\_max\_wait\_time\, and automatically stops registration when the service terminates.
core/src/main/scala/com/pingcap/tispark/safepoint · high confidence
Introduction of TiContext and TiExtensions for TiSpark integration
This change introduces the core integration components \TiContext\ and \TiExtensions\ in the Spark SQL module. \TiContext\ serves as the central bridge to the TiKV cluster, managing the \ClientSession\, initializing statistics and cache invalidation listeners, and providing a \DebugTool\ for balancing region leaders. \TiExtensions\ registers the necessary Spark SQL catalyst rules (parser, resolution, planner, and check rules) to enable TiSpark functionality within a \SparkSession\, including support for authentication, statistics, and telemetry (which is disabled by default).
core/src/main/scala/org/apache/spark/sql · high confidence
Introduction of a new Java-based TiKV client library
The \tikv-client/src/main\ area now contains a new Java client library (\com.pingcap.tikv\) that provides core connectivity and data access capabilities for TiKV. This includes a session management layer (\ClientSession\) for handling connections and catalog caching, a \Snapshot\ class for point reads and table/index scans, and a \TiConfiguration\ class for managing cluster endpoints, timeouts, and TLS settings. The library also introduces a \ReplicaReadPolicy\ to allow users to configure read routing based on store roles (leader, follower, learner) and labels, and a \TiFlashClient\ to support health checks against TiFlash stores via the MPP protocol.
tikv-client/src/main · high confidence
Introduction of the v2 DataSource API implementation for TiDB tables
This change introduces the v2 implementation of the TiDB data source, adding new classes such as TiDBTable, TiDBTableProvider, and TiDBTableScan to support the Spark 3.x Table API. The new TiDBTable class implements SupportsRead, SupportsWrite, and SupportsDelete, enabling batch reads and writes as well as filtered row deletions via the connector catalog interface. The TiDBTableProvider registers the data source under the short name 'tidb' and handles table resolution, while the sink components (TiDBBatchWrite, TiDBDataWrite, TiDBDataWriterFactory) provide the underlying write infrastructure. This v2 path is distinct from the existing v1 implementation, which remains the default for DataFrame writes as noted in the provider code.
core/src/main/scala/com/pingcap/tispark/v2 · high confidence
New TiDB data source write implementation
TiSpark now supports writing data to TiDB via the standard Spark DataSource API using \SaveMode.Append\. This change introduces a new batch write engine (\TiBatchWrite\) that handles inserts and updates, including support for clustered indexes, auto-random columns, and region splitting. It also adds a new \TiDBDelete\ component for deleting rows via the data source, configurable options for TTL and deduplication, and authorization checks for write operations.
core/src/main/scala/com/pingcap/tispark/write · high confidence
New utility classes for TLS, reflection, and system info
Added new utility classes in the utils package: HttpClientUtil for handling HTTP/HTTPS requests with TLS support (including JKS and PKCS\#8 certificate formats), ReflectionUtil for managing Spark version-specific reflection calls, SystemInfoUtil for gathering OS and hardware information, and updated TiUtil to include TLS configuration handling. These changes enhance TiSpark's ability to securely connect to TiKV/TiDB and adapt to different Spark versions.
core/src/main/scala/com/pingcap/tispark/utils · high confidence
Support for URI-based host mapping in TiKV client
Users can now configure host mapping via a string-based URI format (e.g., 'host1:port1;host2:port2') to override network addresses for TiKV PD clients. This new UriHostMapping implementation parses the provided string into a concurrent map and applies the mapping when resolving URIs, enabling scenarios where internal TiKV addresses need to be remapped for external access or testing.
core/src/main/java/com/pingcap/tikv/hostmap · high confidence
TiDB snapshot-aware SQL parsing via TiParser
The Spark connector now includes a TiParser and TiParserFactory in the catalyst parser package. When SQL is parsed, the TiParser refreshes the TiDB metadata catalog before delegating to the underlying parser, and it supports switching to a snapshot catalog if a TiDB snapshot timestamp is configured in the session. This enables users to query TiDB tables with consistent, point-in-time views without manual catalog management.
core/src/main/scala/org/apache/spark/sql/catalyst/parser · high confidence
Removals
Removal of TiStrategy query planning component
The TiStrategy file, which contained the logic for translating Spark SQL logical plans into TiDB coprocessor requests (including aggregation, filter, and projection pushdowns), has been deleted. This removes the specific query optimization and execution strategy that allowed Spark to offload these operations to TiKV.
src/main/scala/org · high confidence
Removal of legacy TiSpark expression and data source components
The legacy \BasicExpression\, \TiDBRelation\, \TiRDD\, and \TiUtils\ files have been removed from the \com.pingcap.tispark\ package. This deletion eliminates the previous implementation of Spark SQL expression translation, table relation handling, and RDD-based data retrieval, indicating a structural shift in how TiSpark processes queries and accesses TiDB data.
src/main/scala/com · high confidence
Behavioural changes
Automate TiPB protocol buffer checkout in build scripts
The build process now includes a dedicated script (tikv-client/scripts/proto.sh) that automatically clones or updates the tipb repository to a specific, fixed commit hash (29e23c6). This ensures that the protocol buffer definitions used during compilation are consistent and pinned to a known version, rather than relying on the latest master branch or manual updates.
tikv-client/scripts · high confidence
Introduce TiHandleRDD, TiRowRDD, and TiRDD base classes for data retrieval
This change introduces three new core components in the TiSpark connector: an abstract \TiRDD\ base class that handles partitioning logic and timezone validation, a \TiHandleRDD\ for retrieving primary key handles from TiKV, and a \TiRowRDD\ for reading table data in columnar batches. These classes form the foundation for how Spark queries TiDB/TiKV data, replacing or refactoring previous internal mechanisms to support more efficient handle-based and chunk-based reads.
core/src/main/scala/org/apache/spark/sql/tispark · high confidence
Optimized aggregation pushdown to TiKV via Average rewrite
The query planner now rewrites Average aggregate expressions into Divide(Sum / Count) forms before pushing them down to TiKV. This change, implemented in the new TiAggregation and TiAggregationImpl components within the catalyst planner, ensures that Sum and Count operations are pushed to the storage layer while handling necessary type promotions (e.g., Long to Decimal) and casts to prevent overflow and maintain result accuracy. A new TiStrategyFactory is also introduced to manage strategy instantiation across Spark versions.
core/src/main/scala/org/apache/spark/sql/catalyst/planner · high confidence
Refactored expression handling with new BasicExpression and TiExprUtils components
The codebase introduces two new utility objects, BasicExpression and TiExprUtils, to manage the translation of Spark Catalyst expressions to TiKV expressions. BasicExpression provides helper methods for literal conversion, type checking, and extracting underlying Ti expressions, while TiExprUtils handles the transformation of grouping, aggregation (including Count, Sum, Min, Max, First), filtering, and sorting expressions into TiDB DAG request formats. This refactoring centralizes expression logic and supports push-down optimizations for aggregate functions.
core/src/main/scala/org/apache/spark/sql/catalyst/expressions · high confidence
TiSpark core module restructured to v2 API with Spark 3.4 support
The TiSpark core module has been reorganized into a new directory structure, introducing a v2 data source implementation via the \DefaultSource\ class that delegates to \TiDBTableProvider\. This change includes a new \MetaManager\ for catalog management, a \TiConfigConst\ object defining configuration keys, and a \TiSparkInfo\ utility that explicitly lists supported Spark versions (3.0 through 3.4) and enforces version compatibility checks. The internal data access layer has been updated, with \TiPartition\ now accepting a sequence of region tasks and an application ID, and the legacy \TiOptions\ class replaced by a placeholder \TiDBRelation\ to support the new v1 path build.
core/src/main/scala/com/pingcap/tispark · high confidence
Updated TLS certificates and keys in PEM format
The configuration files in config/cert/pem have been replaced with a new set of certificates and private keys for the root CA, client, TiDB, TiKV, and PD components. This update ensures that the TLS credentials used for secure communication between these services are current and valid, addressing the expiration of the previous certificates.
config/cert · high confidence
Test coverage
Added LineItem batch write regression tests; Added TLS configuration templates and updated test resource structure; Added TLS integration tests for TiKV client, JDBC, and TiSpark; Added TPC-DS benchmark SQL queries for integration testing; Added TPC-H SQL test resources; Added TiFlash integration test suite; Added concurrency and SQL insert test suites; Added concurrency test for ConcreteBackOffer; Added overflow validation tests for batch write data types; Added random database test infrastructure with clustered index support; Added test coverage for multi-table batch write functionality; Added test coverage for partition write operations and new collation support; Added test fixtures for UTF-8 prefix index scenarios; Added test for wide-table query handling; Added test infrastructure and unit tests for TiConfiguration; Added test logging configuration; Added test resources for resolve lock functionality; Added test suites for TiBatchWrite edge cases and TPC-H data integrity; Added test suites for TiSpark DELETE operations; Added test suites for TiSpark DataSource batch write capabilities; Added test suites for TiSpark data type conversions; Added tests for TiSpark authentication and authorization logic; Added tests for TiSpark collation support; Added tests for TiSpark host mapping configuration; Added tests for batch write TTL and lock timeout scenarios; Added tests for batch write data types and decimal reading logic; Added tests for telemetry usage reporting; Added unit tests for RowIDAllocator shard row ID logic; Added unit tests for TiKV client codec and collation components; Added unit tests for TiKV client type converters and data types; Added unit tests for TiKV meta information parsing and serialization; Added unit tests for TiKV predicate matching and scan analysis; Added unit tests for expression normalization, chunk iteration, and schema inference; Added unit tests for key handling classes; Added unit tests for partition expression rewriting and range partition locator; Added unit tests for the TiKV expression parser; Expanded test coverage for TiSpark SQL features and data types.
Dependencies
Maven build structure reorganized with new module POMs and parent upgrade
The project's Maven build structure has been reorganized: the root pom.xml now inherits from the Apache parent POM (version 18) and is renamed to tispark-parent, while new dedicated POMs are introduced for the assembly, core, core-test, db-random-test, and tikv-client modules. Additionally, Spark wrapper modules are added for versions 3.0 through 3.4 to support multiple Spark releases. The core module explicitly includes oshi-core 6.1.6 and sets Netty to 4.1.75.Final, while the core-test module pins Log4j to 2.17.2 and Guava to 29.0-android.
(dependencies) · high confidence
Housekeeping
Added placeholder file in core-test source
A new empty file named KEEPME was added to the core-test/src directory. This appears to be a placeholder or marker file with no functional impact on the product behavior.
core-test · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Score
- CAI 49 → 49 (-0.2)
- Rubric changed (rubric-2026.09.8 → rubric-2026.09.16) — scores are not directly comparable.
Lenses
- Code Health 87 → 86 (-0.7)
- Architecture 100 → 92 (-8.2)
- Maturity 69 → 69 (+0.0)
- Readiness 45 → 46 (+1.0)
- Security 35 → 35 (+0.0)
New (54)
- Documentation: no installation or build instructions (README.md)
- Duplicated block (10 lines × 4) (spark-wrapper/spark-3.1/src/main/scala/org/apache/spark/sql/extensions/TiStrategy.scala)
- Duplicated block (11 lines × 3) (spark-wrapper/spark-3.0/src/main/scala/org/apache/spark/sql/catalyst/plans/logical/TiBasicLogicalPlan.scala)
- Duplicated block (12 lines × 4) (spark-wrapper/spark-3.1/src/main/scala/org/apache/spark/sql/catalyst/expressions/TiBasicExpression.scala)
- Duplicated block (12 lines × 5) (spark-wrapper/spark-3.0/src/main/scala/org/apache/spark/sql/catalyst/expressions/TiBasicExpression.scala)
- Duplicated block (12 lines × 5) (spark-wrapper/spark-3.0/src/main/scala/org/apache/spark/sql/extensions/TiStrategy.scala)
- Duplicated block (13 lines × 2) (spark-wrapper/spark-3.1/src/main/scala/org/apache/spark/sql/extensions/TiAggregationProjectionV2.scala)
- Duplicated block (13 lines × 5) (spark-wrapper/spark-3.0/src/main/scala/org/apache/spark/sql/extensions/TiStrategy.scala)
- Duplicated block (14 lines × 2) (spark-wrapper/spark-3.1/src/main/scala/org/apache/spark/sql/extensions/TiStrategy.scala)
- Duplicated block (14–15 lines × 5) (spark-wrapper/spark-3.0/src/main/scala/org/apache/spark/sql/extensions/TiStrategy.scala)
- Duplicated block (15 lines × 5) (spark-wrapper/spark-3.0/src/main/scala/org/apache/spark/sql/connector/write/TiDBWriteBuilder.scala)
- Duplicated block (15 lines × 5) (spark-wrapper/spark-3.0/src/main/scala/org/apache/spark/sql/extensions/TiStrategy.scala)
- Duplicated block (15–16 lines × 2) (spark-wrapper/spark-3.3/src/main/scala/org/apache/spark/sql/extensions/TiStrategy.scala)
- Duplicated block (16 lines × 5) (spark-wrapper/spark-3.0/src/main/scala/org/apache/spark/sql/extensions/TiStrategy.scala)
- Duplicated block (18 lines × 3) (spark-wrapper/spark-3.2/src/main/scala/org/apache/spark/sql/connector/write/TiDBWriteBuilder.scala)
- Duplicated block (20–21 lines × 5) (spark-wrapper/spark-3.0/src/main/scala/org/apache/spark/sql/extensions/TiStrategy.scala)
- Duplicated block (22 lines × 2) (core/src/main/scala/com/pingcap/tispark/write/TiBatchWriteTable.scala)
- Duplicated block (24 lines × 5) (spark-wrapper/spark-3.0/src/main/scala/org/apache/spark/sql/extensions/TiStrategy.scala)
- Duplicated block (25 lines × 4) (spark-wrapper/spark-3.1/src/main/scala/com/pingcap/tispark/SparkWrapper.scala)
- Duplicated block (25 lines × 5) (spark-wrapper/spark-3.0/src/main/scala/org/apache/spark/sql/extensions/TiStrategy.scala)
- …and 34 more
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
pingcap/tispark was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 28 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 4395cd4910bd1cb858c57cb478d65b4f631843fa — the exact code this score is about.
- Scored under rubric-2026.09.16 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-2d9048c36d26.