Skip to content
CAI
Software that uses CAICheck a score

snowflakedb/spark-snowflake

52.4

Adequate · 20 September 2026

11.9k

lines of production code

Scala

primary language

1

measurement over time

CAI band scale
CAI lens gauges

What this system is

This system is the official Snowflake Spark Connector, a library that enables Apache Spark applications to read from and write to Snowflake data warehouses. It facilitates data transfer through direct JDBC queries and cloud storage stages (S3, Azure, GCP), supporting various data formats and handling authentication and security features like credential redaction. The connector is designed for cross-compatibility with Spark 3.5 and 4.x, including native support for the Spark VariantType, and is maintained with comprehensive unit and integration test suites to ensure stability and correctness.

How it got here

2014–2016 — Initial connector development and stabilization

8 changes.

This period established the Snowflake Spark Connector from scratch, implementing core infrastructure for reading and writing data via JDBC and cloud storage stages. It focused on building comprehensive integration and unit test suites to verify functionality across various data types and configurations. The work also involved migrating legacy code, securing credential handling, and ensuring compatibility with evolving Spark versions.

2018–2022 — Test coverage expansion

5 changes.

This period focused on significantly expanding test coverage for the Snowflake Spark Connector, particularly within the IO module and integration layers. The work involved adding unit tests for security and upload handling, as well as building new infrastructure for Snowflake-specific integration tests to validate stage operations and utility functions.

2026 — Spark 4.x and 3.5 test infrastructure

6 changes.

This period focused on establishing test infrastructure for Spark 4.x and 3.5 compatibility, including integration tests for native VariantType support and version-specific SQL API handling. It also involved adding placeholder directories for future test suites and resolving Mockito compatibility issues in Scala 2.12 environments to ensure stable test builds.

Features

Initial release of the Snowflake Spark Connector

This change introduces the initial version of the Snowflake Spark Connector, providing the core infrastructure for reading from and writing to Snowflake tables via Spark. It establishes the foundational components including parameter management (Parameters.scala), JDBC interaction (SnowflakeJDBCWrapper.scala), and the Spark Data Source API implementation (SnowflakeRelation.scala). The connector supports data transfer through both direct SELECT queries and cloud storage stages (S3/Azure), handling encryption, compression, and various data formats (CSV, JSON, Parquet). It also includes basic authentication support (OAuth, key pairs) and telemetry logging.

repository · high confidence

Initial repository setup with CI/CD, build, and documentation

The repository is initialized with essential infrastructure files, including a GitHub Actions workflow for integration and cluster testing, a deployment script for publishing to Maven Central, and an SBT-based build system supporting cross-compilation for Spark 3.5 and 4.0. The project adds a Codecov configuration for coverage reporting, a pre-commit hook for secret scanning, and a Scalastyle configuration to enforce code standards. Documentation is provided via a main README, developer notes, and test instructions, alongside example configuration files for Snowflake connections and a standard Apache 2.0 license.

(repo-wide) · high confidence

Spark 3.5 and 4.x compatibility with native Variant support and credential redaction

The connector now includes dedicated source implementations for Spark 3.5 and Spark 4.x. In Spark 4.x, the connector enables native support for the Spark VariantType, allowing VARIANT columns to be read and written directly via JSON staging. For both versions, the connector hardens security by registering sensitive option keys (such as private keys and OAuth secrets) with Spark's plan redaction, ensuring these credentials are masked in logs and the History Server UI.

src/main · high confidence

Behavioural changes

Integration test logging configuration updated

Added new log4j and log4j2 configuration files for the integration test suite to control log verbosity and output. The default test logging is now set to WARN to reduce log volume, while specific configurations are provided for console output, file logging, and compatibility with Spark 3.3, ensuring cleaner test runs and better separation of master/worker logs for security reviews.

src/it/resources · high confidence

Legacy spark-redshift codebase archived and migrated to Snowflake package structure

The legacy \spark-redshift\ codebase has been moved into the \legacy\ directory and restructured under the \net.snowflake.spark.snowflake\ package namespace. This change includes the addition of \Utils.SNOWFLAKE\_SOURCE\_NAME\ to enable compile-time validation of connection strings, and updates to configuration files and documentation (including a new \OLD-README.md\ and tutorial files) to reflect the transition from the Databricks \com.databricks.spark.redshift\ format to the Snowflake-specific implementation.

legacy · high confidence

Test coverage

Add Spark 3.5 integration test infrastructure; Added Mockito compatibility helper for test builds; Added Python test for Spark UDF integration; Added Snowflake-specific integration test infrastructure; Added Spark 4.x unit tests for Variant type handling; Added integration test suites for Snowflake Spark Connector; Added integration tests for SnowflakeSparkUtils; Added integration tests for internal stage file operations; Added placeholder directories for Spark 3.5, 4.0, and 4.1 test suites; Added tests for IO module security, GCP support, and upload abort handling; Added tests for SnowflakeFallbackCatalog error handling and V1Table fallback logic; Added unit tests for Snowflake Spark Connector; Integration tests for Spark 4.x compatibility and native VariantType support.

Dependencies

Upgrade to Spark 4.x support and Snowflake JDBC 4.x

The connector now supports Apache Spark 4.x (including 4.0 and 4.1) and Scala 2.13, requiring Java 17 for compilation and runtime. The Snowflake JDBC driver has been upgraded to version 4.0.2, and the connector version is set to 3.2.2. This update includes specific source directory handling for Spark 4.x, adjusted JVM options for Java 17 module opens, and updated dependency versions for Parquet Avro (1.15.2 for Spark 4.x, 1.13.1 for earlier versions) and Jackson (2.15.2).

(dependencies) · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

Baseline

  • First survey — no prior run to compare against. CAI 52.

Lenses

  • Code Health 87
  • Architecture 96
  • Maturity 40
  • Readiness 59
  • Security 53

Changes since last survey

  • 300 commits — 268 feature/other, 32 fixes

By area

  • src/main — 99 commits
  • src/it — 65 commits
  • (root) — 57 commits
  • (repo) — 52 commits
  • .github/workflows — 13 commits
  • src/test — 5 commits
  • .github/CODEOWNERS — 2 commits
  • .github/docker — 2 commits
  • ClusterTest/run_cluster_test.sh — 2 commits
  • ClusterTest/src — 2 commits
  • project/plugins.sbt — 1 commit

Notable commits

  • fix: Fix Merge Gate (#606)
  • fix: Fix Parquet Issues (#591)
  • fix: Fix cn aws url (#596)
  • fix: Fixed Exception name format and added comments
  • fix: Fixed test spark URL
  • fix: NOSNOW: Revert Spark4 work for the upcoming 3.1.9 release (#652)
  • fix: PRODSEC-1721 fix synk token privilege issue for public repo(#434)
  • fix: Revert "NOSNOW: Revert Spark4 work for the upcoming 3.1.9 release (#652)" (#655)
  • fix: Revert PR 292 for a cleaner release code base. (#314)
  • fix: SNOW-1359291 Fix Proxy Protocol Issue (#555)
  • fix: SNOW-170311 Send specific unsupported operation in generateQueries. Before the enhancement, if generateQueries() meets unsupported operation, it returns None, so a pushdown failure message for "NoSuchElementException" is sent. It is unclear. With this fix, specific info about the unsupported operation is sent.
  • fix: SNOW-170484 Fix runQuery() to return a ResultSet W/O connection
  • fix: SNOW-184015 Fix a regression for writing with external stage
  • fix: SNOW-188235 Fixed test and throwable catch
  • fix: SNOW-199334 Revert SNOW-187770 and bring back 4 deleted public functions
  • fix: SNOW-199722 Revert 2 recent manual reverts. (#294)
  • fix: SNOW-200050 Revert recent pushdown and functionality changes (#295)
  • fix: SNOW-200846 Fix NPE when multiple threads send OOB telemetry message. (#296)
  • fix: SNOW-206588: Fix translations of null literals (#302)
  • fix: SNOW-2207811: Fix tmp aws credentials for external S3 bucket (#622)
  • …and 280 more

Architecture

  • 0 containers · 1 bounded contexts · 0 dependency edges (baseline)

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

snowflakedb/spark-snowflake was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 20 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit ac1a75a16b89c05f9617ede3afca66289c7f7534 — the exact code this score is about.
  • Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-b51f968c9b10.