Skip to content
CAI
Software that uses CAICheck a score

memsql/singlestore-spark-connector

61.8

Adequate · 20 September 2026

6.5k

lines of production code

Scala

primary language

1

measurement over time

CAI band scale
CAI lens gauges

What this system is

This system is the SingleStore Spark Connector, a library that enables data ingestion and query execution between Apache Spark and SingleStore databases. It supports Spark versions 3.1 through 3.5 and provides DataFrame-based APIs for batch inserts and data loading, having replaced legacy RDD implementations. The project includes tooling for automated cluster provisioning with SSL and JWT authentication, as well as CI workflows for publishing AWS Glue containers.

Features

Added support for Spark 3.1 through 3.4

The connector now supports Spark versions 3.1, 3.2, 3.3, and 3.4. This is implemented by adding version-specific source directories (e.g., \scala-sparkv3.1\, \scala-sparkv3.2\, etc.) containing version-adapted expression generators, aggregate extractors, and utility classes to handle API differences across these Spark releases. Additionally, the \DataSourceRegister\ service file has been updated to register both \com.singlestore.spark.DefaultSource\ and \com.memsql.spark.DefaultSource\ to maintain backward compatibility with the previous data source name.

src/main · high confidence

Automated AWS Glue container publishing workflow

The CI pipeline now includes a dedicated workflow to build and publish an AWS Glue container for the SingleStore Spark Connector. This change introduces a Dockerfile based on Amazon Linux that packages the connector JARs and a configuration file defining the connector metadata (version 5.0.2, Spark 3.5.0) for marketplace distribution.

ci · high confidence

Automated SingleStore cluster setup with SSL and JWT authentication

A new \scripts/setup-cluster.sh\ script has been added to automate the provisioning of a SingleStore development environment. This script handles the generation of SSL certificates (CA, server, and client) and keystores, launches the SingleStore container, and configures the cluster to enforce SSL connections for a dedicated \root-ssl\ user. It also sets up JWT authentication by configuring the auth config file and creating a \test\_jwt\_user\ with JWT-based login capabilities, ensuring the cluster is ready for secure integration testing.

scripts · high confidence

Initial public release of the SingleStore Spark Connector

This entry marks the initial public commit of the SingleStore Spark Connector, establishing the project's foundational structure and documentation. The repository now includes the Apache 2.0 license, a comprehensive README detailing configuration for SingleStoreDB Cloud and On-Premises deployments, and a changelog tracking versions up to 5.0.2. Build and development tooling is introduced via \.arcconfig\ for internal review, \.gitignore\, \.java-version\ (Java 1.8), and \.scalafmt.conf\, while legacy build artifacts like the old Makefile and \memsqlrdd\_app.sbt\ are removed to streamline the project setup.

(repo-wide) · high confidence

Removals

Removal of legacy MemSQLRDD class

The legacy \MemSQLRDD\ class, which previously handled direct SQL query execution and partitioning logic for MemSQL, has been removed from the connector library. This change eliminates the custom RDD implementation that managed its own JDBC connections and EXPLAIN-based partition discovery, indicating a shift toward newer connector abstractions (such as DataFrames or the DataSource API) for data ingestion.

connectorLib/src/main/scala/com/memsql/spark/connector/rdd · high confidence

Removal of legacy RDD-based save functionality

The connector has removed the legacy \RDDFunctions\ class and its associated implicit conversion in \package.scala\, which previously allowed saving RDDs of arrays directly to MemSQL via JDBC. Additionally, the internal \NextIterator\ utility class, copied from the Spark codebase, has been deleted. This change eliminates the older RDD-based ingestion path in favor of the DataFrame-based \saveToMemSQL\ API.

connectorLib/src/main/scala/com/memsql/spark/connector · high confidence

Test coverage

Added integration tests for batch insert, data loading, and connector features; Added test configuration for inline mock maker; Added test logging configuration files.

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

Baseline

  • First survey — no prior run to compare against. CAI 62.

Lenses

  • Code Health 86
  • Architecture 100
  • Maturity 58
  • Readiness 60
  • Security 57

Changes since last survey

  • 300 commits — 261 feature/other, 39 fixes

By area

  • (root) — 112 commits
  • src/main — 109 commits
  • src/test — 34 commits
  • .circleci/config.yml — 17 commits
  • (repo) — 8 commits
  • demo/notebook — 7 commits
  • .idea/modules — 4 commits
  • demo/Dockerfile — 2 commits
  • demo/README.md — 2 commits
  • scripts/ssl — 2 commits
  • .github/workflows — 1 commit
  • ci/secring.asc.enc — 1 commit
  • scripts/ensure-test-memsql-cluster.sh — 1 commit

Notable commits

  • fix: Bump version to 2.0.6 and fix tests
  • fix: Fix "CaseWhen" pushdown test
  • fix: Fix for the thread contention issue (#90)
  • fix: Fixed BinaryType in spark connector
  • fix: Fixed CircleCI build
  • fix: Fixed Databricks compatibility (#102)
  • fix: Fixed Databricks incompatibility (#94)
  • fix: Fixed NoSuchMethodException error
  • fix: Fixed Table has reached its quota of 1 reader(s) error
  • fix: Fixed a bug in the README
  • fix: Fixed aggregate pushdowns with udf
  • fix: Fixed bug which caused reading from the wrong result table (#87)
  • fix: Fixed bugs in FromUnixTime
  • fix: Fixed deadlock in error handling
  • fix: Fixed error handling in BatchInsertWriter
  • fix: Fixed last changes that caused to failure
  • fix: Fixed link in README
  • fix: Fixed links in the README
  • fix: Fixed materializedTableCreationTimeoutMS option name
  • fix: Fixed notebooks to contain new docker names
  • …and 280 more

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

memsql/singlestore-spark-connector was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 20 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit 9da5197965d0721f25f9603e8d7b280512749225 — the exact code this score is about.
  • Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-b51f968c9b10.