neo4j/neo4j-spark-connector
72.9
Strong · 20 September 2026
8.2k
lines of production code
Scala
with Python
1
measurement over time
What this system is
This system is the Neo4j Connector for Apache Spark, a library that enables batch and streaming data integration between Spark and Neo4j databases. It provides capabilities for reading and writing nodes and relationships via the modern Spark DataSource V2 API, supporting structured streaming and complex Cypher query execution. The connector handles data type mapping, schema inference, and query optimization while enforcing strict driver modes and offering comprehensive integration testing infrastructure.
How it got here
2016–2020 — Spark Connector V2 migration
13 changes.
The project underwent a major architectural overhaul to replace legacy RDD-based integration with the modern Spark DataSource V2 API. This migration involved rewriting core Scala components to use the Cypher DSL, implementing new batch and streaming read/write mechanisms, and updating build tooling to support Spark 4 and Java 17.
2021–2026 — v6.0 infrastructure and streaming support
13 changes.
This period focused on establishing the v6.0 release infrastructure, including a comprehensive TeamCity CI pipeline, automated release processes, and robust test coverage with parallel execution support. Significant architectural changes were made to the connector, introducing native type support, a refactored DataSource V2 reader, and new structured streaming implementations for both reading and writing data.
Features
Add Neo4j Spark example notebooks
Added two new example notebooks in the examples directory: a data engineering workflow and a data science workflow. The data engineering notebook demonstrates extracting insights using the Neo4j Connector for Apache Spark with AuraDB, including Spark environment configuration and exercises. The data science notebook shows how to combine the Neo4j Spark connector with the Graph Data Science library, supporting both PySpark and PySpark Pandas APIs.
examples · high confidence
Added README template for ZIP release archives
A new template file for ZIP assembly archives has been introduced, providing users with a standardized README that includes the project version, links to source code and release notes, documentation, and support contact information for both Enterprise customers and the community.
src/jreleaser · high confidence
Initial TeamCity build pipeline configuration
The project now includes a TeamCity build pipeline defined in .teamcity/settings.kts, establishing automated builds for the main branch, pull requests, and scheduled compatibility tests. The configuration specifies supported Java (17, 21), Scala (2.13), PySpark (4.0, 4.1), and Neo4j versions, with pull request builds triggering on PR branches and main branch builds triggering on commits. Compatibility tests are scheduled to run daily at 6 AM. An .editorconfig file is also added to enforce Kotlin formatting rules consistent with ktfmt for contributors.
.teamcity · high confidence
New TeamCity build pipeline configuration for version 6.0
The project has introduced a new TeamCity build pipeline configuration located in \.teamcity/builds\, replacing previous CI setups. This Kotlin-based configuration defines a comprehensive build matrix targeting the 6.0 branch, supporting Java 17 and 21, Scala 2.13, and PySpark versions 4.0 and 4.1. The pipeline includes distinct stages for pull request checks (including a whitelist check and PR-specific validation), semantic code analysis via Semgrep, Maven packaging, unit tests, and integration tests for both Java and Python. It also features a dedicated release build type that handles versioning, JReleaser publishing to Maven Central, and Slack notifications, while optimizing for performance with build caching and disk space requirements.
.teamcity/builds · high confidence
Register Neo4j Spark connector as a data source and expose version
The connector is now automatically discoverable by Spark SQL as a data source via the META-INF/services registration for org.neo4j.spark.DataSource, and a new properties file exposes the connector version for runtime identification.
src/main/resources · high confidence
Repository initialization with modern build and release tooling
The repository has been initialized with a new structure and tooling. A Maven Wrapper (version 3.3.4) is now included to ensure consistent builds, alongside a \maven-release.sh\ script and \jreleaser.yml\ configuration to automate versioning, signing, and deployment to Maven Central. The project now enforces Conventional Commits via \@commitlint/config-conventional\ and a custom \dangerfile.mjs\ that validates commit messages and PR descriptions. Documentation has been migrated to Antora, and the README now specifies that version 6.0.0+ requires Scala 2.13 and supports building for Spark 4.
(repo-wide) · high confidence
Behavioural changes
Enforce commit message format and code style on TeamCity and project files
The pre-commit hook now automatically formats code using Maven wrappers (mvnw) for both the main project and the .teamcity directory, ensuring consistent style via sortpom and spotless. Additionally, a new commit-msg hook enforces commit message standards by running commitlint via pnpm before allowing commits.
.husky · high confidence
Migrate from RDD-based API to Spark DataSource V2
The Neo4j Spark connector has replaced its legacy RDD-based data access methods (CypherRDD, CypherRowRDD, CypherResultRdd, CypherTupleRDD) with the modern Spark DataSource V2 API. Users now interact with Neo4j data through the standard Spark DataFrame/Dataset interfaces via the 'neo4j' data source, which supports both batch and streaming read/write operations, schema inference, and external metadata. The old RDD classes and their supporting configuration classes (Neo4jConfig, BoltConfig, DummyPartition) have been removed.
src/main/scala/org/neo4j/spark · high confidence
Migration from legacy Cypher-Bolt-RDDs to Spark Connector API
The legacy Cypher-Bolt-RDD integration has been removed in favor of the modern Spark Connector API. The \Neo4jSparkContext\ class, which previously provided RDD-based query methods (such as \query\, \queryTuple\, \queryRow\, and \queryDF\), has been deleted. To support the new connector's requirements, a new \ReflectionUtils\ utility has been added to handle compatibility with Spark's \Aggregation\ interface, specifically managing the transition between \groupByColumns\ and \groupByExpressions\.
src/main/java · high confidence
New Spark Connector V2 writer implementation
The writer component in \src/main/scala/org/neo4j/spark/writer\ has been replaced with a new implementation based on the Spark Connector V2 API. This change introduces \Neo4jWriterBuilder\ to handle batch and streaming write construction, \Neo4jBatchWriter\ to manage batch execution and index waiting, and dedicated \Neo4jDataWriter\ and \Neo4jDataWriterFactory\ classes for row-level data writing. The new builder also enforces validation that prevents the use of the 'ErrorIfExists' save mode for relationships, requiring 'Append' instead.
src/main/scala/org/neo4j/spark/writer · high confidence
New data conversion architecture with native type support and legacy mode
The converter package has been restructured into dedicated DataConverter and TypeConverter components, introducing native support for Spark types such as byte, short, decimal, timestamp without timezone (TimestampNTZ), and day-time/year-month intervals (mapped to Neo4j durations). By default, timestamps are now converted as zoned datetimes and local datetimes as timestamp-without-timezone; users can opt back into the previous behavior by enabling the legacy type conversion option.
src/main/scala/org/neo4j/spark/converter · high confidence
New structured streaming read and write implementations
The streaming module now includes new core components for structured streaming: \Neo4jMicroBatchReader\ manages micro-batch offsets and partition planning, \BaseStreamingPartitionReader\ and \Neo4jStreamingPartitionReader\ handle data retrieval with explicit range filtering via \$stream.from\ and \$stream.to\ parameters (replacing the deprecated \$stream.offset\), and \Neo4jStreamingWriter\ along with \Neo4jStreamingDataWriterFactory\ enable data writing. This change introduces the underlying infrastructure for streaming operations, including the enforcement of new streaming parameter syntax and the removal of legacy offset usage in queries.
src/main/scala/org/neo4j/spark/streaming · high confidence
New utility classes and strict driver wrapper for warning enforcement
The connector now includes a new \StrictDriver\ wrapper that enforces strict query mode by checking for Cypher warning notifications on every consumed result and throwing a \CypherWarningException\ if warnings are found. Additionally, a new \Neo4jUnknownCommitOutcomeException\ is introduced to handle cases where a transaction commit outcome is unknown and replay is not permitted. Supporting utilities include \Neo4jImplicits\ for property path parsing and parameter name generation, \Neo4jUtil\ for internal field definitions and filter mapping, \PrimitiveValueDeserializer\ for JSON deserialization, and \ValidationUtil\ for input validation.
src/main/scala/org/neo4j/spark/util · high confidence
Refactored Spark DataSource V2 reader implementation
The partition reader logic has been restructured to use a new DataSource V2 architecture. A new \BasePartitionReader\ abstract class now centralizes connection management, retry logic, and query execution, while \Neo4jPartitionReader\ implements the specific partition reading interface. A \Neo4jPartitionReaderFactory\ is introduced to instantiate these readers, and \Neo4jScanBuilder\ now implements V2 pushdown interfaces (filters, aggregates, columns, limits, and TopN) to allow the Spark optimizer to push down operations to Neo4j.
src/main/scala/org/neo4j/spark/reader · high confidence
Rewrite of Scala connector internals to use Cypher DSL and new Spark Read API
The Scala connector source files have been completely rewritten to replace the previous implementation with a new architecture. This includes a new \Neo4jScan\ class implementing the Spark \Scan\ and \Batch\ interfaces for reading, a \CypherRenderer\ that dynamically selects the Cypher dialect based on the Neo4j server version or user options, and an \AutomaticAliaser\ to handle query result column naming. The \SchemaService\ has been updated to use the Cypher DSL for schema resolution, and the \DriverCache\ now uses a reference-counting strategy to manage driver lifecycle. Additionally, new streaming components (\Neo4jOffset\, \Neo4jStreamingPartitionReaderFactory\) and a \DataWriterMetrics\ system for tracking write operations have been introduced.
scala · high confidence
Rewritten Cypher query generation and relationship write mapping
The service layer for generating Cypher queries and mapping data for writes has been refactored to use the Neo4j Cypher DSL. Read queries are now embedded as CALL subqueries to handle result aliasing, and write operations for nodes and relationships are generated using DSL builders instead of string concatenation. The relationship write mapping strategy now supports a 'NATIVE' mode that explicitly structures source node, target node, and relationship data, alongside the existing 'KEYS' strategy, allowing for more precise control over how relationship data is written to Neo4j.
src/main/scala/org/neo4j/spark/service · high confidence
Test coverage
Added Python integration tests for the Spark-Neo4j pipeline; Added integration tests for Neo4j Spark connector; Added test configuration and resources for parallel execution and Neo4j SSO integration; Added test coverage for Neo4j Spark connector utilities; Added tests for Cypher query service and schema handling; Added tests for Spark data source type reading and refactored test suite structure; New test infrastructure and support utilities for integration tests; New test support utilities for assertions and parallel execution.
Dependencies
Project restructured for Connector 6.0 with updated build tooling and dependencies
The project has been reorganized to align with the Connector 6.0 release, including a major update to the Maven \pom.xml\ (upgrading to \org.neo4j.connectors:spark\ version 6.1.0-SNAPSHOT, Java 17, Scala 2.13.18, Spark 4.1.2, and Neo4j Java Driver 6.2.1). Documentation examples for Java, Python, and Scala have been added to provide users with updated dependency configurations. Additionally, the build infrastructure now uses \pnpm\ (v12.3.4) for Node.js dependencies, introducing \@commitlint\ and \husky\ for commit linting, while the TeamCity configuration has been migrated to a Kotlin-based DSL.
(dependencies) · high confidence
Updated Maven wrapper and distribution to version 3.3.4 and 3.9.16
The Maven wrapper configuration has been updated to use wrapper version 3.3.4 and the Apache Maven distribution version 3.9.16. This ensures that builds executed via the wrapper will download and use the specified Maven binary, aligning the local development and CI environments with the newer release.
.mvn · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Baseline
- First survey — no prior run to compare against. CAI 73.
Lenses
- Code Health 94
- Architecture 100
- Maturity 66
- Readiness 79
- Security 71
Changes since last survey
- 300 commits — 266 feature/other, 34 fixes
By area
- (root) — 161 commits
- src/main — 29 commits
- .teamcity/builds — 28 commits
- src/test — 23 commits
- docs/modules — 17 commits
- common/src — 14 commits
- .github/workflows — 6 commits
- spark-3/src — 5 commits
- .teamcity/settings.kts — 4 commits
- spark/src — 3 commits
- .github/dependabot.yml — 2 commits
- docs/publish.yml — 2 commits
- (repo) — 1 commit
- .husky/pre-commit — 1 commit
- .mvn/wrapper — 1 commit
- docs/package-lock.json — 1 commit
- scripts/python — 1 commit
- test-support/src — 1 commit
Notable commits
- fix: Revert "build: skip signature verification"
- fix: Revert "ci: make release announcement to team-spark too (#762)"
- fix: chore: streamline build & fix warnings (#971)
- fix: ci: fix branch specifiers
- fix: ci: fix compat build scheduling and notifications (#737)
- fix: ci: fix for 6 semgrep check (#842)
- fix: ci: fix snyk test id
- fix: ci: fix snyk token
- fix: ci: fix version parsing for calver (#799)
- fix: fix: add maven assembly plugin (#900)
- fix: fix: add undici as dev dependency
- fix: fix: apply transaction config to auto-commit transactions (#775)
- fix: fix: constraints inner escaping of names (#948)
- fix: fix: do not overflow when computing partition size (#724)
- fix: fix: docs action publish.yaml
- fix: fix: handle alwaysTrue and alwaysFalse (#1102)
- fix: fix: labels containing spaces causes slowdown (#1114)
- fix: fix: neo4j options should always validate (#961)
- fix: fix: omit scriptResult clause when script is not provided (#951)
- fix: fix: race condition in DriverCache (#1043)
- …and 280 more
Architecture
- 0 containers · 1 bounded contexts · 0 dependency edges (baseline)
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
neo4j/neo4j-spark-connector was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 20 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 81ba90749e01822c144a523f23074abb108a585d — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-b51f968c9b10.