spotify/scio
70.7
Strong · 27 September 2026
84.5k
lines of production code
Scala
with Java
3
measurements over time
What this system is
This system is a Scala library (Scio) that provides a high-level API for building and running Apache Beam data processing pipelines. It offers extensive integration with various data storage and processing systems, including Google Cloud services (BigQuery, Bigtable, Datastore, Spanner, Pub/Sub), databases (Cassandra, JDBC, Snowflake, Neo4j, Redis), and file formats (Avro, Parquet, Protobuf, CSV). The library also includes utilities for efficient data operations such as Sort-Merge Bucket joins, dynamic file routing, and interactive REPL development, alongside comprehensive tooling for automated code migration across API versions.
How it got here
2015–2017 — Apache Beam migration and feature expansion
56 changes.
The project migrated from the legacy Google Cloud Dataflow SDK to Apache Beam, removing all deprecated core libraries, examples, and test infrastructure. This period focused on rebuilding the codebase with modern Beam APIs, introducing new I/O modules for Avro, Parquet, JDBC, and TensorFlow, and expanding the example suite to cover complex data processing patterns.
2018–2020 — Google Cloud I/O and serialization expansion
67 changes.
This period focused on consolidating Google Cloud Platform integrations into a new scio-google-cloud-platform module, adding comprehensive I/O support for BigQuery, Bigtable, Datastore, and Spanner. Concurrently, the codebase expanded its core capabilities with robust serialization improvements, including Magnolia-based schema derivation, Kryo serializer updates, and new batched async DoFns for efficient data processing.
2021–2026 — IO expansion and migration tooling
47 changes.
This period focused on expanding Scio's I/O capabilities by adding support for Neo4j, Snowflake, Iceberg, and dynamic file destinations, while significantly enhancing Parquet and SMB features. Concurrently, the project invested heavily in automated migration tooling via Scalafix rules to guide users through API changes across multiple versions, alongside extensive integration testing for new and existing connectors.
Features
Add Beam SQL and Minimal Word Count example pipelines
The scio-examples module now includes two new Java example pipelines: a Beam SQL example demonstrating how to define schemas, create input data, and execute SQL queries (including simple filtering and aggregation) over PCollections, and a Minimal Word Count example showing how to read text files, split lines into words, count occurrences, and write results to text files.
scio-examples/src/main/java/org/apache/beam/examples · high confidence
Add BigQuery and Bigtable service wrappers for Google Cloud Platform integration
New wrapper classes have been added to the scio-google-cloud-platform module to facilitate interactions with Google Cloud services. BigQueryServicesWrapper exposes package-private methods for BigQuery operations such as creating tables, checking if tables are empty, inserting rows, and retrieving table metadata. Similarly, BigtableServiceHelper wraps the Bigtable implementation, handling the translation of Bigtable options to veneer settings and providing a simplified interface for opening writers to specific tables.
scio-google-cloud-platform/src/main/java/org · high confidence
Add Bigtable helper utilities and package syntax
This change introduces new helper objects for Google Cloud Bigtable within the \scio-bigtable\ package. \BTOptions\ provides a builder for Bigtable connection options, while \Mutations\ and \Rows\ offer convenience methods for constructing Bigtable mutations (such as \SetCell\, \DeleteFromColumn\, \DeleteFromFamily\, and \DeleteFromRow\) and rows with optional timestamps. The \package.scala\ file exposes these utilities via the \AllSyntax\ import, enabling users to easily create Bigtable data structures in Scio pipelines.
scio-google-cloud-platform/src/main/scala/com/spotify/scio/bigtable · high confidence
Add Coder instances for Redis mutation types
The scio-redis module now provides implicit Coder instances for various Redis mutation types (such as Append, Set, IncrBy, DecrBy, SAdd, LPush, RPush, PFAdd, and ZAdd) for both String and byte array data. This enables these mutations to be serialized and deserialized correctly within Scio pipelines, supporting the underlying RedisIO write operations.
scio-redis/src/main/scala/com/spotify/scio/redis/instances · high confidence
Add Managed Iceberg IO syntax extensions
Introduces new syntax extensions for reading and writing Iceberg tables via Beam's Managed IO. Users can now use \sc.iceberg(...)\ on \ScioContext\ to read tables and \sc.saveAsIceberg(...)\ on \SCollection\ to write tables, with support for catalog configuration, table properties, sort and partition fields, and streaming-specific options like triggering frequency and direct write byte limits. The implementation also handles automatic inference of the \keep\ configuration from Magnolify RowType schemas to ensure correct field selection during read and write operations.
scio-managed · high confidence
Add Neo4j IO support for Scio
This change introduces a new Neo4j integration module (\scio-neo4j\) that enables reading from and writing to Neo4j databases via Cypher queries. Users can now use \scioContext.neo4jCypher\ to read data into an SCollection and \sCollection.saveAsNeo4j\ to write data out, supporting both parameterized read queries and UNWIND-based batch writes. The implementation leverages Apache Beam's Neo4jIO under the hood and includes syntax extensions for both ScioContext and SCollection.
scio-neo4j · high confidence
Add Parquet TensorFlow support
Users can now read and write Parquet files containing TensorFlow examples by importing the new \com.spotify.scio.parquet.tensorflow\ package, which exposes the necessary syntax and APIs for handling TensorFlow-specific data formats within Scio.
scio-parquet/src/main/scala/com/spotify/scio/parquet/tensorflow · high confidence
Add Parquet metadata extraction for file collections
Users can now extract detailed metadata (schema, block statistics, and key-value metadata) from Parquet files using the new \parquetMetadata\ method on \SCollection\[String\]\ (file paths) and \SCollection\[ReadableFile\]\. This feature allows inspecting Parquet file structure and content statistics without reading the actual data records.
scio-parquet/src/main/scala/com/spotify/scio/parquet/syntax · high confidence
Add Redis read and write syntax extensions for ScioContext and SCollection
Users can now read from and write to Redis directly within Scio pipelines using new implicit syntax extensions. ScioContext gains a \redis\ method to read key-value pairs based on a pattern, while SCollection gains a \saveAsRedis\ method to write Redis mutation objects, enabling seamless integration with Redis data stores.
scio-redis/src/main/scala/com/spotify/scio/redis/syntax · high confidence
Add Snowflake I/O support for Scio
This change introduces a new \scio-snowflake\ module that enables reading from and writing to Snowflake tables and queries within Scio pipelines. Users can now use \sc.snowflakeQuery\ and \sc.snowflakeTable\ to read data into an SCollection, and \scollection.saveAsSnowflake\ to write data out, leveraging Beam's Snowflake IO under the hood with configurable connection options, staging buckets, and write dispositions.
scio-snowflake · high confidence
Add SortedBucketIOUtil for test ID generation
A new utility object, SortedBucketIOUtil, has been added to the scio-smb module to facilitate test identification for SortedBucketIO operations. This utility provides methods to generate unique test IDs for read, write, and transform output operations by delegating to the SmbIO testId implementation, ensuring consistent identification in test scenarios.
scio-smb/src/main/scala/org · high confidence
Add Sparkey-backed side inputs and Scala 2.12 compatibility layer
Users can now use Sparkey-backed collections as large side inputs in Scio via the new \.asLargeMapSideInput\ and \.asLargeSetSideInput\ methods, which rely on the newly introduced \SparkeyMapBase\ and \SparkeySetBase\ traits that enforce read-only access by throwing \NotImplementedError\ on mutation operations. Additionally, a compatibility shim in \kantan.codecs\ adds an \iterator\ method to \ResourceIterator\ to support Scala 2.12.
scio-extra/src/main/scala-2.12 · high confidence
Add deterministic random sampling utilities
Added new \RandomSampler\ and \XORShiftRandom\ components in \scio-core\ to support deterministic sampling. The \RandomSampler\ provides \BernoulliSampler\ and \PoissonSampler\ classes that accept an optional seed, allowing users to configure \SCollection.sample\ and similar operations to produce reproducible results. The underlying \XORShiftRandom\ implementation replaces the standard Java RNG for improved performance and deterministic seeding capabilities.
scio-core/src/main/scala/com/spotify/scio/util/random · high confidence
Add dynamic destination syntax for saving text files
Users can now write SCollection data to multiple file paths determined at runtime using the new \saveAsDynamicTextFile\ method. This feature, exposed via the \scio-core\ dynamic syntax package, allows specifying a destination function that maps each element to a target file path, along with standard write options like sharding, compression, and optional headers or footers.
scio-core/src/main/scala/com/spotify/scio/io/dynamic/syntax · high confidence
Add experimental annotation for unstable public APIs
A new \@experimental\ annotation has been added to the \com.spotify.scio.annotations\ package. This annotation allows developers to mark public classes, methods, or fields as subject to incompatible changes or removal in future releases, providing a clear signal to users that the stability of the marked API is not guaranteed.
scio-core/src/main/scala/com/spotify/scio/annotations · high confidence
Add gRPC lookup and batched lookup capabilities to Scio
The scio-grpc module now exposes new extension methods on SCollection to perform asynchronous gRPC lookups. Users can call grpcLookup for single-request lookups, grpcLookupStream for server-streaming responses, and the new grpcBatchLookup and grpcLookupBatchStream methods to batch multiple requests into a single gRPC call, reducing network overhead. These methods support optional caching via a CacheSupplier and are available through the com.spotify.scio.grpc package object.
scio-grpc · high confidence
Add new Scio example applications
The scio-examples module now includes five new runnable examples: MinimalWordCount, WordCount (with metrics), DebuggingWordCount (with assertions), WindowedWordCount (with windowing logic), and MinimalWordCountPipelineOptionsExample (with typed arguments). These provide users with reference implementations for common Scio patterns, including basic word counting, metric tracking, debugging with PAssert, windowed processing, and typed pipeline options.
scio-examples/src/main/scala/com/spotify/scio/examples · high confidence
Add streaming word extraction, TF-IDF, and Wikipedia session examples
Three new Scio examples are added to the complete examples directory: StreamingWordExtract demonstrates a streaming pipeline that extracts words from text and writes them to BigQuery; TfIdf computes Term Frequency-Inverse Document Frequency metrics from a corpus of text files using the FileSystems API; and TopWikipediaSessions analyzes Wikipedia edit logs to identify top user sessions using windowing and sampling techniques.
scio-examples/src/main/scala/com/spotify/scio/examples/complete · high confidence
Add support for writing Avro records to dynamic Parquet destinations
Users can now write Scio SCollection of Avro records to Parquet files using dynamic destinations via the new \saveAsDynamicParquetAvroFile\ method. This feature, exposed through the \com.spotify.scio.parquet.avro.dynamic\ package, allows specifying a destination function to determine output paths, along with standard write parameters like schema, compression, and file prefixes/suffixes.
scio-parquet/src/main/scala/com/spotify/scio/parquet/avro/dynamic · high confidence
Add support for writing Parquet files to dynamic destinations
Users can now write typed Parquet files to dynamic destinations by providing a function that determines the output path for each record. This change introduces the \saveAsDynamicTypedParquetFile\ method on \SCollection\, allowing flexible file routing while supporting standard write parameters such as prefix, suffix, compression, and shard count.
scio-parquet/src/main/scala/com/spotify/scio/parquet/types/dynamic · high confidence
Add syntax helpers for reading and writing TensorFlow Parquet Example files
New implicit syntax extensions are added to ScioContext and SCollection, enabling users to read Parquet files containing TensorFlow Example records via scio.parquetExampleFile and write them via saveAsParquetExampleFile. These methods accept parameters for schema projection, Hadoop configuration, compression, sharding, and metadata, providing a convenient API for TensorFlow model data I/O.
scio-parquet/src/main/scala/com/spotify/scio/parquet/tensorflow/syntax · high confidence
Added Cassandra compatibility and bulk write utility classes
The scio-cassandra module now includes two new Java utility classes to support Cassandra integration. CompatUtil provides helpers for handling DataStax Java Driver compatibility, including retrieving protocol versions and serializing data types. CqlBulkRecordWriterUtil offers a workaround to expose the package-private CqlBulkRecordWriter constructor, enabling bulk record writing to Cassandra tables with configurable hosts, ports, authentication, and schema settings.
scio-cassandra/cassandra3/src/main/java · high confidence
Added SQL Server integration test infrastructure and sample data
This change introduces the necessary components to support integration testing against SQL Server. It adds a \SqlServer\ configuration object defining the connection details for the test instance, a \PopulateTestData\ utility to seed the database with sample employee records, and sample data files (Avro, CSV, JSON, and Avro schema) for other integration tests. These resources are now available in the compile scope to facilitate the new SQL Server test scenarios.
integration/src/main · high confidence
Added Scala 2.13-specific JMapWrapper utility
A new \JMapWrapper\ utility has been added for Scala 2.13 to provide immutable wrappers around \java.util.Map\ instances. This component bridges the gap between Java's mutable collections and Scala's idiomatic immutable maps, ensuring consistency when wrapping Beam API data structures by using the Scala 2.13 \scala.jdk.CollectionConverters\ library.
scio-core/src/main/scala-2.13/com/spotify/scio/util · high confidence
Added game data injector example for Pub/Sub and file output
The scio-examples module now includes a new Java-based Injector utility that simulates mobile game usage data (teams, scores, and robot activity) and publishes it to a Google Cloud Pub/Sub topic or writes it to a local file. This addition provides a concrete example of generating synthetic event streams for testing Beam pipelines, with built-in support for service account authentication and automatic topic creation.
scio-examples/src/main/java/org/apache/beam/examples/complete/game/injector · high confidence
Added serialization support for Cassandra DataType
Users can now serialize and deserialize Cassandra DataType objects within Scio pipelines. This is achieved by introducing a new DataTypeExternalizer that registers custom Kryo serializers for Guava's ImmutableList and ImmutableSet, ensuring these types are handled correctly during distributed data processing.
scio-cassandra/cassandra3/src/main/scala · high confidence
Automated migration rules for Scio versions 0.7 through 0.15
This release introduces a comprehensive suite of Scalafix migration rules covering API changes from version 0.7.0 up to 0.15.0. The rules automate the upgrade process for various breaking changes, including refactoring the BigQuery client API, updating Avro coder and schema handling, migrating PubSub IO specializations, and adjusting SMB join parameters. Specific fixes address changes in join naming conventions, Bigtable and Datastore IO signatures, TensorFlow proto imports, and the removal of deprecated features like skewed joins and logical type suppliers, ensuring code remains compatible with newer Scio versions.
scalafix/rules/src/main/scala · high confidence
Expanded support for Java, Algebird, and Zstd serialization coders
The Scio core library now includes dedicated coder implementations for a broader range of data types, improving out-of-the-box serialization support. This change adds coders for Java standard library types (including java.time, java.util collections, and java.sql.Timestamp), Algebird statistical types (such as Moments, CMS, and TopK), and Google Protobuf messages. Additionally, it introduces a Zstd compression coder for efficient data encoding and a custom coder for Guava BloomFilters. These additions allow users to handle these specific types in their pipelines without needing to define custom serialization logic.
scio-core/src/main/scala/com/spotify/scio/coders/instances · high confidence
Initial release of the scio-google-cloud-platform BigQuery module
This change introduces the new \scio-google-cloud-platform\ module, which consolidates BigQuery-specific functionality previously scattered across the codebase. It provides a comprehensive API for interacting with BigQuery, including typed table annotations, schema parsing, and system property configuration for caching, timeouts, and authentication (including service account impersonation). The module also adds support for the BigQuery Storage API, BIGNUMERIC and JSON data types, and DML statements. For testing, it includes a \MockBigQuery\ environment that supports mocking tables, wildcard tables, and live DML execution, along with utilities for resolving \$LATEST\ partitioned tables in queries.
scio-google-cloud-platform/src/main/scala/com/spotify/scio/bigquery · high confidence
Initial scalafix migration infrastructure and rules for scio 0.7, 0.13, and 0.14
This change introduces the foundational scalafix tooling for the scio project, including SBT build configuration (sbt 1.12.13, sbt-scalafix 0.14.7) and a project matrix structure to test migration rules in isolation. It adds specific migration rules and corresponding input/output test cases for upgrading to scio versions 0.7, 0.13, and 0.14. For version 0.7, the \FixAvroIO\ rule updates \AvroIO\ calls to include explicit type parameters (e.g., \AvroIO\[InputClass\]\) and adjusts imports. For version 0.13, the \FixTfParameter\ rule simplifies the \predict\ and \predictWithSigDef\ method signatures by removing the second type parameter. For version 0.14, the \FixAvroCoder\ rule adds necessary Avro imports in test contexts.
scalafix · high confidence
Initial support for Google Cloud Datastore I/O operations
This change introduces the \scio-google-cloud-platform\ module, providing new syntax extensions for reading from and writing to Google Cloud Datastore. Users can now use \sc.datastore()\ and \sc.typedDatastore()\ on \ScioContext\ to read queries as \SCollection\s, and \saveAsDatastore()\ on \SCollection\s to write data. The implementation supports both raw \com.google.datastore.v1.Entity\ types and typed entities via Magnolify's \EntityType\, allowing for flexible integration with Datastore datasets.
scio-google-cloud-platform/src/main/scala/com/spotify/scio/datastore · high confidence
Initial support for Google Cloud Spanner integration
This change introduces the \scio-google-cloud-platform\ module, adding the ability to read from and write to Google Cloud Spanner within Scio pipelines. Users can now use \sc.spannerTable\ and \sc.spannerQuery\ on \ScioContext\ to read data as \Struct\ collections, and call \saveAsSpanner\ on \SCollection\[Mutation\]\ to write mutations. The implementation includes necessary Coder instances for Spanner types (ReadOperation, Struct, Mutation, MutationGroup) to ensure correct serialization, and provides client helpers for database and admin operations.
scio-google-cloud-platform/src/main/scala/com/spotify/scio/spanner · high confidence
Initial support for writing Parquet files with Avro schemas
This change introduces the core components for writing Parquet files using Avro schemas in Scio. It adds a new \ParquetAvroSink\ class that handles the actual writing process, supporting configurable compression codecs and row group sizes via configuration. Additionally, a \package.scala\ file is added to expose the Parquet Avro API, including aliases for Projection and Predicate types and importing the necessary syntax for users to easily integrate Parquet Avro writes into their pipelines.
scio-parquet/src/main/scala/com/spotify/scio/parquet/avro · high confidence
Introduce BigQuery typed API with macro-generated converters
The \scio-google-cloud-platform\ module now includes the \BigQueryType\ macro annotations and supporting infrastructure (including \BigQueryTag\, \ConverterProvider\, \SchemaProvider\, and \TypeProvider\). This adds the capability to automatically generate Scala case classes and companion objects for BigQuery tables, schemas, and queries, providing built-in converters between Scala types and BigQuery/Avro representations.
scio-google-cloud-platform/src/main/scala/com/spotify/scio/bigquery/types · high confidence
Introduce Bigtable syntax extensions for ScioContext and SCollection
This change adds the Bigtable syntax package (\com.spotify.scio.bigtable.syntax\), providing convenient extension methods for reading from and writing to Google Cloud Bigtable. \ScioContextOps\ adds \typedBigtable\ and \bigtable\ methods to load data as typed collections or raw \Row\ objects, supporting features like key-range partitioning, row filtering, and configurable buffer sizes. \SCollectionSyntax\ adds \saveAsBigtable\ methods to write mutations or typed data, including support for flow control and bulk batching. \RowSyntax\ provides helper methods on \Row\ objects to easily extract cell data, family maps, and latest values.
scio-google-cloud-platform/src/main/scala/com/spotify/scio/bigtable/syntax · high confidence
Introduce DataflowResult for independent job status and metrics retrieval
Users can now create a DataflowResult instance independently from a running pipeline execution by providing a project ID, region, and job ID. This new class wraps the Dataflow Pipeline Job and exposes methods to fetch the current Job details and JobMetrics via the Dataflow API, allowing for better observability and status checking of existing Dataflow jobs without needing the original ScioContext or PipelineResult.
scio-core/src/main/scala/com/spotify/scio/runners · high confidence
Introduce JDBC API package with AllSyntax import
Users can now import the entire JDBC API surface using a single wildcard import (\import com.spotify.scio.jdbc.\_\). This change adds a new \package.scala\ file that extends \AllSyntax\, consolidating the available JDBC syntax and operations into the main package namespace for easier access.
scio-jdbc/src/main/scala/com/spotify/scio/jdbc · high confidence
Introduce Redis write mutations with TTL support
The scio-redis module now includes a new RedisMutator implementation that handles write operations (such as SET, APPEND, INCRBY, SADD, LPUSH, RPUSH, PFADD, and ZADD) within Redis transactions. This change enables users to perform these mutations atomically and supports setting a time-to-live (TTL) on keys using the PSETEX command for SET operations and PEXPIRE for others, ensuring data expiration is handled consistently during batch writes.
scio-redis/src/main/scala/com/spotify/scio/redis · high confidence
Introduce Scala API for Sort-Merge Bucket (SMB) operations
This change adds the core Scala API for the Sort-Merge Bucket module, introducing \SmbIO\ as the primary interface for reading and writing sorted bucket data and \SortMergeTransform\ to handle the underlying merge logic. Users can now leverage these components to perform efficient sort-merge joins and writes within Scio pipelines, with \SmbIO\ handling test ID generation and tap creation, while \SortMergeTransform\ exposes builders for defining transformation functions and side inputs.
scio-smb/src/main/scala/com/spotify/scio/smb · high confidence
Introduce Scio REPL with interactive shell and file I/O commands
The Scio REPL is now available as an interactive shell (ScioShell) that provides a custom prompt, automatic imports for Scio and BigQuery, and a pre-configured BigQuery client. Users can manage Scio contexts via new magic commands (:newScio, :newLocalScio, :scioOpts) and perform direct file I/O operations (read/write for Avro, Text, CSV, and TSV) through the new IoCommands class. The REPL also includes a custom classloader for runtime class handling and a convenience method (runAndCollect) on SCollection to execute pipelines and collect results interactively.
scio-repl/src/main/scala · high confidence
Introduce Sort-Merge Bucket (SMB) API for efficient joins and grouping
This change introduces the Sort-Merge Bucket (SMB) syntax layer in Scio, enabling users to perform highly efficient joins and group-by-key operations on pre-sorted data without a full shuffle. The new API adds \sortMergeJoin\ and \sortMergeGroupByKey\ methods to \ScioContext\, allowing joins between datasets where keys are extracted from bucket metadata, as well as a \saveAsSortedBucket\ method on \SCollection\ to write data into SMB-compatible buckets. These features are marked as experimental and include support for secondary sort keys and test-mode simulation.
scio-smb/src/main/scala/com/spotify/scio/smb/syntax · high confidence
Introduce Sorted Bucket Merge (SMB) module with Avro, JSON, and Iceberg support
The scio-smb module is introduced, providing a Sorted Bucket Merge implementation for efficient large-scale joins. This change adds core infrastructure including file operations for Avro and JSON formats, an Iceberg encoder for key types, and utilities for bucket metadata and iterator buffering. It also includes patched Beam classes (such as CoGbkResultUtil) to resolve classloader conflicts and enable serialization of SMB-specific components.
scio-smb/src/main/java · high confidence
Introduce dedicated schema instance traits for Scala, Java, and Joda types
The \scio-core\ schemas module now organizes its type support into distinct traits (\ScalaInstances\, \JavaInstances\, \JodaInstances\, and \LowPrioritySchemaDerivation\) consolidated in \AllInstances\. This change explicitly adds schema definitions for Java standard types (such as \java.lang.Integer\, \java.math.BigDecimal\, and \java.util.List\), Joda Time's \ReadableInstant\, and Scala collections (including \List\, \Vector\, \Map\, and \Set\), while also integrating automatic case-class schema derivation via Magnolia 1.
scio-core/src/main/scala/com/spotify/scio/schemas/instances · high confidence
Introduce extensible BigQuery type mapping via OverrideTypeProvider
Added a new validation package in scio-google-cloud-platform that allows users to customize how BigQuery schema fields map to Scala types. The change introduces an OverrideTypeProvider interface and a finder mechanism that loads a custom implementation via the 'override.type.provider' system property, enabling advanced type mapping scenarios beyond the default behavior.
scio-google-cloud-platform/src/main/scala/com/spotify/scio/bigquery/validation · high confidence
Introduce macro-based Avro type generation and conversion
The scio-avro module now provides a macro-driven system for mapping Avro schemas to Scala case classes and vice versa. Users can use annotations like @AvroType.fromSchema, @AvroType.fromPath, and @AvroType.fromSchemaFile to automatically generate case classes from Avro schemas or files at compile time. The @AvroType.toSchema annotation allows case classes to be serialized to Avro. This includes macro implementations in TypeProvider, converter generation in ConverterProvider, schema resolution in SchemaProvider, and utility functions in SchemaUtil and MacroUtil, enabling seamless integration between Avro data formats and Scala types.
scio-avro/src/main/scala/com/spotify/scio/avro/types · high confidence
Introduce new BigQuery client library with typed API and service account impersonation
Adds a new BigQuery client implementation in the \scio-google-cloud-platform\ module, providing a structured API for interacting with BigQuery. This includes support for service account impersonation via the \bigquery.act\_as\ system property, a type-safe API for reading and writing rows using \BigQueryType\ annotations, and the ability to create typed tables with explicit descriptions. The client also introduces a caching mechanism for schemas and table references, and supports loading data in CSV, JSON, and Avro formats with configurable options.
scio-google-cloud-platform/src/main/scala/com/spotify/scio/bigquery/client · high confidence
Introduce schema pretty-printing and materialization utilities
Added a new PrettyPrint utility to the schemas package that formats Beam Schema field definitions into a readable, table-like string representation, including support for nested row types. Additionally, introduced the SchemaMaterializer object, which provides core logic for converting Scio Schema definitions into Apache Beam Schema objects and generating the necessary serialization/deserialization functions (encoders and decoders) to map between Scio types and Beam Row objects.
scio-core/src/main/scala/com/spotify/scio/schemas · high confidence
Introduce syntax-based API for Parquet Avro read and write operations
This change adds new implicit syntax extensions for ScioContext and SCollection to simplify interacting with Parquet files containing Avro records. Users can now read Parquet data as Avro GenericRecords via the new \parquetAvroFile\ method on ScioContext, which supports schema projection and filter predicates, and write Avro IndexedRecords to Parquet using the \saveAsParquetAvroFile\ method on SCollection, which allows configuration of compression, sharding, and metadata. The implementation introduces a \ParquetAvroFile\ wrapper class to handle the read-side logic, including safe conversion to SCollection and support for map/flatMap transformations.
scio-parquet/src/main/scala/com/spotify/scio/parquet/avro/syntax · high confidence
Introduces utility classes for joins, caching, and call-site tracking
This change adds several new internal utility components to the Scio core library. It introduces ArtisanJoin and MultiJoin to provide optimized, single-shuffle implementations for joining and co-grouping multiple SCollection inputs, including support for various join types (inner, left, right, outer). A new Cache abstraction is added, offering implementations backed by Caffeine, Guava, and ConcurrentHashMap to manage side-input and other internal caches. Additionally, CallSites is introduced to improve transform naming and debugging by accurately tracking call locations, while FilenamePolicySupplier and ScioUtil provide centralized logic for file output naming and temporary location resolution.
scio-core/src/main/scala/com/spotify/scio/util · high confidence
JDBC syntax extensions for ScioContext and SCollection
This change introduces new implicit syntax extensions for ScioContext and SCollection to simplify JDBC operations. Users can now use sc.jdbcSelect() to read data from JDBC queries with configurable fetch size, parallelization, and custom data sources, and sc.jdbcShardedSelect() to perform sharded reads from tables or materialized views for improved performance. Additionally, SCollection gains a saveAsJdbc() method to write data to JDBC databases, supporting batch size configuration, retry strategies, auto-sharding, and custom data source providers.
scio-jdbc/src/main/scala/com/spotify/scio/jdbc/syntax · high confidence
New Beam cookbook examples for data analysis patterns
Added new Java example files in the scio-examples cookbook directory demonstrating common Apache Beam data analysis patterns. These include BigQueryTornadoes for counting events by month, DistinctExample for deduplicating text lines, FilterExamples for filtering with side inputs and mean calculations, JoinExamples for combining multiple data sources, MaxPerKeyExamples for finding maximum values per key, and TriggerExample for streaming windowing and triggering behaviors.
scio-examples/src/main/java/org/apache/beam/examples/cookbook · high confidence
New BigQuery I/O code snippets for documentation
Added a new \Snippets.java\ file containing code examples for reading from and writing to BigQuery, including usage of table specs, queries (including Standard SQL), and schema definitions, intended for use in web documentation.
scio-examples/src/main/java/org/apache/beam/examples/snippets · high confidence
New BigQuery syntax extensions for ScioContext, SCollection, and TableRow
This change introduces a new \scio-google-cloud-platform/src/main/scala/com/spotify/scio/bigquery/syntax\ package containing several new implicit syntax extensions. \ScioContextSyntax\ and \SCollectionSyntax\ add methods for reading and writing BigQuery data, including support for the Storage API, typed queries, and saving as TableRow JSON files. \MagnolifySyntax\ provides a new API for typed BigQuery operations using Magnolify, while \TableRowSyntax\ adds typed getters (e.g., \getBoolean\, \getInt\) to \TableRow\ for easier data extraction. \FileStorageSyntax\ adds a \tableRowJsonFile\ method to read TableRow JSON files, and \TableReferenceSyntax\ adds an \asTableSpec\ helper.
scio-google-cloud-platform/src/main/scala/com/spotify/scio/bigquery/syntax · high confidence
New Bigtable utility for dynamic node scaling
Added BigtableUtil and ChannelPoolCreator classes to the scio-google-cloud-platform module, providing utilities to manage Google Cloud Bigtable clusters. Users can now programmatically update the number of nodes in Bigtable clusters (optionally targeting specific clusters by name) to optimize throughput and cost during batch jobs, and retrieve current cluster sizes.
scio-google-cloud-platform/src/main/java/com · high confidence
New HourlyTeamScore example for windowed gaming analytics
Added the HourlyTeamScore example, which extends the existing UserScore pipeline to demonstrate fixed-windowing and element timestamping. This new tool allows users to calculate sum-of-scores per team within specific time windows, with optional filtering to exclude late-arriving data outside a defined start and stop time range.
scio-examples/src/main/java/org/apache/beam/examples/complete/game · high confidence
New JDBC lookup infrastructure with optional password support
The scio-jdbc module now includes a new abstract JdbcDoFn class that enables synchronous JDBC lookups by managing DataSource and Connection lifecycle (setup, teardown, and bundle start). To support this, a new CloudSqlOptions interface has been added to define connection parameters, specifically making the Cloud SQL database password optional while keeping the username, database name, and instance connection name as required fields. A corresponding CloudSqlOptionsRegistrar ensures these options are automatically registered with the Beam PipelineOptions system.
scio-jdbc/src/main/java · high confidence
New Kryo serializers for Protobuf, Beam, gRPC, Joda Time, and Java Path types
This change introduces a set of new Kryo serializers in the scio-core module to handle serialization of common types that previously lacked efficient or correct handling. Specifically, it adds serializers for Google Protobuf ByteString, Apache Beam Coder instances, gRPC Status objects, Java NIO Path, and Joda Time types (DateTime, LocalDate, LocalDateTime, LocalTime). It also includes a renamed KVSerializer that now supports Apache Beam's KV type instead of the legacy Cloud Dataflow SDK, ensuring compatibility with the current Beam version. These additions improve serialization reliability and performance for these data types within Scio pipelines.
scio-core/src/main/scala/com/spotify/scio/coders/instances/kryo · high confidence
New PubSub admin and coder utilities in scio-google-cloud-platform
The scio-google-cloud-platform module now includes PubSub administration helpers and Coder instances. PubSubAdmin provides methods to ensure topics and subscriptions exist and to retrieve their configurations via gRPC, while the coders package exposes an implicit Coder for PubsubMessage using Beam's PubsubMessageWithAttributesCoder.
scio-google-cloud-platform/src/main/scala/com/spotify/scio/pubsub · high confidence
New RemoteFileUtil and TransformingCache utilities for Scio
Added RemoteFileUtil, a new utility class for handling remote file systems within DoFns, which supports checking file existence, downloading single or batched URIs in parallel using a Guava-backed cache, deleting local copies, and uploading files with optional MIME type specification. Also added TransformingCache, an abstract cache wrapper that allows key transformation via an injective function, enabling caching strategies where the stored key differs from the lookup key.
scio-core/src/main/java/com/spotify/scio/util · high confidence
New SCollection transform syntax traits for parallelism, file downloads, and safe operations
The \scio-core\ library introduces a new \com.spotify.scio.transforms.syntax\ package that consolidates several SCollection extension traits. Users can now access \mapWithParallelism\, \flatMapWithParallelism\, \filterWithParallelism\, and \collectWithParallelism\ via \SCollectionParallelismSyntax\ to control concurrent DoFn threads. \SCollectionFileDownloadSyntax\ adds \mapFile\ and \flatMapFile\ for downloading URI elements as local paths. \SCollectionSafeSyntax\ provides \safeFlatMap\ to handle exceptions by routing faulty elements to an error side output. Additionally, \SCollectionPipeSyntax\ enables piping elements through external commands, and \SCollectionWithResourceSyntax\ offers resource-aware variants like \mapWithResource\ and \flatMapWithResource\. These are aggregated in the new \AllSyntax\ trait.
scio-core/src/main/scala/com/spotify/scio/transforms/syntax · high confidence
New SMBMultiJoin utility for multi-source sorted-bucket joins
Added the SMBMultiJoin utility class, which provides sortMergeCoGroup methods capable of joining up to six sorted-bucket sources (A through F) in a single operation. This allows users to perform efficient sorted-merge joins across multiple datasets without chaining multiple binary joins, with support for specifying target parallelism and handling test-mode execution.
scio-smb/src/main/scala/com/spotify/scio/smb/util · high confidence
New Scala-native async and resource-handling DoFns in scio-core
The \scio-core/src/main/scala/com/spotify/scio/transforms\ package now includes new Scala-idiomatic DoFns: \ScalaAsyncDoFn\ and \ScalaAsyncLookupDoFn\ for handling external service calls via Scala \Future\, a \ScalaAsyncBatchLookupDoFn\ for batched lookups, and \Parallel\*\ DoFns (\ParallelCollectFn\, \ParallelFilterFn\, \ParallelMapFn\, \ParallelFlatMapFn\) that limit parallelism using \ParallelLimitedFn\. Additionally, \WithResourceDoFns\ (\CollectFnWithResource\, \MapFnWithResource\, \FlatMapFnWithResource\, \FilterFnWithResource\) allow DoFns to manage external resources (e.g., DB connections) with lifecycle management, and \JavaAsyncConverters\ provides a \RichAsyncLookupDoFnTry\ implicit class to convert Java \AsyncLookupDoFn.Try\ to Scala \Try\. These changes expand the available transform primitives for async operations and resource management in Scio pipelines.
scio-core/src/main/scala/com/spotify/scio/transforms · high confidence
New Scio and Kryo pipeline options with registrar support
Users can now configure Scio-specific behaviors and Kryo serialization details via new PipelineOptions interfaces. ScioOptions exposes settings for blocking job completion, metrics storage location, custom application arguments, chained cogroup validation, nullable coder usage, and Zstd dictionary mappings, while also exposing Scio and Scala versions. KryoOptions allows tuning buffer sizes, reference tracking, and registration requirements. The new ScioOptionsRegistrar ensures these options are automatically discovered by the Beam pipeline options framework.
scio-core/src/main/java/com/spotify/scio/options · high confidence
New automation and code-generation scripts for CI, site builds, and Scio version bumps
The scripts directory now includes several new tools to streamline development and release workflows. The bump\_scio.sh script automates updating the Scio and Beam versions in downstream repositories (such as big-data-rosetta-code, featran, and ratatool) and the homebrew formula by fetching the latest release metadata and submitting pull requests. CI and build processes are supported by ci\_repl.sh for running REPL integration tests, gha\_setup.sh for configuring SBT options via environment variables, and make-site.sh for generating and publishing the documentation site. Additionally, Python scripts (multijoin.py, smb\_multijoin.py, tuplecoders.py) are added to generate Scala code for multi-join transforms and tuple coders, while gen\_schemas.sh and associated JSON files manage BigQuery schema caching.
scripts · high confidence
New batched and resource-managed async DoFns for Scio transforms
The \scio-core\ transforms package now includes \BatchDoFn\ for buffering and emitting elements in weight-based batches while preserving windowing, and a new family of batched async lookup DoFns (\BaseAsyncBatchLookupDoFn\, \GuavaAsyncBatchLookupDoFn\, \JavaAsyncBatchLookupDoFn\) that support configurable batch sizes, max pending requests, and optional request deduplication with caching. These batched lookups are backed by \DoFnWithResource\, which introduces a \ResourceType\ enum (PER\_CLASS, PER\_INSTANCE, PER\_CLONE) to control how external resources (e.g., clients) are shared and automatically closed (via AutoCloseable) across DoFn clones. Existing async DoFns (\GuavaAsyncDoFn\, \JavaAsyncDoFn\, \GuavaAsyncLookupDoFn\, \JavaAsyncLookupDoFn\) are retained for single-element async operations, and \FutureHandlers\ provides a unified abstraction for Guava ListenableFuture and Java CompletableFuture callbacks with a 10-minute default timeout. A new \UnmatchedRequestException\ is added to surface errors when batch responses do not match requests, and \PipeDoFn\/\ProcessUtil\ support external command execution with setup/teardown commands.
scio-core/src/main/java/com/spotify/scio/transforms · high confidence
New common options and utilities for Beam examples
Added new Java interfaces and utility classes in the scio-examples module to standardize configuration and resource management for Beam examples. ExampleBigQueryTableOptions, ExamplePubsubTopicOptions, and ExamplePubsubTopicAndSubscriptionOptions provide structured pipeline options with sensible defaults (e.g., deriving table names from job names) for BigQuery and Pub/Sub resources. ExampleOptions adds generic example settings like job retention and worker counts. ExampleUtils implements setup and teardown logic for these external resources, including automatic creation/deletion of Pub/Sub topics and subscriptions and BigQuery tables, along with pipeline cancellation hooks.
scio-examples/src/main/java/org/apache/beam/examples/common · high confidence
New cookbook examples for BigQuery, joins, and distinct operations
Added a set of new cookbook examples in the \scio-examples\ module to demonstrate common Scio patterns. These include \BigQueryTornadoes\ and \StorageBigQueryTornadoes\ for reading and writing BigQuery data (using both standard and Storage APIs), \JoinExamples\ covering left outer, side-input, hash, and skewed joins, \CombinePerKeyExamples\ and \MaxPerKeyExamples\ for aggregation, and \DistinctExample\ and \DistinctByKeyExample\ for deduplication.
scio-examples/src/main/scala/com/spotify/scio/examples/cookbook · high confidence
New dynamic IO syntax package for dynamic destinations
A new \com.spotify.scio.io.dynamic\ package has been introduced to support dynamic file IO destinations. This package acts as the entry point for the dynamic syntax, allowing users to import \com.spotify.scio.io.dynamic.\_\ to access the \AllSyntax\ extensions required for writing to dynamic file paths.
scio-core/src/main/scala/com/spotify/scio/io/dynamic · high confidence
New extra examples for Annoy, Avro, Beam, Bigtable, Cloud SQL, DistCache, DoFn, JavaConverters, Metrics, Protobuf, Redis, SafeFlatMap, SideInput/Output, GZip, Stateful, TableRow JSON, Tap, and Template usage
The \scio-examples/src/main/scala/com/spotify/scio/examples/extra\ directory now includes a comprehensive set of new example files demonstrating various Scio capabilities. These include \AnnoyExamples\ for vector similarity search, \AvroInOut\ and \ProtobufExample\ for schema-based I/O, \BeamExample\ for mixing Beam Java SDK transforms with Scio, \BigtableExample\ for Google Cloud Bigtable operations, \CloudSqlExample\ for JDBC connections to Cloud SQL, \DistCacheExample\ for distributed caching, \DoFnExample\ for custom Beam DoFns, \JavaConvertersExample\ for Beam IO options, \MetricsExample\ for counters and gauges, \RedisExamples\ for Redis read/write/lookup, \SafeFlatMapExample\ for error handling, \SideInOutExample\ for side inputs/outputs, \SingleGZipFileExample\ for compressed output, \StatefulExample\ for stateful processing, \TableRowJsonInOut\ for BigQuery JSON, \TapOutputExample\ and \TapsExample\ for intermediate data handling, and \TemplateExample\ for Dataflow templates.
scio-examples/src/main/scala/com/spotify/scio/examples/extra · high confidence
New hash-based join and approximate filter APIs
The \scio-core\ values package now includes a new \hash\ package with implicit classes to create \ApproxFilter\ (e.g., Bloom filters) from \Iterable\ and \SCollection\ types, including methods to create them as \SideInput\s for use in joins. Additionally, \PairHashSCollectionFunctions\ provides hash-based join operations (\hashJoin\, \hashLeftOuterJoin\, \hashFullOuterJoin\, \hashIntersectByKey\) that replicate a small right-hand-side collection to all workers as a side input, enabling efficient joins when the right side fits in memory.
scio-core/src/main/scala/com/spotify/scio/values · high confidence
New macro utilities for Java bean validation, Magnolia integration, and system property registration
The scio-macros module introduces three new macro-based capabilities: IsJavaBean validates that a type is a Java bean by checking for matching getter and setter methods at compile time; MagnoliaMacros provides a custom derivation path for Magnolia1 that strips non-serializable annotations and cleans up outer references to ensure Coder serialization works correctly; and SysPropsMacros allows objects to automatically register their variables as system properties by mixing in the SysProps trait via a macro annotation.
scio-macros/src/main/scala/com/spotify/scio · high confidence
New scio-extra utilities for CSV, time-series iterators, collections, and Breeze
This change introduces several new capabilities in scio-extra: a CSV IO package (including dynamic destination support via saveAsDynamicCsvFile), iterator windowing utilities (timeSeries with fixed, session, and sliding windows), collection helpers (top and topByKey), and Breeze Semigroups for aggregation. It also adds Sparkey map/set base traits, Sparkey coders, Annoy URI handling, BigQuery Avro-to-TableRow/Schema converters, and ZetaSketch HyperLogLog++ approximate distinct count syntax for SCollection.
scio-extra/src/main/scala · high confidence
New subprocess execution example with C++ integration
Added a new example in the \subprocess\ package that demonstrates how to execute external C++ binaries (Echo.cc, EchoAgain.cc) from within a Beam pipeline. The change introduces a complete support layer including \SubProcessPipelineOptions\ for configuration (source path, concurrency, timeouts), a \SubProcessKernel\ to manage process execution and I/O redirection, and utility classes (\FileUtils\, \CallingSubProcessUtils\) to handle downloading executables from GCS to the worker, managing concurrency via semaphores, and uploading log files back to storage.
scio-examples/src/main/java/org/apache/beam/examples/subprocess · high confidence
New utility classes for writing game data to BigQuery
Added three new utility classes to the game examples: GameConstants, which defines shared constants like timestamp attributes and date formatters; WriteToBigQuery, a generic PTransform that converts input collections into BigQuery table rows using configurable field definitions and lambda functions; and WriteWindowedToBigQuery, a subclass of WriteToBigQuery that preserves windowing context during row generation, enabling writes that require access to window information.
scio-examples/src/main/java/org/apache/beam/examples/complete/game/utils · high confidence
Project initialization and repository structure setup
The repository has been initialized with essential project infrastructure files. This includes a comprehensive .gitignore to exclude build artifacts and IDE files, configuration for the Scala Steward dependency bot (.scala-steward.conf) to manage updates and pin specific versions, and formatting/linting rules for Scalafmt (.scalafmt.conf) and Scalafix (.scalafix.conf). Additionally, the project now provides a Code of Conduct (CODE\_OF\_CONDUCT.md), Contribution Guidelines (CONTRING.md), a Security Policy (SECURITY.md), and an updated README.md that reflects the current Scio API, features, and quick-start instructions.
(repo-wide) · high confidence
Re-enable scio-examples with new Java complete examples
The scio-examples module is re-enabled, introducing a new set of Java-based 'Complete' examples in the org.apache.beam.examples.complete package. This includes a README documenting the examples and new source files such as StreamingWordExtract.java, which demonstrates a streaming pipeline reading text, tokenizing words, and writing to BigQuery, and TfIdf.java, which computes TF-IDF scores using joins and side inputs. These additions provide users with reference implementations for complex data processing tasks like streaming word extraction and search table generation.
scio-examples/src/main/java/org/apache/beam/examples/complete · high confidence
Scala 2.13 REPL implementation
Added Scala 2.13-specific REPL support by introducing \ScioGenericRunner\ and \ILoop\ compatibility layers. This enables the Scio REPL to function correctly on Scala 2.13 by configuring the classloader, REPL settings, and interpreter initialization specific to that version.
scio-repl/src/main/scala-2.13 · high confidence
Support for sharded JDBC reads
Users can now perform sharded reads from JDBC tables to improve performance on large datasets. This change introduces a new \JdbcShardedSource\ that automatically determines the range of a specified shard column, partitions the data into multiple shards, and executes parallel queries. The implementation supports sharding by various data types, including numeric ranges (Long, Int, Double, etc.) and string-based identifiers (UUIDs, Base64, and SQL Server uniqueidentifiers), allowing efficient parallelization of table scans.
scio-jdbc/src/main/scala/com/spotify/scio/jdbc/sharded · high confidence
Support for writing Avro and Protobuf files to dynamic destinations
Users can now write SCollection data to Avro or Protobuf files where the output path is determined dynamically at runtime via a destination function. This change introduces \saveAsDynamicAvroFile\ for both SpecificRecord and GenericRecord types, as well as \saveAsDynamicProtobufFile\, allowing flexible file routing within the scio-avro module.
scio-avro/src/main/scala/com/spotify/scio/avro/dynamic · high confidence
Support for writing TensorFlow Example records to dynamic Parquet destinations
Users can now save collections of TensorFlow Example records to dynamic file destinations using the new \saveAsDynamicParquetExampleFile\ method. This feature introduces a \ParquetExampleSink\ that writes data using the TensorFlow Parquet writer, allowing users to specify a schema, compression codec, and custom metadata for each output file. The implementation is exposed via the \com.spotify.scio.parquet.tensorflow.dynamic\ package, which provides syntax extensions for SCollection operations.
scio-parquet/src/main/scala/com/spotify/scio/parquet/tensorflow/dynamic · high confidence
Support for writing to dynamic BigQuery table destinations
Users can now write Scio collections to BigQuery tables determined at runtime using the new \saveAsBigQuery\ methods in the \com.spotify.scio.bigquery.dynamic\ package. This feature introduces \DynamicBigQueryOps\, \DynamicTableRowBigQueryOps\, and \DynamicTypedBigQueryOps\ (the latter deprecated in favor of the Magnolify API), allowing users to specify a function that maps elements to target \TableDestination\s. The implementation leverages Beam's \DynamicDestinations\ and supports configuration for write/create dispositions and extended error info, returning a \ClosedTap\ that includes side outputs for write results.
scio-google-cloud-platform/src/main/scala/com/spotify/scio/bigquery/dynamic · high confidence
TensorFlow Example Parquet I/O support
Added Java classes in scio-parquet to read and write Apache Parquet files containing TensorFlow Example protos. The new TensorflowExampleParquetInputFormat, OutputFormat, Reader, and Writer components handle schema conversion between TensorFlow metadata and Parquet schemas, support column projection pushdown, and allow storing TensorFlow schemas in Parquet metadata for schema-aware reconstruction at read time.
scio-parquet/src/main/java · high confidence
TensorFlow TFRecord I/O and syntax support
This change introduces the core infrastructure for reading and writing TensorFlow TFRecord files within Scio. It adds a Scala-based TFRecordCodec for handling record parsing and compression detection, alongside a Java FileBasedSink to integrate with Apache Beam's file writing pipeline. On the user-facing side, new syntax extensions are provided for ScioContext (to read TFRecord, Example, and SequenceExample files, including schema-aware variants via DistCache) and SCollection (to save data as TFRecord files with configurable compression, sharding, and prefix/suffix options).
scio-tensorflow/src/main · high confidence
Type-safe Parquet I/O syntax extensions for ScioContext and SCollection
This change introduces type-safe syntax extensions for reading and writing Parquet files in Scio. Users can now use \typedParquetFile\ on \ScioContext\ to read Parquet data directly into case classes, supporting optional filtering via \FilterPredicate\ and Hadoop \Configuration\. Similarly, \saveAsTypedParquetFile\ on \SCollection\ allows writing case classes to Parquet with configurable parameters including shard count, compression, suffix, prefix, temporary directory, filename policy, and custom metadata. These methods leverage Magnolify's \ParquetType\ for automatic schema derivation and encoding.
scio-parquet/src/main/scala/com/spotify/scio/parquet/types/syntax · high confidence
Removals
Removal of BigQuery TableRow helper classes and implicits
The BigQuery module has removed the \RichTableRow\ class and the \Implicits\ trait, along with the \package.scala\ definitions that provided implicit conversions and a \TableRow\ type alias. Users can no longer rely on the automatic enrichment of \TableRow\ objects with convenient type-casting methods (such as \getBoolean\, \getInt\, etc.) or the specific \TableRow\ type alias previously exposed in the \com.spotify.cloud.bigquery\ package.
bigquery · high confidence
Removal of core Scala Dataflow library components
The core Scala Dataflow library has been removed from the codebase. This change deletes the \com.spotify.cloud.dataflow\ package and its sub-packages, eliminating the \SCollection\ wrapper, \DataflowContext\ orchestration, argument parsing (\Args\), and various utility functions for side inputs, side outputs, and accumulators. It also removes the associated Kryo-based coders (\KryoAtomicCoder\, \AvroSerializer\), the \RichCoderRegistry\, and the testing infrastructure (\JobTest\, \TestDataManager\). Additionally, the Java \FloatCoder\ class has been deleted.
core/src/main · high confidence
Removal of legacy Java Dataflow SDK examples
The \examples/src/main/java\ directory has been cleared of several legacy Java example pipelines that relied on the deprecated \com.google.cloud.dataflow.sdk\ library. Specifically, the files \AutoComplete.java\, \BigQueryTornadoes.java\, \CombinePerKeyExamples.java\, \DatastoreWordCount.java\, \DeDupExample.java\, \FilterExamples.java\, \JoinExamples.java\, \MaxPerKeyExamples.java\, and \PubsubFileInjector.java\ have been deleted. These examples demonstrated concepts such as BigQuery I/O, Datastore integration, Pub/Sub injection, and various transforms (e.g., \Combine.perKey\, \RemoveDuplicates\, \CoGroupByKey\) using the older SDK API, and their removal indicates a cleanup of outdated sample code.
examples/src/main/java · high confidence
Removed Scala Dataflow examples
Deleted all Scala example programs in examples/src/main/scala, including WordCount, WindowingWordCount, AutoComplete, BigQueryTornadoes, DistCacheExample, and others. These files are no longer available for users to reference or run.
examples/src/main/scala · high confidence
Behavioural changes
Add Scala 2.12 compatibility shims for REPL and CSV reading
The Scio REPL now includes compatibility wrappers for Scala 2.12 to support newer library APIs and internal interpreter changes. This adds a \CsvReaderOps\ implicit class that exposes an \iterator\ method on \kantan.csv.CsvReader\ instances, ensuring consistent iteration behavior. It also introduces a custom \ILoop\ abstraction that wraps the Scala interpreter, managing initialization commands and output flushing, which supports the refactored REPL lifecycle.
scio-repl/src/main/scala-2.12/com/spotify/scio/repl/compat · high confidence
Added Kryo serialization for Bigtable MutateRowsException
The Scio Google Cloud Platform module now includes a dedicated Kryo registrar and serializer for \MutateRowsException\. This change ensures that Bigtable mutation errors are correctly serialized during distributed processing, preventing them from being lost or incorrectly handled as generic exceptions when used as causes in other throwables.
scio-google-cloud-platform/src/main/scala/com/spotify/scio/coders · high confidence
Added Kryo serializer for Java Traversable collections in Scala 2.13
A new Kryo serializer (JTraversableSerializer) and associated collection builder factories have been added to the Scala 2.13-specific code path to handle serialization of Java Traversable, Iterable, and Collection types. This change ensures that these collection types are correctly serialized and deserialized using Kryo within the Scio coders system for Scala 2.13, addressing compatibility requirements for the updated Scala version.
scio-core/src/main/scala-2.12/com/spotify/scio/coders/instances/kryo, scio-core/src/main/scala-2.13/com/spotify/scio/coders/instances/kryo · high confidence
Added URL redirection templates for examples and Scaladoc
New HTML template files have been added to the site documentation structure to handle URL redirections. The \examples.st\ template now redirects users from the root examples path to \examples/index.html\, and the \scaladoc.st\ template redirects Scaladoc access to the specific API documentation path \api/com/spotify/scio/index.html\. These changes ensure that legacy or root URLs correctly forward users to the updated documentation locations.
_site/src/main/paradox/\template · high confidence
Added immutable Java-to-Scala Map wrappers
Introduced JMapWrapper in the Scala 2.12 utility package to provide immutable wrappers for java.util.Map instances. This addresses the inconsistency where standard conversions return mutable maps, ensuring that Beam API interactions use idiomatic, immutable Scala collections instead.
scio-core/src/main/scala-2.12/com/spotify/scio/util · high confidence
Avro schema instances and Protobuf-to-Avro utilities moved to scio-avro
The Avro schema implementation (AvroInstances) and Protobuf utility functions (ProtobufUtil) have been relocated from scio-core to the scio-avro module. This change introduces implicit Schema instances for Avro SpecificRecord and GenericRecord types, along with utilities to convert Protobuf messages into Avro GenericRecords, making these capabilities available only when the scio-avro dependency is included.
scio-avro/src/main/scala/com/spotify/scio/avro/schemas, scio-avro/src/main/scala/com/spotify/scio/protobuf · high confidence
Build infrastructure refactored and upgraded to SBT 1.12
The build system has been significantly restructured and modernized. The legacy Build.scala file was removed and replaced with modular Scala files (BuildCredentials, CheckBeamDependencies, Exclude, JavaOptions, ScalacOptions, SoccoIndex) that handle credential detection, dependency conflict checking, exclusion rules, and compiler options. The SBT version was upgraded from 0.13.8 to 1.12.13, and plugins were updated to their latest versions (e.g., sbt-typelevel 0.8.7, sbt-scalafix 0.14.7, sbt-assembly 2.3.1). A new CheckBeamDependencies plugin was added to detect and warn about version mismatches between project dependencies and Apache Beam's expected versions, and exclusion rules were refined to prevent conflicts with libraries like gcsio, metrics-core, and Envoy control plane APIs.
project · high confidence
Centralized example data paths for Scio examples
The Scio examples now use a centralized \ExampleData\ object to define input file paths (such as Shakespeare texts, Wikipedia edits, and traffic sensor data) and BigQuery table references. This change consolidates data location definitions, ensuring that example programs reference consistent, public Google Cloud Storage and BigQuery resources rather than hardcoding paths individually.
scio-examples/src/main/scala/com/spotify/scio/examples/common · high confidence
Customized Scala REPL execution for Scio
The Scio REPL now uses a custom runner (ScioGenericRunner) that configures the Scala compiler environment specifically for Scio usage. This includes automatically adding the Scio classpath, enabling the paradise macro plugin (required for BigQuery macros) by detecting it in the classpath or using the assembly jar, and enforcing synchronous execution and class-based REPL output to resolve known issues. This ensures that macro expansion and class loading work correctly within the interactive Scio shell.
scio-repl/src/main/scala-2.12/com/spotify/scio/repl · high confidence
Elasticsearch integration switches to Jackson JSON serialization
The Elasticsearch IO module now uses Jackson as the default JSON serialization backend, replacing the previous implementation. This change introduces a new JsonpMapperFactory interface and a default JacksonJsonpMapper configuration, ensuring that Scala and Java types (including java.time) are correctly serialized and deserialized. The update also includes new CoderInstances to handle BulkOperation serialization via Kryo and adds corresponding tests to verify round-trip encoding and Jackson-based mapping behavior.
scio-elasticsearch/common · high confidence
Improved local GCS authentication for Parquet reads
The Parquet module now includes a new utility (GcsConnectorUtil) that automatically configures Hadoop credentials when reading GCS paths locally. It searches for credentials in the GOOGLE\_APPLICATION\_CREDENTIALS environment variable or the standard application default credentials file, falling back to an unauthenticated mode for unit testing. This ensures that local path validation works correctly without interfering with credentials in Dataflow workers.
scio-parquet/src/main/scala/com/spotify/scio/parquet · high confidence
Integrate Avro datum factory for improved CharSequence handling
The scio-avro module now integrates a custom Avro datum factory that forces CharSequence implementations to String, resolving issues where the default Utf8 type failed to implement equals for joins or SMB keys. This change includes new internal components (AvroDatumFactory, AvroFileStorage, AvroSysProps) to manage datum reading/writing and system properties, and patches logical type conversions to ensure compatibility with Avro 1.8 generated code.
scio-avro/src/main/scala/com/spotify/scio/avro · high confidence
Integrate Avro datum factory in scio-avro
The scio-avro module has been refactored to use the Apache Beam AvroDatumFactory API for reading and writing Avro data. This change introduces new IO implementations (AvroTypedIO, AvroMagnolifyTypedIO, ObjectFileIO) that delegate to a unified GenericRecordIO, allowing for more flexible and standard-compliant Avro serialization. Users benefit from improved type safety and consistency in Avro I/O operations, with the legacy AvroTyped API deprecated in favor of the new typed IOs.
repository · high confidence
Introduce module-specific Kryo registrar and Avro serialization utilities
The scio-avro module now includes a dedicated Kryo registrar (AvroKryoRegistrar) that explicitly registers serializers for Avro GenericRecord and SpecificRecord types, ensuring consistent serialization behavior within the module. Additionally, new utility classes (AvroBytesUtil, AvroSerializer) and coder implementations (SlowGenericRecordCoder, SpecificFixedCoder) have been added to handle Avro-specific encoding, decoding, and schema caching, supporting the integration of Avro datum factories and improving the robustness of Avro data handling in Scio pipelines.
scio-avro/src/main/scala/com/spotify/scio/coders · high confidence
Introduce new Args parser and Beam-based Metrics API
This change replaces the previous argument handling with a new \Args\ class that parses command-line flags into typed values (strings, integers, booleans, etc.) and exposes the internal argument map. It also introduces a new metrics system via \ScioMetrics\ and \ScioResult\, allowing users to create and retrieve Beam Counter, Distribution, and Gauge metrics from pipeline results, replacing the older accumulator-based approach.
scio-core/src/main/scala/com/spotify/scio · high confidence
Migration rules for Scio API changes (v0.7–v0.12)
The scalafix migration rules have been reorganized into their proper packages to support automated upgrades for Scio versions 0.7 through 0.12. These rules handle specific behavioral changes and deprecations, including: replacing \BigQueryClient\ with the new \BigQuery\ client API (v0.7); updating system property access to use typed \CoreSysProps\ and \BigQuerySysProps\ (v0.7); adding missing imports for Avro (v0.7); wrapping BigQuery table references in \Table.Spec\ or \Table.Ref\ (v0.8); renaming join methods to use \Outer\ suffixes and \rhs\ parameter names (v0.8); replacing \sc.close()\ with \sc.run()\ (v0.8); updating \ScioIO\ write signatures to return \Tap\ directly (v0.8); fixing syntax imports for JDBC and AsyncLookup (v0.8); changing \saveAsTfExampleFile\ to \saveAsTfRecordFile\ (v0.8); updating Coder propagation to use curried Ordering (v0.10); migrating BigQuery save methods to \saveAsBigQueryTable\ with \Table.Ref\ and \toTableSchema\ (v0.12); and refactoring PubSub IO to use the new \read\/\write\ pattern with \PubsubIO\ params (v0.12).
(repo-wide) · high confidence
New Parquet read implementation with SplittableDoFn and configuration options
The Parquet read module has been refactored to introduce a new \ParquetRead\ API that supports both legacy HadoopFormatIO and a new SplittableDoFn-based read path. The new implementation defaults to using SplittableDoFn for improved performance when Dataflow Runner V2 is enabled, while allowing users to opt-in or opt-out via configuration. It also introduces granular control over split sizes (file vs. row group) and filter application (page vs. record level), and adds support for reading TensorFlow Example protos alongside standard Avro and typed Scala reads.
scio-parquet/src/main/scala/com/spotify/scio/parquet/read · high confidence
New macro-based Kryo registration and improved coder fallback warnings
The coder macro system now includes a \@KryoRegistrar\ annotation that automatically mixes the \AnnotatedKryoRegistrar\ trait into classes extending \IKryoRegistrar\ (requiring names ending in 'KryoRegistrar'), streamlining custom Kryo registration. Additionally, the fallback coder logic has been enhanced to provide more specific warnings for discouraged types like \GenericRecord\ and to respect a \MacroSettings.showCoderFallback\ flag, allowing users to control the verbosity of Kryo fallback messages during compilation.
scio-macros/src/main/scala/com/spotify/scio/coders · high confidence
Refactored Avro syntax with Magnolify integration and deprecated legacy APIs
The Avro syntax layer in scio-avro has been reorganized into dedicated \ScioContextSyntax\ and \SCollectionSyntax\ traits, introducing a new \typedAvroFileMagnolify\ method that leverages the Magnolify library for type-safe Avro serialization. The previous typed Avro methods (\typedAvroFile\ and \saveAsTypedAvroFile\) are now deprecated in favor of this Magnolify-based approach. Additionally, the API now supports passing a custom \AvroDatumFactory\ for both reading and writing operations on GenericRecord and SpecificRecord, allowing for more flexible serialization control, while the experimental \parseAvroFile\ method remains available for untyped parsing.
scio-avro/src/main/scala/com/spotify/scio/avro/syntax · high confidence
Refactored coder internals and require explicit imports for Kryo fallback
The coder subsystem has been restructured to improve error reporting and reduce memory overhead. New internal utilities in BeamCoders safely unwrap nested Beam coders (including Zstd compression) to extract key-value and tuple coders from SCollection and PCollection instances, ensuring informative errors if extraction fails. CoderDerivation now uses Magnolia for auto-deriving coders for case classes and sealed traits, with specific handling to prevent issues with inner classes. Custom coder implementations (SingletonCoder, DisjunctionCoder, RecordCoder) have been added to optimize serialization for specific patterns. Additionally, the implicit Kryo fallback coder is no longer automatically available; users must now explicitly import it from com.spotify.scio.coders.kryo to use Kryo serialization as a fallback.
scio-core/src/main/scala/com/spotify/scio/coders · high confidence
Refactored in-memory and file storage implementations
The in-memory sink and file storage logic in scio-core have been refactored to improve testability and robustness. The InMemorySink now explicitly requires a test context and uses a TrieMap to store SCollection data, ensuring empty collections return an iterable rather than throwing exceptions. The new FileStorage object replaces previous implementations by leveraging Apache Beam's FileSystems API for listing and reading files, adding support for compressed text files via CompressorStreamFactory, and introducing an isDone method that accurately tracks write progress by parsing shard patterns.
scio-core/src/main/scala/com/spotify/scio/io · high confidence
Registration of Scalafix migration rules for versions 0.7 through 0.15
The service provider configuration for Scalafix rules has been updated to include a comprehensive list of migration rules covering Apache Beam versions 0.7.0 through 0.15.0. This registration enables automatic code fixes for deprecated APIs and structural changes across multiple major releases, including specific migrations for Avro I/O, BigQuery client refactoring, TensorFlow imports, Bigtable I/O, and join name consistency.
scalafix/rules/src/main/resources · high confidence
Scalafix migration rules for Scio 0.14 API changes
This update adds a suite of Scalafix rules for the 0.14 migration, located in the \scalafix/input-0\_14\ and \scalafix/output-0\14\ directories. The \FixAvroCoder\ rule automatically imports Avro coders from \com.spotify.scio.avro.\\ and updates \Coder\ companion object calls to direct imports. The \FixGenericAvro\ rule reorders arguments in \saveAsAvroFile\ to match the new signature. The \FixQuery\ rule migrates deprecated \HasQuery.query\ calls to \queryRaw\. The \FixSMBCharSequenceKey\ rule replaces \CharSequence\ key types with \String\ in Sort-Merge-Bucket operations. Additionally, \FixLogicalTypeSupplier\ removes deprecated Parquet configuration options, \FixAvroSchemasPackage\ updates schema imports, and \FixDynamicAvro\ adds necessary dynamic Avro imports.
_scalafix/input-0\_14, scalafix/output-0\14 · high confidence
Streaming API aliases moved to dedicated package
The streaming-related type aliases for Apache Beam's \AccumulationMode\ (specifically \ACCUMULATING\_FIRED\_PANES\ and \DISCARDING\_FIRED\PANES\) have been moved into a new dedicated \com.spotify.scio.streaming\ package object. Users should now import these aliases from \com.spotify.scio.streaming.\\ instead of the previous location, simplifying access to streaming-specific Beam windowing configurations.
scio-core/src/main/scala/com/spotify/scio/streaming · high confidence
Updated scalafix rules for v0.13.0 API migrations
This release includes updated scalafix rules for version 0.13.0 to handle two distinct API changes. First, the \FixSkewedJoins\ rule now transforms calls to \skewedJoin\, \skewedLeftOuterJoin\, and \skewedFullOuterJoin\ to use the new named parameters (e.g., \hotKeyMethod\, \cmsEps\, \cmsDelta\) and the \HotKeyMethod.Threshold\ type, replacing the previous positional arguments. Second, the \FixTaps\ rule updates various tap constructors (such as \SpecificRecordTap\, \ObjectFileTap\, \GenericRecordTap\, \TextTap\, and \TFRecordFileTap\) to explicitly accept \ReadParam\ instances (e.g., \AvroIO.ReadParam()\, \TextIO.ReadParam()\, \TFRecordIO.ReadParam()\) as additional arguments.
_scalafix/input-0\_13, scalafix/output-0\13 · high confidence
Updated tuple and Java collection wrapper coders for Scala 2.12 and 2.13
The coder instances for Scala 2.12 and 2.13 have been updated to use specialized TupleCoders (Tuple2, Tuple3, etc.) instead of the previous PairCoder approach, which reduces memory footprint and ensures proper getCoderArguments generation for tuple types. Additionally, the JavaCollectionWrappers module now correctly identifies Scala 2.13's specific wrapper class names (e.g., \JavaCollectionWrappers$JIterableWrapper\ vs \Wrappers$JIterableWrapper\), fixing a regression where JIterableWrapper coders failed to match in Scala 2.13.
scio-core/src/main/scala-2.12/com/spotify/scio/coders/instances, scio-core/src/main/scala-2.13/com/spotify/scio/coders/instances · high confidence
Updated v0.15 migration rules for Bigtable, Datastore, and TensorFlow imports
The v0.15 scalafix migration rules have been updated to handle specific API changes. For Bigtable, the \FixBigtableIO\ rule now migrates calls using \Seq\[ByteKeyRange\]\ to the \BTOptions\ overload, correctly preserving named arguments (e.g., \keyRanges = ...\) during the transformation. For Datastore, the \FixDatastoreIO\ rule now adds the required type parameter \\[Entity\]\ to \DatastoreIO\ instantiations. Additionally, the \FixTensorflowProtoImports\ rule updates imports from the nested \proto.example\ package to the top-level \proto\ package for TensorFlow classes.
_scalafix/input-0\_15, scalafix/output-0\15 · high confidence
Vendor Socco plugin resources
The socco-plugin now includes its own bundled frontend assets, specifically the 'tooltips' library (version 0.1.0) and associated CSS stylesheets (base.css, tooltips.css). This change vendors the tooltip JavaScript and CSS directly into the plugin's resources, ensuring the plugin's UI components (such as code comments and identifiers) have the necessary styling and interactive tooltip behavior without relying on external runtime dependencies for these specific UI elements.
socco-plugin · high confidence
Fixes
Add implicit coder for BigQuery TableRow
Users can now seamlessly work with BigQuery TableRow objects in Scio pipelines, as an implicit Coder instance is provided via the new CoderInstances trait. This change ensures that TableRow data can be correctly serialized and deserialized during distributed processing without requiring manual coder configuration.
scio-google-cloud-platform/src/main/scala/com/spotify/scio/bigquery/coders, scio-google-cloud-platform/src/main/scala/com/spotify/scio/bigquery/instances · high confidence
Test coverage
Add RuleSuite test entry point for scalafix tests; Added Apache 2.0 license file to test resources; Added Avro schema definitions for SMB tests; Added Avro schema for integration tests; Added JMH benchmarks for Scio performance testing; Added Java unit tests for scio-smb core components; Added Parquet test utilities for predicates and projections; Added SMB integration tests and benchmarks; Added TestUtil helper for generating unique test IDs; Added comprehensive test coverage for Avro IO, Coder, and Type system; Added comprehensive tests for BigQuery type conversion and schema generation; Added integration and unit tests for Scio examples; Added integration test for Parquet read operations; Added integration tests for Annoy, Sparkey, and Voyager side inputs; Added integration tests for Avro and JDBC IO; Added integration tests for BigQuery Storage API and BigQuery Type annotations; Added integration tests for BigQuery client, IO, typed writes, and partition utilities; Added integration tests for BigQuery, Cassandra, Spanner, and DistCache; Added integration tests for Dataflow runner result handling; Added integration tests for Elasticsearch IO with v8 client; Added integration tests for IcebergIO with REST catalog and dynamic table properties; Added integration tests for Neo4j IO operations; Added integration tests for ScioContext configuration and job graph generation; Added integration tests for SortMergeBucket parity; Added test coverage for Bigtable integration components; Added test coverage for extra Scio examples; Added test coverage for scio-extra modules; Added test fixtures and test suites for scio-core; Added test resources for TensorFlow SavedBundleModel support; Added test support classes for JMH benchmarks; Added tests for BigQuery validation type overrides; Added tests for GCP Kryo serialization of Bigtable exceptions; Added tests for JDBC I/O and connection handling; Added tests for JDBC sharding logic and string codecs; Added tests for Parquet SMB components; Added tests for PubsubIO subscription, topic, and timestamp attributes; Added tests for SmbIO and version parity; Added tests for TensorFlow IO and metadata schema support; Added tests for cookbook examples; Added tests for game example components; Added tests for scio-parquet TensorFlow, Avro, and dynamic IO features; Added unit tests for Beam cookbook examples; Added unit tests for Bigtable utilities and bulk writer; Added unit tests for Cassandra IO and DataType serialization; Added unit tests for game example pipeline components; Initial test suite for scio-google-cloud-platform BigQuery, Datastore, and Spanner IOs; New testing utilities and assertions in scio-test/core; Removal of legacy Dataflow coder test suite; Removed deprecated PipelineTest trait; Removed example job tests; Removed legacy SCollection test suite; Removed unused test matchers.
Dependencies
Upgrade to Apache Beam 2.76.0 and sync ecosystem dependencies
The build system has been upgraded to Apache Beam 2.76.0, which drives updates to the entire dependency tree including Hadoop 3.4.3, Spark 3.5.0, and Flink 1.19.0. Core libraries have been bumped to align with the new Beam release: Guava to 33.1.0-jre, Jackson to 2.18.8, and Protobuf to 4.33.2. Additionally, the project now uses the Google Cloud Libraries BOM (version 26.85.0) for consistent GCP client versions and updates the BigQuery API client to v2-rev20260612.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Score
- CAI 48 → 71 (+22.6)
- Rubric changed (rubric-2026.08.15 → rubric-2026.09.15) — scores are not directly comparable.
Lenses
- Code Health 93 → 86 (-6.4)
- Architecture 94 → 99 (+5.3)
- Maturity 46 → 60 (+14.5)
- Readiness 28 → 86 (+58.3)
- Security 80 → 72 (-8.2)
Resolved (34)
- Coverage not measured — test suite did not build
- Dimension evaluation failed
- Duplicated block (11 lines × 2) (scio-core/src/main/java/com/spotify/scio/util/RemoteFileUtil.java)
- Duplicated block (13 lines × 2) (scio-smb/src/main/java/org/apache/beam/sdk/extensions/smb/SortedBucketIO.java)
- Duplicated block (14 lines × 2) (scio-examples/src/main/java/org/apache/beam/examples/complete/game/LeaderBoard.java)
- Duplicated block (14 lines × 2) (scio-grpc/src/main/java/com/spotify/scio/grpc/GrpcBatchDoFn.java)
- Duplicated block (15 lines × 3) (scio-smb/src/main/java/org/apache/beam/sdk/extensions/smb/SortedBucketSource.java)
- Duplicated block (177 lines × 3) (scio-examples/src/main/java/org/apache/beam/examples/common/ExampleUtils.java)
- Duplicated block (19 lines × 2) (scio-elasticsearch/common/src/main/java/org/apache/beam/sdk/io/elasticsearch/ElasticsearchIO.java)
- Duplicated block (21 lines × 2) (scio-smb/src/main/java/org/apache/beam/sdk/extensions/smb/JsonSortedBucketIO.java)
- Duplicated block (21 lines × 2) (scio-smb/src/main/java/org/apache/beam/sdk/extensions/smb/TensorFlowBucketIO.java)
- Duplicated block (27 lines × 2) (scio-smb/src/main/java/org/apache/beam/sdk/extensions/smb/ParquetBucketMetadata.java)
- Duplicated block (28 lines × 2) (scio-smb/src/main/java/org/apache/beam/sdk/extensions/smb/SortedBucketSink.java)
- Duplicated block (30 lines × 2) (scio-examples/src/main/java/org/apache/beam/examples/subprocess/kernel/SubProcessKernel.java)
- Duplicated block (46 lines × 2) (scio-smb/src/main/java/org/apache/beam/sdk/extensions/smb/AvroSortedBucketIO.java)
- Duplicated block (5 lines × 2) (scio-examples/src/main/java/org/apache/beam/examples/complete/game/GameStats.java)
- Duplicated block (5 lines × 2) (scio-examples/src/main/java/org/apache/beam/examples/subprocess/kernel/SubProcessKernel.java)
- Duplicated block (5 lines × 2) (scio-smb/src/main/java/org/apache/beam/sdk/extensions/smb/SortedBucketIO.java)
- Duplicated block (6 lines × 2) (scio-examples/src/main/java/org/apache/beam/examples/complete/TfIdf.java)
- Duplicated block (6 lines × 2) (scio-examples/src/main/java/org/apache/beam/examples/snippets/Snippets.java)
- …and 14 more
New (345)
- ClassTooLong: ElasticsearchIO (scio-elasticsearch/common/src/main/java/org/apache/beam/sdk/io/elasticsearch/ElasticsearchIO.java)
- ClassTooLong: SortedBucketIO (scio-smb/src/main/java/org/apache/beam/sdk/extensions/smb/SortedBucketIO.java)
- ClassTooLong: SortedBucketScioContext (scio-smb/src/main/scala/com/spotify/scio/smb/syntax/SortMergeBucketScioContextSyntax.scala)
- ClassTooLong: SortedBucketSink (scio-smb/src/main/java/org/apache/beam/sdk/extensions/smb/SortedBucketSink.java)
- ClassTooLong: SortedBucketSource (scio-smb/src/main/java/org/apache/beam/sdk/extensions/smb/SortedBucketSource.java)
- ClassTooLong: SortedBucketTransform (scio-smb/src/main/java/org/apache/beam/sdk/extensions/smb/SortedBucketTransform.java)
- ConverterProvider.fromAvroInternal (cyclomatic 22) (scio-google-cloud-platform/src/main/scala/com/spotify/scio/bigquery/types/ConverterProvider.scala)
- Duplicated block (10 lines × 16) (scio-smb/src/main/scala/com/spotify/scio/smb/util/SMBMultiJoin.scala)
- Duplicated block (10 lines × 2) (scio-core/src/main/scala/com/spotify/scio/util/FunctionsWithWindowedValue.scala)
- Duplicated block (10 lines × 2) (scio-core/src/main/scala/com/spotify/scio/values/PairHashSCollectionFunctions.scala)
- Duplicated block (10 lines × 2) (scio-examples/src/main/scala/com/spotify/scio/examples/complete/game/GameStats.scala)
- Duplicated block (10 lines × 2) (scio-smb/src/main/scala/com/spotify/scio/smb/syntax/SortMergeBucketScioContextSyntax.scala)
- Duplicated block (10 lines × 2) (scio-smb/src/main/scala/com/spotify/scio/smb/syntax/SortMergeBucketScioContextSyntax.scala)
- Duplicated block (10 lines × 2) (scio-snowflake/src/main/scala/com/spotify/scio/snowflake/SnowflakeIO.scala)
- Duplicated block (10 lines × 26) (scio-smb/src/main/scala/com/spotify/scio/smb/util/SMBMultiJoin.scala)
- Duplicated block (10 lines × 4) (scio-core/src/main/scala/com/spotify/scio/util/MultiJoin.scala)
- Duplicated block (10 lines × 6) (scio-smb/src/main/scala/com/spotify/scio/smb/util/SMBMultiJoin.scala)
- Duplicated block (10–11 lines × 19) (scio-smb/src/main/scala/com/spotify/scio/smb/util/SMBMultiJoin.scala)
- Duplicated block (11 lines × 2) (scalafix/rules/src/main/scala/fix/v0_14_0/FixAvroCoder.scala)
- Duplicated block (11 lines × 2) (scio-examples/src/main/scala/com/spotify/scio/examples/complete/TrafficMaxLaneFlow.scala)
- …and 325 more
Changes since last survey
- 29 commits — 24 feature/other, 5 fixes
By area
- (root) — 18 commits
- .github/workflows — 3 commits
- project/plugins.sbt — 3 commits
- integration/src — 1 commit
- project/build.properties — 1 commit
- scio-extra/src — 1 commit
- scio-managed/src — 1 commit
- scio-parquet/src — 1 commit
Notable commits
- fix: Fix IcebergIOIT implicit ambiguity; compile integration tests on PR builds (#6001)
- fix: Fix ambiguous implicit for RowField[Instant] in IcebergIOIT
- fix: Fix concurrent Sparkey closure cleaning (#6004)
- fix: Fix fat jar merge failure from Beam 2.76 envoy proto collision (#6002)
- fix: Fix/publish gh site workflow (#6005)
- change: Bump actions/checkout from 6 to 7 (#5947)
- change: Bump actions/setup-java from 5 to 6 (#5995)
- change: Expand IcebergIO write API to cover all write options (#5986)
- change: Generate scaladoc for scio-managed (#6006)
- change: Update caffeine to 3.2.4 (#5961)
- change: Update circe-core, circe-generic, ... to 0.14.16 (#5967)
- change: Update cloud-sql-connector-jdbc-sqlserver, ... to 1.28.6 (#5963)
- change: Update elasticsearch-java to 8.19.18 (#5960)
- change: Update jedis to 7.5.3 (#5980)
- change: Update jna to 5.19.1 (#5969)
- change: Update magnolify to 0.9.8; adopt portable Timestamp encoding for IcebergIO (#6000)
- change: Update metrics-core to 4.2.39 (#5968)
- change: Update munit to 1.3.3 (#5975)
- change: Update mysql-connector-j to 9.7.0 (#5964)
- change: Update neo4j-java-driver to 4.4.26 (#5972)
- …and 9 more
Architecture
- Containers 0 added · 0 removed · contexts 2 added · 0 removed · edges 2 added · 0 removed
Added bounded contexts (2)
- repository
- scalafix
Added dependency edges (2)
- repository → scalafix
- scalafix → repository (coupling)
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
spotify/scio was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 27 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 4f6c6d5b3e1ff896a6d0f151354a5ff59e41c1b7 — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-d00c643c3f66.