twitter/scalding
64.0
Adequate · 27 September 2026
47.3k
lines of production code
Scala
with Java
4
measurements over time
What this system is
This system is a Scala-based distributed data processing library that provides a type-safe API for defining and executing data pipelines. It supports multiple execution backends, including Cascading, Apache Beam, and Spark, allowing jobs to run on various cluster managers. The library facilitates reading from and writing to diverse data sources such as HBase, Parquet, Avro, and relational databases, while offering utilities for serialization, mathematical operations, and interactive development.
How it got here
2012–2013 — Scalding architecture modernization
25 changes.
This period focused on a comprehensive architectural overhaul of the Scalding codebase, migrating the build system to modern sbt and replacing legacy core components with a new modular structure. Key developments included the introduction of the Execution API for composable job planning, the implementation of a type-safe Typed API, and the addition of a new interactive REPL. The work also involved significant improvements to data handling through typed sources, matrix mathematics utilities, and extensive updates to test coverage and CI infrastructure.
2014–2015 — macro-based serialization and Parquet integration
33 changes.
This period focused on enhancing Scalding's type safety and performance through extensive use of Scala macros, particularly for serialization, case class mapping, and Parquet schema generation. Significant features included typed Parquet sources and sinks with automatic schema handling, JDBC support via scalding-db, and improved reducer estimation strategies leveraging historical data. The work also expanded testing infrastructure and added support for various data sources like HBase and Thrift structs.
2016–2022 — multi-backend expansion and core refactoring
15 changes.
This period focused on expanding Scalding's execution capabilities by introducing new backends for Spark, Apache Beam, and Cats, alongside a foundational graph library called Dagon. The core codebase underwent significant refactoring to improve join logic, input estimation, and Parquet I/O stability, while a new macro-based quotation module was added for compile-time code analysis.
Features
Add Avro schema type mappings and write helpers
The scalding-avro module now includes \AvroSchemaType\, providing implicit schema definitions for primitive types (Boolean, Int, Long, Float, Double, String, ByteBuffer), collections (Array, Iterable, Map), and Avro SpecificRecords. Additionally, the \package.scala\ exposes \writePackedAvro\ and \writeUnpackedAvro\ methods, allowing users to easily write TypedPipes to Avro files using either packed or unpacked formats with explicit schema control.
scalding-avro/src · high confidence
Add FlowListenerPromise to bridge Cascading flows with Scala Futures
A new FlowListenerPromise class has been added to the cascading\_interop package to facilitate running Scalding flows without a Job. This component wraps a Cascading Flow with a Scala Promise, allowing callers to start the flow and receive a Scala Future that completes with the result of a mapping function upon successful completion, or fails with specific exceptions (such as FlowStopException) if the flow is stopped, steps fail, or runtime exceptions occur.
_scalding-core/src/main/java/com/twitter/scalding/cascading\interop · high confidence
Add Parquet-Scrooge module for reading and writing Scrooge Thrift structs
This change introduces the \scalding-parquet-scrooge\ module, providing sources and sinks to read and write Parquet files containing Scrooge-generated Thrift structs. It includes Java support classes (\ParquetScroogeInputFormat\, \ParquetScroogeOutputFormat\, \ScroogeReadSupport\, \ScroogeWriteSupport\) and Scala sources (\ParquetScrooge\, \DailySuffixParquetScrooge\, etc.) that integrate with Scalding. The implementation handles schema conversion via \ScroogeStructConverter\ and includes a workaround for PARQUET-346 to ensure compatibility with older Parquet files missing specific metadata.
scalding-parquet-scrooge · high confidence
Add Scrooge OrderedSerialization support for Thrift types
This change introduces a new \scalding-thrift-macros\ module that provides compile-time macro generation of \OrderedSerialization\ instances for Scrooge-generated Thrift types. It enables Scrooge structs, unions, and enums to be used as keys in Scalding jobs by implementing the necessary serialization, deserialization, hashing, and comparison logic via generated code.
scalding-thrift-macros · high confidence
Add filter predicate support to Parquet sources
Parquet sources now support column projection and row-level filtering via the new HasFilterPredicate trait, allowing users to specify filter predicates that are pushed down to the Parquet reader for improved performance.
scalding-parquet/src/main/scala/com/twitter/scalding/parquet · high confidence
Add macro-based Parquet tuple read/write support
A new macro implementation in the scalding-parquet module now automatically generates Parquet schema definitions and read/write support for Scala case classes. This allows users to serialize and deserialize case classes containing primitive fields, nested case classes, and collection types (List, Set, Map) to Parquet files without manually defining schema mappings or conversion logic.
scalding-parquet/src/main/scala/com/twitter/scalding/parquet/tuple/macros · high confidence
Add tutorial example for ExecutionApp
A new Scala tutorial file (ExecutionTutorial.scala) has been added to the execution-tutorial module, demonstrating how to use ExecutionApp to perform a MapReduce word count. The example shows how to read input, process data using TypedPipe, and write results to a local file using toIterableExecution, providing a concrete reference for users implementing standalone execution jobs.
tutorial/execution-tutorial · high confidence
Added Trollop command-line argument parsing library
The \scripts/lib\ directory now includes the Trollop library (version 1.16.2), a Ruby gem for parsing command-line arguments. This addition provides the necessary infrastructure for scripts to define, parse, and handle command-line options, flags, and arguments consistently.
scripts/lib · high confidence
Adds typed API aliases and RichPathFilter convenience methods
The scalding-core package now exposes aliases for the Typed API (such as TDsl, TypedPipe, TypedSink, and Grouped) to simplify imports, and introduces a RichPathFilter implicit conversion that allows users to combine Hadoop PathFilters using and, or, and not methods.
scalding-core/src/main/scala/com/twitter · high confidence
Autogenerated Scala code for tuple operations and joins
Added Ruby code generators in scalding-core/codegen that produce Scala source files for tuple converters, setters, adders, mappable traits, typed sources/sinks, function implicits, and multi-join/flatten operations. These generators create typeclass-style implicits and helper methods for tuples up to arity 22, enabling cleaner N-way join syntax via MultiJoin and automatic flattening of nested value tuples, while limiting implicit enrichment to arity 6 to control compile times.
scalding-core/codegen · high confidence
Initial Beam backend planner for Scalding
A new \BeamBackend\ module has been introduced to provide the foundational planning logic for executing Scalding jobs on Apache Beam. This change adds the \BeamPlanner\ object, which translates Scalding's \TypedPipe\ operations—such as filtering, mapping, reducing, and co-grouping—into a sequence of \BeamOp\ instructions. This serves as the core translation layer that enables Scalding programs to run on Beam-based runners like Google Cloud Dataflow.
scalding-beam/src/main · high confidence
Initial import of the Dagon graph library
The scalding-dagon module now includes the core Dagon library, providing a directed acyclic graph (DAG) implementation with support for node rewriting via rules, memoized evaluation, and heterogeneous maps. This import also introduces Scala version compatibility shims to support both Scala 2.12 and 2.13+ environments, along with the corresponding test suites for the new graph and caching components.
scalding-dagon · high confidence
Introduce Scalding REPL with interactive pipe execution and Hadoop shell access
Adds a new Scalding REPL that allows users to interactively define, run, and inspect TypedPipes directly in the console. The REPL provides implicit conversions to execute pipes locally via \save\, \snapshot\, \toIterator\, and \dump\, and exposes Hadoop's \FsShell\ for file system operations. It supports switching between Local and HDFS modes, automatically manages temporary JARs for compiled REPL code, and sources custom \.scalding\_repl\ configuration files from the current directory hierarchy.
scalding-repl/src/main · high confidence
Introduce scalding-base as a minimal dependency module
A new \scalding-base\ module has been added to provide a minimal dependency base for the Scalding library. This module introduces core foundational components, including a \Config\ wrapper for job settings, a \CancellationHandler\ for managing asynchronous stop operations, and a \FutureCache\ for caching futures in execution contexts. It also establishes the \KeyedList\ and \KeyedPipe\ abstractions for typed data operations, implements \MultiJoin\ and \LookupJoin\ for complex join semantics, and provides utility classes such as \CumulativeSum\, \HashEqualsArrayWrapper\, and \OptimizationPhases\ to support typed pipe planning and data manipulation.
scalding-base · high confidence
Introduce scalding-db JDBC macros for Scala case class to SQL mapping
Adds a new \scalding-db\ module that provides Scala macros to automatically generate SQL column definitions, Cascading fields, and JDBC read/write converters from Scala case classes. Users can now define a case class (e.g., \ExampleDBRecord\) and use \DBTypeDescriptor\ or \ColumnDefinitionProvider\ to get type-safe mappings for reading from and writing to relational databases, including support for annotations like \@size\, \@varchar\, \@text\, and \@date\ to control SQL types. The module also includes specific extensions for Vertica database compatibility and handles nullable columns and nested case classes.
scalding-db · high confidence
Introduce scalding-quotation sub-project for compile-time code quotation
This change introduces the new \scalding-quotation\ sub-project, which provides Scala macro-based capabilities to capture and inspect source code structure at compile time. The implementation includes \Liftables\ for lifting values into quasiquote trees, \ProjectionMacro\ for extracting projection information from function bodies, \QuotedMacro\ for wrapping method calls in a \Quoted\ representation, and \TextMacro\ for parsing source text to handle parameter extraction. This enables users to perform static analysis and transformation of Scalding jobs by accessing the underlying AST and source text during compilation.
scalding-quotation/src/main · high confidence
Introduce typed Parquet tuple scheme with macro-generated converters
This change adds a new \TypedParquetTupleScheme\ and supporting converter classes (e.g., \ParquetTupleConverter\, \ParquetReadSupport\, \ParquetWriteSupport\) in the \scalding-parquet\ module, enabling users to read and write Parquet files using strongly-typed Scala tuples and case classes. The implementation includes primitive and collection field converters, an \Option\ converter, and integration with Parquet's \ReadSupport\/\WriteSupport\ APIs, allowing automatic schema generation and type-safe data mapping via macros for case classes.
scalding-parquet/src/main/scala/com/twitter/scalding/parquet/tuple/scheme · high confidence
Introduce typed and untyped Parquet tuple sources and sinks
Added new \ParquetTupleSource\ and \TypedParquet\ APIs to the \scalding-parquet\ module, enabling users to read and write Parquet files using both explicit \Fields\ definitions and Scala type-based schemas. The untyped \ParquetTupleSource\ (including \DailySuffixParquetTuple\, \HourlySuffixParquetTuple\, and \FixedPathParquetTuple\) allows specifying fields and supports filter predicate pushdown, while the typed \TypedParquet\ and \TypedParquetSink\ objects provide a macro-friendly interface for creating sources and sinks from Scala case classes or tuples, automatically handling read/write support and schema mapping.
scalding-parquet/src/main/scala/com/twitter/scalding/parquet/tuple · high confidence
Introduction of customizable DateParser and CalendarOps
The scalding-date module now provides a new \DateParser\ trait and companion object, allowing users to define custom date parsing logic and chain multiple parsers together, while the previous Natty dependency has been removed. Additionally, a new \CalendarOps\ object is introduced to provide utility methods for truncating \Calendar\ and \Date\ objects to specific fields.
scalding-date/src/main · high confidence
Introduction of the Execution API and modularized core components
This change introduces the Execution API, a new monadic abstraction for composing multi-step computations, branching, and intermediate service calls, replacing the previous Job-based flow model. It includes the core Execution, CFuture, and CPromise classes for handling cancellable asynchronous operations, along with ExecutionOptimizationRules to optimize the resulting DAG. The update also modularizes the codebase by extracting Args and RangedArgs into a dedicated scalding-args module, moving Mode and JobStats into scalding-base, and adding a new maple MemorySourceTap for in-memory data handling.
repository · high confidence
Introduction of typed source and sink abstractions with transformation support
This change introduces the core \TypedSource\ and \TypedSink\ traits in the \com.twitter.scalding.typed\ package, providing a type-safe API for reading and writing data in Scalding jobs. \TypedSource\ adds an \andThen\ method to chain transformations after reading, while \TypedSink\ provides a \contraMap\ method to transform data before writing. The update also includes \BijectedSourceSink\ for handling type bijections, \PartitionSchemed\ and \PartitionUtil\ to support partitioned data sources, \MemorySink\ for in-memory testing, and \TDsl\ to expose implicit conversions that bridge standard Cascading Pipes with the new typed API.
scalding-core/src/main/scala/com/twitter/scalding/typed · high confidence
Macro support for case class tuple conversion and type descriptors
The macros module now includes macro-based implementations for converting case classes to and from Cascading tuples, as well as generating TypeDescriptors. This adds implicit and direct methods for TupleSetter, TupleConverter, and TypeDescriptor that automatically flatten case class fields (including nested case classes and Options) into tuple columns, enabling seamless serialization and deserialization of case class data in Scalding pipelines.
scalding-core/src/main/scala/com/twitter/scalding/macros · high confidence
Macro-based Parquet schema generation and writing for Scala case classes
The Scalding Parquet module now uses compile-time macros to automatically generate Parquet schemas and write logic for Scala case classes. This change introduces \ParquetSchemaProvider\ and \WriteSupportProvider\ to handle type mapping for primitives, nested case classes, and collection types (List, Set, Map), replacing previous manual or less flexible schema definition approaches.
scalding-parquet/src/main/scala/com/twitter/scalding/parquet/tuple/macros/impl · high confidence
New BDD DSL for testing Scalding pipe transformations
Added a new Behavior-Driven Development (BDD) DSL in the \scalding-core/src/main/scala/com/twitter/scalding/bdd\ package that allows users to test individual pipe transformations in isolation. This feature introduces \PipeOperationsConversions\ for the Fields API and \TypedPipeOperationsConversions\ for the Typed API, enabling developers to write modular tests using Given/When/Then syntax. This complements the existing \JobTest\ class by focusing on sub-step decomposition rather than end-to-end job testing.
scalding-core/src/main/scala/com/twitter/scalding/bdd · high confidence
New Hadoop test platform with local cluster management
Added a new test infrastructure in the \com.twitter.scalding.platform\ package that provides a \LocalCluster\ for running Scalding jobs against a local MiniDFSCluster and MiniMRCluster. This includes lifecycle management via \HadoopPlatformTest\ and \HadoopSharedPlatformTest\ traits, a \HadoopPlatform\ trait for defining test sources and sinks, and utilities like \MakeJar\ to handle classpath dependencies. The \LocalCluster\ implementation uses a file-based mutex to prevent race conditions when multiple test processes run concurrently, and configures Hadoop settings such as speculative execution and IPC ping intervals to ensure stable test execution.
scalding-hadoop-test/src/main · high confidence
New Maple Taps and Schemes for HBase, Local, and Memory Data Sources
This change introduces a suite of new data source and sink components for the Maple library, enabling users to read from and write to HBase clusters, local file systems in Cascading's local mode, and in-memory tuple collections. The update adds HBaseTap and HBaseScheme for HBase integration, LocalTap for local execution support, and MemorySinkTap, StdoutTap, and TupleMemoryInputFormat for in-memory and debugging workflows.
maple · high confidence
New Spark backend implementation with counter support and iterator utilities
The scalding-spark module introduces a new Spark backend implementation. This includes a new \Iterators\ utility for partitioning iterators into sequential key groups, a \SparkCounters\ class for tracking execution statistics via Spark accumulators, and a comprehensive test suite (\SparkBackendTests\, \SparkCountersTests\) validating basic operations, joins, writes, and counter behavior.
scalding-spark · high confidence
New configurable reducer estimation strategies
Scalding now includes a new reducer estimation module that provides multiple strategies for determining the number of reducers. The InputSizeReducerEstimator calculates reducers based on total input size and a configurable bytes-per-reducer target (defaulting to 4 GB, with support for human-readable config values like '128m'). The RatioBasedEstimator refines this estimate by applying a historical ratio of map output bytes to input bytes, filtered by an input size threshold to ensure relevance. Additionally, the RuntimeReducerEstimator uses historical job execution times to estimate reducers, supporting both input-scaled and non-input-scaled modes, and allows users to choose between mean or median aggregation schemes for task and job times. A step strategy orchestrates these estimators, applies a configurable maximum reducer cap (default 5000), and respects explicit reducer settings unless overridden.
_scalding-core/src/main/scala/com/twitter/scalding/reducer\estimation · high confidence
New hRaven-based memory and reducer estimators
The new scalding-hraven module adds estimators that query hRaven job history to improve resource planning. HRavenMemoryHistoryService and HRavenSmoothedMemoryEstimator use historical counters (committed heap, physical memory, GC time, CPU) to estimate memory needs, while HRavenReducerHistoryService, HRavenRatioBasedEstimator, and HRavenRuntimeBasedEstimator use historical job metadata (task type, status, timing) to better estimate the number of reducers based on mapper-reducer input data ratios.
scalding-hraven · high confidence
New macro-based OrderedSerialization provider for binary ordering
The serialization module now includes a new macro implementation (OrderedSerializationProviderImpl) and a BinaryOrdering trait that automatically generate OrderedSerialization instances for various types. This enables more efficient binary serialization and ordering for case classes, sealed traits, options, and other common Scala types without requiring manual implementation of serialization logic.
scalding-serialization/src/main/scala/com/twitter/scalding/serialization/macros/impl · high confidence
New mathematics library with matrix, combinatorics, and histogram utilities
The \scalding-core\ module now includes a new \mathematics\ package providing distributed mathematical utilities. This adds a \Matrix\ class with support for block matrices, row/column vectors, and algebraic operations via Algebird, along with extension methods to convert Scalding pipes into these structures. It also introduces \Combinatorics\ for generating combinations, permutations, and weighted sums, a \Histogram\ class for statistical analysis (mean, standard deviation, percentiles), and a \Poisson\ random number generator.
scalding-core/src/main/scala/com/twitter/scalding/mathematics · high confidence
New ordered serialization providers for Scala types
The ordered serialization macro system now includes dedicated providers for a wider range of Scala types, enabling more efficient and stable serialization. New providers handle primitives (Boolean, Byte, Short, Char, Int, Long, Float, Double), String, ByteBuffer, Unit, Option, Either, case classes, case objects, and sealed traits. Additionally, a StableKnownDirectSubclasses utility ensures consistent ordering for sealed trait serialization across compilations, and an ImplicitOrderedBuf fallback allows user-defined implicit OrderedSerialization instances to be used for opaque types.
_scalding-serialization/src/main/scala/com/twitter/scalding/serialization/macros/impl/ordered\serialization/providers · high confidence
New scalding-cats module provides Cats typeclass instances for Scalding types
A new \scalding-cats\ module has been added, introducing the \HellCats\ object which provides implicit Cats typeclass instances (such as \Functor\, \MonoidK\, \FunctorFilter\, and \Semigroupal\) for Scalding's \TypedPipe\ and \CoGroupable\ types. It also implements \Async\ and \Effect\ typeclasses for Scalding's \Execution\ type, enabling users to leverage Cats Effect abstractions for asynchronous execution and error handling within Scalding pipelines. The module includes comprehensive property-based tests to verify the correctness of these typeclass laws.
scalding-cats · high confidence
New specialized serialization typeclasses and utilities
The scalding-serialization module introduces a new set of specialized typeclasses (Hasher, Reader, Writer) and supporting utilities (JavaStreamEnrichments, MurmurHashUtils, PositionInputStream, UnsignedComparisons) to facilitate efficient, low-boxing serialization of primitive types, collections, and tuples. This includes a base Serialization trait with built-in equivalence and hashing laws, as well as OrderedSerialization implementations for strings and tuples, enabling more performant and type-safe data serialization within Scalding jobs.
scalding-serialization/src/main/scala/com/twitter/scalding/serialization · high confidence
New time-pathed, typed, and codec-based source classes
This change introduces several new source classes in the \com.twitter.scalding.source\ package to improve type safety and data handling. It adds \DailySuffixTsv\, \DailySuffixTypedTsv\, \DailySuffixCsv\, \DailySuffixMostRecentCsv\, \HourlySuffixTsv\, \HourlySuffixTypedTsv\, and \HourlySuffixCsv\ to support reading delimited data from time-partitioned directories (daily and hourly). It also introduces \CodecSource\ for writing types to Hadoop \WritableSequenceFile\ using a codec, \TypedSequenceFile\ for typed sequence file I/O, \NullSink\ for discarding output, and \CheckedInversion\/\MaxFailuresCheck\ to handle injection inversion errors with configurable failure limits.
scalding-core/src/main/scala/com/twitter/scalding/source · high confidence
New tutorials for Execution, Avro, JSON, Matrix operations, and REPL usage
The tutorial directory now includes several new examples and documentation files. A new Execution tutorial (Execution.md) explains the composable Execution API for planning and running Scalding jobs. New source files demonstrate specific data formats and operations: AvroTutorial0.scala and JsonTutorial0.scala show reading/writing Avro and JSON data, while MatrixTutorial0-6.scala cover matrix operations like outdegree, cofollows, filtering, intersection, cosine similarity, Jaccard similarity, and TF-IDF. A new ReplTutorial1.scala and WONDERLAND.md provide a guide for using the Scalding REPL. Additionally, Tutorial5.scala has been updated to use a configurable words input file instead of a hardcoded system path, and a .scalding\_repl file was added to verify REPL initialization during testing.
tutorial · high confidence
New utility class for safe ASCII byte extraction
Added a new \Undeprecated\ class in the serialization package that provides a public static method \getAsciiBytes\ for extracting bytes from a string. This utility is designed for ASCII data and is intended to be used by internal macros after verifying ASCII compliance, leveraging a pattern from Kryo to improve performance while suppressing deprecation warnings that Scala cannot handle directly.
scalding-serialization/src/main/java · high confidence
New versioned data store and typed LZO sources in scalding-commons
This release introduces a new versioned data store infrastructure in scalding-commons, including the Java classes VersionedStore and VersionedTap which manage data versions, cleanup, and tap creation, alongside the Scala VersionedKeyValSource for writing key-value pairs into these versioned stores. It also adds a suite of typed LZO-compressed sources (LzoTypedText, LzoGenericSource, LzoTraits) and schemes (LzoGenericScheme, CombinedSequenceFileScheme) that enable reading and writing LZO-compressed data with type safety, including support for Protobuf, Thrift, and generic binary converters.
scalding-commons · high confidence
Removals
Removal of example jobs from the examples package
The \MergeTest\, \PageRank\, and \WordCountJob\ example files have been removed from the \src/main/scala/com/twitter/scalding/examples\ directory. These files are no longer included in the library distribution.
src/main/scala/com/twitter/scalding/examples · high confidence
Removal of legacy core Scalding source files
The \src/main/scala/com/twitter/scalding\ directory has been cleared of its legacy implementation files, including \Args.scala\, \DateRange.scala\, \FieldConversions.scala\, \GeneratedConversions.scala\, \GroupBuilder.scala\, \Job.scala\, \KryoHadoopSerialization.scala\, \MemoryTap.scala\, \Mode.scala\, \Operations.scala\, and \RichPipe.scala\. This change removes the previous command-line argument parsing, date handling, field conversion implicits, group-by aggregation logic, job execution, Kryo serialization, in-memory testing taps, execution modes, and core pipe operations, indicating a major architectural shift or migration to a new codebase structure.
src/main/scala/com/twitter/scalding · high confidence
Architecture
Scalding core refactored into new modular source files
The scalding-core library has been reorganized into a new set of source files, introducing dedicated modules for core execution logic (ExecutionContext, ExecutionUtil), DSL enhancements (Dsl, FunctionImplicits), serialization support (BijectedOrderedSerialization, CascadingTokenUpdater), and various pipe operations (FoldOperations, JoinAlgorithms, Operations). This change also adds new abstractions for job execution (CascadeJob), configuration handling (HfsConfPropertySetter), and internal utilities (GeneratedMappable, GeneratedTupleAdders, IntegralComparator, LibJarsExpansion, MemoryTap), effectively restructuring the codebase to improve modularity and separation of concerns without altering the external API surface.
scalding-core/src/main/scala/com/twitter/scalding · high confidence
Behavioural changes
Fix input size estimation for glob patterns and improve record reading stability
The \GlobHfs\ tap now correctly calculates the total input size for paths containing glob patterns by iterating through matched files, preventing \IOException\ errors during size estimation. Additionally, the \ScaldingHfs\ tap uses a custom \HadoopTupleEntrySchemeIterator\ that wraps record readers with \MeasuredRecordReader\ for better tracking, addressing issues with \TupleEntrySchemeIterator\ and problematic record readers.
scalding-core/src/main/java/com/twitter/scalding/tap · high confidence
Introduce required ordered serialization for binary comparators
Scalding now supports a new mode that enforces the use of binary comparators for group-by and co-group operations, which allows keys to be compared in their serialized form to reduce serialization/deserialization overhead. This change adds the \CascadingBinaryComparator\ wrapper, \RequiredBinaryComparators\ trait with macro-based implicit \OrderedSerialization\ generation, and configuration modes (\Fail\ or \Log\) to validate that all sorting selectors in the flow use these binary comparators. It also includes specific Kryo serializers for Scalding types like \RichDate\, \DateRange\, and \Args\, and updates \KryoHadoop\ to register these and handle boxed classes appropriately.
scalding-core/src/main/scala/com/twitter/scalding/serialization · high confidence
Major overhaul of scald.rb script and addition of CI test harness
The scald.rb entry-point script has been significantly refactored to support modern Scala versions (2.10, 2.11, 2.12) and dynamic dependency resolution via SBT, replacing hardcoded paths and static JAR references with a configurable system that reads from build.sbt and allows module selection (e.g., --avro, --json). Concurrently, a suite of new shell scripts has been added to the CI pipeline: run\_test.sh handles test compilation, execution, and Codecov coverage uploads with retry logic; testValidator.sh ensures all build targets are covered in the Travis configuration; and dedicated scripts (test\_tutorials.sh, test\_matrix\_tutorials.sh, test\_repl\_tutorial.sh, etc.) verify that Scalding tutorials and REPL functionality work correctly across different modes. The legacy scalding\_gen.rb script has been removed.
scripts · high confidence
Migrate code formatting from Scalariform to Scalafmt
The project has switched its code formatting tool from Scalariform to Scalafmt (version 3.5.1). This change introduces a new \.scalafmt.conf\ configuration file that defines formatting rules, such as a 110-character column limit and specific rewrite rules (e.g., \AvoidInfix\, \SortImports\). It also configures the Scala dialect target, using \scala211\ for general files and \scala212\ for \build.sbt\ to ensure compatibility. This affects how source code is automatically formatted and linted in the repository.
(repo-wide) · high confidence
Optimized comparison for ordered serialization collections
The serialization macros now use a partial quicksort approach to compare traversable collections, avoiding the overhead of a full sort. This change improves performance for ordered serialization by reducing comparison complexity from O(N log N) to O(N + M) for unsorted inputs, while maintaining correct ordering semantics for users relying on these serialization paths.
_scalding-serialization/src/main/scala/com/twitter/scalding/serialization/macros/impl/ordered\_serialization/runtime\helpers · high confidence
Parquet I/O now uses a patched input format to fix split handling
The Scalding Parquet module now uses a custom \ScaldingDeprecatedParquetInputFormat\ instead of the upstream Apache Parquet class. This change copies the implementation from Parquet 1.12.0-RC1 to include a specific fix for task-side metadata and split handling (addressing Apache Parquet issue \#844) while waiting for the official library update. This patched input format is now the standard for reading Parquet files in both Thrift (\ParquetTBaseScheme\) and Tuple (\ParquetTupleScheme\) workflows, ensuring correct data ingestion and split calculation.
scalding-parquet/src/main/java · high confidence
Refactor CoGroup and HashJoin logic into dedicated Cascading backend joiners
The Cascading backend implementation for typed CoGroup and HashJoin operations has been restructured to use dedicated joiner classes (CoGroupedJoiner, DistinctCoGroupJoiner, and HashJoiner). This change moves the join logic out of anonymous functions and into reusable components that leverage MultiJoinFunction and Externalizer for better serialization and performance, specifically ensuring that left-side iterators are handled correctly for Hadoop compatibility.
_scalding-core/src/main/scala/com/twitter/scalding/typed/cascading\backend · high confidence
Refactored ordered serialization macros for improved length calculation and trait handling
The ordered serialization macro implementation has been restructured to introduce a new \CompileTimeLengthTypes\ hierarchy (including \ConstantLengthCalculation\, \FastLengthCalculation\, \MaybeLengthCalculation\, and \NoLengthCalculationAvailable\) that allows for more granular and efficient computation of serialized payload sizes. This change is accompanied by the introduction of \ProductLike\ and \SealedTraitLike\ helper objects, which now handle the generation of comparison, hashing, and serialization logic for case classes and sealed traits respectively. These updates enable the serialization framework to better optimize binary comparisons by leveraging early-exit byte-level checks and to correctly manage the serialization of sealed trait hierarchies with type indices.
_scalding-serialization/src/main/scala/com/twitter/scalding/serialization/macros/impl/ordered\serialization · high confidence
Fixes
Workaround for Parquet Thrift metadata bug (PARQUET-346)
This change introduces a workaround for Apache Parquet issue PARQUET-346, which causes Thrift record reading to fail when the Parquet file metadata is missing specific struct/union type information. The new \Parquet346TBaseScheme\ and \Parquet346StructTypeRepairer\ classes automatically repair this missing metadata by defaulting missing types to UNION before passing them to the Thrift converter, allowing Scalding jobs to successfully read older or malformed Parquet-Thrift files that would previously throw decoding exceptions.
scalding-parquet/src/main/scala/com/twitter/scalding/parquet/thrift · high confidence
Test coverage
Added ReplTest for Scalding REPL functionality; Added Thrift fixtures for macro testing; Added Thrift schema fixtures for Scalding Parquet Scrooge tests; Added Thrift schema for test fixtures; Added comprehensive test coverage for scalding-date components; Added comprehensive test suite for scalding-core; Added property-based tests for serialization components; Added property-based tests for serialization macros; Added serialization and comparison benchmarks; Added test coverage for scalding-args argument parsing and range validation; Added tests for Beam backend operations; Added tests for JsonLine JSON parsing and serialization; Added tests for Parquet input/output and scheme integration; Added tests for Parquet source filter and column projection behavior; Added tests for TypedParquetTuple read/write and filter pushdown; Added tests for memory and reducer estimators; Added tests for partitioned Parquet Thrift source writing; Added tests for the Execution API and distributed cache support; Added unit tests for Parquet tuple macro schema generation; Initial test suite for scalding-quotation macro; Removal of legacy test suite.
Dependencies
Major dependency and build system overhaul
The build configuration has been significantly updated to support Scala 2.11 and 2.12, upgrading core libraries such as Algebird to 0.13.4, Chill to 0.8.4, Cascading to 2.1.2, and Avro to 1.8.2. The project also integrates Beam 2.29.0 for Dataflow support, adds a new scalding-cats module for Cats typeclasses, and migrates the build tooling to sbt 1.5.6 with updated code coverage and assembly merge strategies.
(dependencies) · high confidence
Migrate build system to sbt 1.5.4 with modern plugin configuration
The project build has been upgraded from the legacy sbt 0.7.4 to sbt 1.5.4. This migration replaces the old Scala-based plugin definitions (Plugins.scala) and project structure (Project.scala) with a modern plugins.sbt file, introducing updated versions for key tools such as sbt-assembly (0.14.6), sbt-microsites (1.3.4), and sbt-scalafmt (2.4.6). Additionally, a new build script (scalding-dagon.scala) enables Scala version-specific source folder resolution for 2.12 and 2.13+, and a dedicated Travis CI log4j configuration (travis-log4j.properties) is added to control build output verbosity.
project · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
This is the PUBLIC form of this artifact. Findings are listed in full, but the details of SECURITY findings — which rule fired, in which file, on which line, and how to fix it — are deliberately withheld, and any secret-scanner results are excluded entirely. Where detail is absent here it was REMOVED FOR PUBLICATION; it is not missing from the analysis. The complete artifact is available from the repository owner.
Score
- CAI 42 → 64 (+22.2)
- Rubric changed (rubric-2026.08.15 → rubric-2026.09.15) — scores are not directly comparable.
Lenses
- Code Health 100 → 87 (-12.8)
- Architecture 90 → 99 (+9.2)
- Maturity 51 → 49 (-2.6)
- Readiness 25 → 77 (+52.3)
- Security 43 → 72 (+28.5)
Resolved (22)
- Context/problem and consequences are both absent; the body is a single decision statement with no framing (scalding-core/src/test/resources/com/twitter/scalding/test_filesystem/test_data/2013/03/2013-03.txt)
- Context/problem and consequences are both absent; the body is a single decision statement with no rationale (scalding-core/src/test/resources/com/twitter/scalding/test_filesystem/test_data/2013/07/2013-07.txt)
- Context/problem and consequences are both absent; the body is a single decision statement with no rationale (scalding-core/src/test/resources/com/twitter/scalding/test_filesystem/test_data/2013/08/2013-08.txt)
- Coverage not measured — test suite did not build
- Dimension evaluation failed
- Duplicated block (5 lines × 2) (scalding-parquet-scrooge/src/main/java/com/twitter/scalding/parquet/scrooge/ScroogeStructConverter.java)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- LLM evaluation failed
- No artifact signing
- No context/problem and no consequences; the only visible element is a title '2013-04' with no decision (scalding-core/src/test/resources/com/twitter/scalding/test_filesystem/test_data/2013/04/2013-04.txt)
- No exposed public API
- No tests found
- Test reliability not included
- early-stage repository — too little history to judge knowledge freshness
- …and 2 more
New (208)
- AbsoluteDuration.fromMillisecs (cognitive 16) (scalding-date/src/main/scala/com/twitter/scalding/AbsoluteDuration.scala)
- BeamPlanner.plan (cyclomatic 26) (scalding-beam/src/main/scala/com/twitter/scalding/beam_backend/BeamBackend.scala)
- CascadingBackend.compile (cognitive 24) (scalding-core/src/main/scala/com/twitter/scalding/typed/cascading_backend/CascadingBackend.scala)
- CascadingBackend.compile (cyclomatic 36) (scalding-core/src/main/scala/com/twitter/scalding/typed/cascading_backend/CascadingBackend.scala)
- ColumnDefinitionProviderImpl.getColumnFormats (cyclomatic 27) (scalding-db/src/main/scala/com/twitter/scalding/db/macros/impl/ColumnDefinitionProviderImpl.scala)
- Dag.ensureRec (cognitive 22) (scalding-dagon/src/main/scala/com/twitter/scalding/dagon/Dag.scala)
- Documentation: no installation or build instructions (docs/src/main/tut/resources_for_learners.md)
- Documentation: no project overview (docs/src/main/tut/resources_for_learners.md)
- Documentation: no usage examples (docs/src/main/tut/resources_for_learners.md)
- Dormant codebase
- Duplicated block (10–16 lines × 2) (scalding-serialization/src/main/scala/com/twitter/scalding/serialization/macros/impl/ordered_serialization/providers/CaseObjectOrderedBuf.scala)
- Duplicated block (11 lines × 2) (scalding-core/src/main/scala/com/twitter/scalding/RichPipe.scala)
- Duplicated block (11 lines × 2) (scalding-core/src/main/scala/com/twitter/scalding/reducer_estimation/RuntimeReducerEstimator.scala)
- Duplicated block (11–12 lines × 2) (scalding-base/src/main/scala/com/twitter/scalding/typed/OptimizationRules.scala)
- Duplicated block (16–18 lines × 3) (scalding-serialization/src/main/scala/com/twitter/scalding/serialization/macros/impl/ordered_serialization/providers/CaseClassOrderedBuf.scala)
- Duplicated block (18 lines × 2) (scalding-beam/src/main/scala/com/twitter/scalding/beam_backend/BeamBackend.scala)
- Duplicated block (18 lines × 2) (scalding-core/src/main/scala/com/twitter/scalding/typed/PartitionSchemed.scala)
- Duplicated block (22–23 lines × 3) (scalding-serialization/src/main/scala/com/twitter/scalding/serialization/macros/impl/ordered_serialization/providers/CaseClassOrderedBuf.scala)
- Duplicated block (27–31 lines × 2) (scalding-serialization/src/main/scala/com/twitter/scalding/serialization/macros/impl/ordered_serialization/providers/EitherOrderedBuf.scala)
- Duplicated block (28 lines × 2) (scalding-serialization/src/main/scala/com/twitter/scalding/serialization/macros/impl/ordered_serialization/SealedTraitLike.scala)
- …and 188 more
Architecture
- Containers 0 added · 0 removed · contexts 1 added · 0 removed · edges 0 added · 0 removed
Added bounded contexts (1)
- repository
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
twitter/scalding was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 27 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 6d4cfd0e2753d4f8a88c4671d7a1a0796a24a91e — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-d00c643c3f66.