Skip to content
CAI
Software that uses CAICheck a score

twosigma/flint

63.0

Adequate · 27 September 2026

20.9k

lines of production code

Scala

with Python

4

measurements over time

CAI band scale
CAI trend line
CAI lens gauges

What this system is

This release delivers a comprehensive Python API for time-series analysis, featuring a new \FlintContext\ and \TimeSeriesDataFrame\ that enable seamless integration with PySpark. The core Scala engine has been significantly refactored to support high-performance, zero-copy data exchange via Apache Arrow, alongside a robust set of new statistical summarizers including OLS regression, exponential smoothing, and weighted correlations. Additionally, the release upgrades the project to Spark 2.4.3 and Scala 2.12, while introducing configurable window operations and improved partition preservation checks.

Features

Add AIC and BIC to linear regression models

Linear regression models now expose methods to estimate the Akaike Information Criterion (AIC) and Bayesian Information Criterion (BIC), alongside the log-likelihood of the model parameters. These statistical measures allow users to compare the relative quality of different linear regression models fitted to the same data. The implementation also updates copyright headers and removes the deprecated Kahan summation utility.

src/main/scala/com/twosigma/flint/math · high confidence

Add Kahan summation implementation

A new Kahan summation class has been added to the math package, implementing the Neumaier modification of the Kahan summation algorithm to minimize floating-point errors when adding a large amount of doubles. The class provides methods for adding, subtracting, and retrieving the accumulated sum.

src/main/java · high confidence

Add shift functionality to time windows

Introduced the \ShiftTimeWindow\ trait and \AbsoluteTimeWindow\ case class to support time shifting in time series windows. The implementation includes \safeMinus\ and \safePlus\ helper methods to prevent value overflow when expanding windows beyond \Long.MinValue\ and \Long.MaxValue\ boundaries, ensuring robust handling of edge cases in window calculations.

src/main/scala/com/twosigma/flint/timeseries/window · high confidence

Add weather and stock data examples for time-series analysis

Added example notebooks and CSV data files (weather.csv, spy.csv) demonstrating how to perform time-series analysis using the FlintContext. The example shows loading CSV data, parsing dates, joining datasets, and running linear regression on stock and weather data.

python/examples · high confidence

Added Travis CI integration scripts for Python tests

New shell scripts and configuration files have been added to the \python/travis\ directory to automate the setup and execution of Python tests on Travis CI. This includes \prepare\_python\_tests.sh\ to download and configure Spark 2.1.1, \run\_python\_tests.sh\ to execute the test suite, and supporting configuration files (\spark-defaults.conf\, \spark\_log4j.properties\) to define the Spark environment and logging levels.

python/travis · high confidence

Added example notebook and data

A new example notebook (Flint Example.ipynb) and accompanying data file (sp500.csv) have been added to the example directory, providing a ready-to-run demonstration of the library's functionality.

example · high confidence

Added utility classes for timing and safe finalization

New utility classes have been added to the FlInT codebase to support internal operations. A new Timer object provides a method to measure the average execution time of code blocks. A Utils object introduces a tryWithSafeFinally helper that executes a block of code followed by a finally block, ensuring that exceptions in the finally block do not suppress the original exception. Additionally, a LinkedListHolder class wraps a Java LinkedList to provide methods for dropping elements based on a predicate and performing left and right folds. These utilities support internal performance measurement and robust error handling.

src/main/scala/com/twosigma/flint/util · high confidence

Added window summarization logic and batch summarizer interface

The \SummarizeWindows\ object was added to handle windowed aggregations, introducing specialized iterators (\WindowIterator\, \WindowJoinIterator\, \WindowBatchIterator\) that support both overlapped and non-overlapped windowing strategies. Additionally, the \WindowBatchSummarizer\ trait was introduced to define the interface for batch-based summarization, allowing for efficient state management and row addition/subtraction during window processing.

src/main/scala/com/twosigma/flint/rdd/function/window · high confidence

Initial Python package structure and build configuration for ts-flint

The Python package for ts-flint has been initialized with the necessary build and configuration files, including setup.py, setup.cfg, and MANIFEST.in, enabling the package to be built and installed via standard Python tools. The package declares dependencies on pandas, pyarrow, and other libraries, and includes a versioneer script to manage versioning via git tags. Additionally, a .gitignore file is added to exclude build artifacts and test coverage files from version control.

python · high confidence

Initial release of the Python Flint bindings

The \python/ts/flint\ package is introduced, providing Python bindings for the Flint time-series library. This includes a \FlintContext\ entry point for reading time-series data, a \TimeSeriesDataFrame\ class that extends PySpark's DataFrame with time-aware operations, and a \TimeSeriesGroupedData\ wrapper for grouped data. The release also adds support for user-defined functions (UDFs) via a custom \udf\ decorator, a \SummarizerFactory\ for statistical summarization, and utilities for clock generation and data serialization.

python/ts · high confidence

Introduce Arrow-based data serialization for Spark

Added new Scala classes (ArrowConverters, ArrowReader, ArrowUtils, ArrowWriter) that enable converting Spark InternalRows to and from Apache Arrow format. This provides a high-performance, zero-copy data exchange mechanism between Spark and Arrow-compatible systems, supporting primitive types, strings, dates, and timestamps.

src/main/scala/com/twosigma/flint/arrow · high confidence

Introduce Arrow-based window batch summarizer for timeseries data

Added a new \ArrowWindowBatchSummarizer\ implementation that processes timeseries window aggregations using Apache Arrow's memory model. This new component introduces a \WindowBatchSummarizerState\ to manage left and right row batches, handling index rebasing and rendering output as internal rows. This change enables more efficient batch processing for windowed aggregations by leveraging Arrow's columnar format.

src/main/scala/com/twosigma/flint/timeseries/window/summarizer · high confidence

Introduce DataFrame conversion and order/partition-preserving checks

Added DFConverter to enable zero-copy conversion between TimeSeriesRDD and DataFrames using Catalyst InternalRow. Added OrderPreservingOperation and PartitionPreservingOperation objects to check if DataFrame operations preserve order and partitions respectively. Added TimestampCast expressions for converting between Timestamp and Long with microsecond precision. Updated CatalystTypeConverters to use InternalRow instead of GenericInternalRow.

src/main/scala/org · high confidence

Introduce PythonApi annotation for Python bindings

A new Java annotation, @PythonApi, has been added to the Flinst codebase to mark code intended for use by Python bindings. This annotation includes optional fields for a message and an 'unstable' flag, providing metadata about the stability and usage of the annotated elements.

src/main/scala/com/twosigma/flint/annotation · high confidence

Introduce new ReadBuilder API for time-series data reading

A new builder-based API for reading time-series data from Parquet files has been introduced, featuring a \ReadBuilder\ class that allows users to configure read parameters such as time ranges, column selections, and sorting options. The \Parameters\ class manages these settings, including support for expanding time ranges via an \expand\ method. This change provides a more structured and flexible way to configure data source reads, particularly benefiting Python bindings through the \@PythonApi\ annotations.

src/main/scala/com/twosigma/flint/timeseries/io · high confidence

New and updated summarizers support left-subtractable operations

This change introduces support for left-subtractable summarizers, enabling efficient removal of data points from aggregations. New summarizers for dot product, exponential weighted moving average, geometric mean, ordinary least squares (OLS) regression, product, quantiles, weighted mean, z-score, and correlation have been added to the \subtractable\ package. Existing summarizers (Correlation, NthCentralMoment, NthMoment, Rows, Sum) have been refactored to extend \LeftSubtractableSummarizer\ and implement the \subtract\ method. The \RowsSummarizer\ implementation has been optimized by switching from \Vector\ to \ArrayDeque\ for better performance. Additionally, a new \LeftSubtractableOverlappableSummarizer\ trait has been introduced to support overlapped subtraction.

src/main/scala/com/twosigma/flint/rdd/function/summarize/summarizer/subtractable · high confidence

New summarizers and composability for time-series and regression analysis

Added new summarizers for exponential smoothing, weighted covariance/correlation, extremes (min/max), and a stack/composite pattern for combining multiple summarizers. The OLS regression summarizer was removed and replaced with improved handling in the regression utilities, which now include log-likelihood, Akaike, and Bayes information criteria calculations. The base Summarizer trait gained a close() method for resource cleanup, and the FlippableSummarizer trait was introduced to support the Flipper algorithm for windowed operations.

src/main/scala/com/twosigma/flint/rdd/function/summarize/summarizer · high confidence

New time-series generation and clocking infrastructure

The library introduces a new time-series generation and clocking infrastructure to support flexible time-series data creation and manipulation. A new \Clocks\ object provides \uniform\ and \random\ methods to generate \TimeSeriesRDD\ instances with evenly or unevenly sampled time ticks, respectively. This is supported by a \Clock\ abstract class and its \UniformClock\ and \RandomClock\ implementations in the \clock\ subpackage. Additionally, a \TimeSeriesGenerator\ class allows users to create random \TimeSeriesRDD\ instances with customizable schemas, cycle distributions, and column generation functions. The \CycleColumn\ trait and its implicits enable the application of user-defined functions over time cycles, facilitating the addition of new columns based on cycle data. A \TimeSeriesStore\ object and trait provide a unified interface for converting DataFrames and OrderedRDDs into a normalized store, handling partitioning and normalization logic. These changes also include a \TimeType\ abstraction to handle time conversions between internal representations (nanoseconds) and SQL types (long/timestamp).

src/main/scala/com/twosigma/flint/timeseries · high confidence

New time-series summarizers for statistical and composite operations

The \src/main/scala/com/twosigma/flint/timeseries/summarize/summarizer\ directory was refactored to introduce a wide range of new summarizer implementations. These include statistical aggregations such as covariance, dot product, geometric mean, standard deviation, variance, skewness, and kurtosis. The update also adds support for exponential smoothing and weighted moving averages, alongside composite and stackable summarizer factories that allow combining multiple summarizers. Additionally, an \ArrowSummarizer\ was introduced to handle Arrow batch processing, and a \PredicateSummarizer\ was added to filter rows based on conditions before summarization.

src/main/scala/com/twosigma/flint/timeseries/summarize/summarizer · high confidence

Behavioural changes

Configurable time column type and batch size limits

A new configuration file, FlintConf.scala, introduces two configurable parameters for window operations: the maximum batch size for summarizing windows (default 500,000) and the time column type, which now defaults to 'timestamp' (supporting both 'long' and 'timestamp' values).

src/main/scala/com/twosigma/flint · high confidence

Improved performance and fixed bugs in RDD join operations

The join functions (LeftJoin, FutureLeftJoin, SymmetricJoin) have been refactored to improve performance, particularly for large numbers of partitions. The implementation now uses Java's \HashMap\ and \ArrayDeque\ instead of Scala's \mutable.Map\ and \mutable.Queue\ for better efficiency. A critical bug in \FutureLeftJoin\ that caused incorrect purging of old rows from the right table has been fixed. Additionally, \RangeMergeJoin\ now requires its input splits to be sorted by range, ensuring correct intersection calculations.

src/main/scala/com/twosigma/flint/rdd/function/join · medium confidence

Refactor OrderedRDD construction to use a partitioning type enum

The API for creating an OrderedRDD has been updated to accept a KeyPartitioningType enum (UnSorted, Sorted, NormalizedSorted) instead of separate boolean flags or dedicated factory methods. This change simplifies the interface and allows for more efficient handling of sorted and normalized RDDs by explicitly distinguishing between unsorted, sorted, and normalized states, which improves performance for large numbers of partitions.

src/main/scala/com/twosigma/flint/rdd · high confidence

Refactor internal row manipulation utilities and schema handling

The \InternalRowUtils\ object was introduced to centralize functions for manipulating Catalyst \InternalRow\ objects, including optimized implementations for \concat2\ and \addOrUpdate\ operations. The \Schema\ object was moved to the \row\ package and refactored to use a new \DuplicateColumnsException\ for validation, while also adding an \append\ method and updating \prependTimeAndKey\ to accept a \TimeType\ parameter.

src/main/scala/com/twosigma/flint/timeseries/row · medium confidence

Refactor timeseries summarize API with new ColumnList and Summarizer abstractions

The timeseries summarize API has been refactored to introduce new abstractions for column selection and summarizer composition. A new \ColumnList\ trait and \Summarizer\/\SummarizerFactory\ traits have been added to define input column requirements and enable composition of overlappable summarizers. The previous \Summary\ object, which exposed methods like \count\, \sum\, \mean\, and \OLSRegression\ directly, has been removed in favor of this new structure, which allows for more flexible column filtering and summarizer combination.

src/main/scala/com/twosigma/flint/timeseries/summarize · high confidence

Refactored grouping iterators and intervalization logic

The \GroupByKeyIterator\ was renamed to \SummarizeByKeyIterator\ and refactored to use a \Summarizer\ for memory-bounded summarization, allowing \TaskContext\ to be null and supporting empty \TimeSeriesRDD\. In \Intervalize\, the \intervalize\ function was updated to use a broadcast join for the clock and added \inclusion\ and \rounding\ parameters to control interval boundaries and rounding behavior.

src/main/scala/com/twosigma/flint/rdd/function/group · medium confidence

Refactored summarization to use tree-based aggregation and support overlapped summarizers

The summarization logic was refactored to use a new TreeAggregate/TreeReduce implementation for multi-level tree aggregation across partitions, improving performance for large numbers of partitions. The SummarizeWindows class was removed in favor of a new Summarize implementation that supports a configurable aggregation depth. Additionally, support for composing overlappable summarizers was added via OverlappableCompositeSummarizer, and the OverlappableSummarizer trait was updated to use addOverlapped instead of add.

src/main/scala/com/twosigma/flint/rdd/function/summarize · high confidence

Fixes

The copyright headers in the Hadoop module source files were updated to reflect the 2017 year. Additionally, the Hadoop object was refactored to use Grizzled-SLF4J for logging instead of the previous Spark Logging trait, ensuring consistent logging behavior across the module.

src/main/scala/com/twosigma/flint/hadoop · high confidence

The copyright headers in the regression summarizer files (LagWindow.scala and LagWindowQueue.scala) have been updated from 2015-2016 to 2017, and trailing whitespace has been removed.

src/main/scala/com/twosigma/flint/rdd/function/summarize/summarizer/regression · high confidence

Test coverage

Added Python unit tests for Flint; Added test automation scripts for Scala and Python; Added test data for time series summarizers and joins.

Dependencies

Upgrade to Spark 2.4.3 and Scala 2.12

The project now targets Spark 2.4.3 (previously 1.6.1/2.0.2/2.1.0/2.2/2.3) and Scala 2.12 (previously 2.11.8). This update includes a range of dependency changes: adding pyarrow, grizzled-slf4j, scalacheck, and jackson modules; removing spark-csv, commons-csv, avro, and scala-logging; and updating play-json, scalatest, and httpclient versions. The build configuration also removes automatic header insertion and adds an assemblyNoTest command alias.

(dependencies) · medium confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

Score

  • CAI 52 → 63 (+11.3)
  • Rubric changed (rubric-2026.08.15 → rubric-2026.09.15) — scores are not directly comparable.

Lenses

  • Code Health 100 → 92 (-7.7)
  • Architecture 94 → 99 (+4.6)
  • Maturity 54 → 49 (-5.2)
  • Readiness 25 → 61 (+35.9)
  • Security 100 → 86 (-13.8)

Resolved (6)

  • Coverage not measured — test suite did not build
  • Dimension evaluation failed
  • LLM evaluation failed
  • No exposed public API
  • No tests found
  • Test reliability not included

New (82)

  • Dependency hygiene PARTLY measured — Python dependencies read, no exact pin to grade for currency
  • Dormant codebase
  • Duplicated block (10–11 lines × 2) (src/main/scala/com/twosigma/flint/rdd/function/window/SummarizeWindows.scala)
  • Duplicated block (13–14 lines × 2) (src/main/scala/com/twosigma/flint/rdd/function/summarize/summarizer/subtractable/NthCentralMomentSummarizer.scala)
  • Duplicated block (15–16 lines × 2) (src/main/scala/com/twosigma/flint/timeseries/TimeSeriesRDD.scala)
  • Duplicated block (26 lines × 2) (src/main/scala/com/twosigma/flint/rdd/function/window/SummarizeWindows.scala)
  • Duplicated block (5 lines × 2) (src/main/scala/com/twosigma/flint/rdd/function/group/Intervalize.scala)
  • Duplicated block (5 lines × 2) (src/main/scala/com/twosigma/flint/rdd/function/join/FutureLeftJoin.scala)
  • Duplicated block (5 lines × 2) (src/main/scala/com/twosigma/flint/rdd/function/window/SummarizeWindows.scala)
  • Duplicated block (7 lines × 2) (src/main/scala/com/twosigma/flint/rdd/function/summarize/Summarize.scala)
  • Duplicated block (7 lines × 2) (src/main/scala/com/twosigma/flint/rdd/function/window/SummarizeWindows.scala)
  • Duplicated block (8 lines × 2) (src/main/scala/com/twosigma/flint/rdd/function/summarize/summarizer/subtractable/OLSRegressionSummarizer.scala)
  • Duplicated block (8 lines × 2) (src/main/scala/com/twosigma/flint/rdd/function/window/SummarizeWindows.scala)
  • Duplicated block (8 lines × 2) (src/main/scala/com/twosigma/flint/rdd/function/window/SummarizeWindows.scala)
  • Duplicated block (9 lines × 2) (src/main/scala/com/twosigma/flint/rdd/function/window/SummarizeWindows.scala)
  • FileTooLong: flint/dataframe.py (python/ts/flint/dataframe.py)
  • FileTooLong: python/versioneer.py (python/versioneer.py)
  • FileTooLong: timeseries/TimeSeriesRDD.scala (src/main/scala/com/twosigma/flint/timeseries/TimeSeriesRDD.scala)
  • FileTooLong: window/SummarizeWindows.scala (src/main/scala/com/twosigma/flint/rdd/function/window/SummarizeWindows.scala)
  • FixmeComment (src/main/scala/com/twosigma/flint/rdd/PartitionsIterator.scala)
  • …and 62 more

Architecture

  • Containers 0 added · 0 removed · contexts 1 added · 0 removed · edges 0 added · 0 removed

Added bounded contexts (1)

  • repository

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

twosigma/flint was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 27 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit e4d0abae03e619cc6f52fb212d33977bf5aceb52 — the exact code this score is about.
  • Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-d00c643c3f66.