sryza/spark-timeseries
57.0
Adequate · 27 September 2026
7.9k
lines of production code
Scala
with Python
4
measurements over time
What this system is
Features
Add Python time-series models to SparkTS
Added Python implementations for several time-series models, including ARIMA, AR-GARCH, Autoregression, AutoregressionX, EWMA, GARCH, and RegressionARIMA. Each model accepts Numpy arrays as input and provides methods for fitting, forecasting, and sampling. A base PyModel class is introduced to handle time-dependent effects, allowing reuse across model types.
python/sparkts/models · high confidence
Added Python API documentation for sparkts
The python/source directory now contains the Sphinx configuration and reStructuredText files required to generate API documentation for the sparkts package. This includes the main index, a modules overview, and specific documentation entries for the datetimeindex, timeseriesrdd, and utils modules, enabling users to browse the Python API reference.
python/source · high confidence
Added Yahoo financial data parser
Users can now parse Yahoo Finance CSV files into TimeSeries objects. The new YahooParser object provides methods to convert raw CSV text or files from a directory into structured time series data, utilizing Java 8's java.time API for date and zone handling.
src/main/scala/com/cloudera/sparkts/parsers · high confidence
Added core time series data structures and utilities
Introduced foundational classes for time series analysis, including DateTimeIndex and its implementations (UniformDateTimeIndex, IrregularDateTimeIndex) for managing temporal data. Added Frequency classes (e.g., BusinessDayFrequency, DayFrequency) to define regular and irregular time intervals. Implemented TimeSeries and TimeSeriesRDD to represent and process collections of time series data. Added utility objects for matrix conversions (MatrixUtil), lag operations (Lag), and resampling (Resample). Also included plotting utilities (EasyPlot) for autocorrelation analysis and a Kryo registrator for efficient serialization.
src/main/scala/com/cloudera/sparkts · high confidence
Initial Python API for Spark Time Series
The \python/sparkts\ module is introduced, providing a Python interface for Spark time series operations. This includes \DateTimeIndex\ for managing temporal indices with support for uniform and irregular time series, and \TimeSeriesRDD\ for distributed time series data. The API enables converting between Pandas DataFrames/Series and Spark RDDs, performing operations like differencing, filling missing values, and computing return rates. Frequency classes (e.g., \HourFrequency\, \BusinessDayFrequency\) are also added to support time series indexing.
python/sparkts · high confidence
Initial Python package and build infrastructure
Added the initial Python package structure, including setup.py for packaging and dependency management (pandas, numpy, nose, unittest2), a MANIFEST.in to include the required JAR file, and a Makefile to build Sphinx-based Python documentation.
python · high confidence
New Java API for time series and date-time indices
Added new Java API classes—DateTimeIndexFactory, JavaTimeSeriesFactory, and JavaTimeSeriesRDDFactory—to the com.cloudera.sparkts.api.java package. These classes provide factory methods for creating UniformDateTimeIndex, IrregularDateTimeIndex, HybridDateTimeIndex, and union operations, as well as creating JavaTimeSeries and JavaTimeSeriesRDD instances from samples, vectors, DataFrames, CSV files, and Python data. This enables Java users to construct and manipulate time series data structures directly.
src/main/java · high confidence
New Java API for time series operations
Added Java wrappers for time series processing, including \JavaTimeSeries\ and \JavaTimeSeriesRDD\ classes. These provide Java-friendly access to core time series operations such as lagging, differencing, quotienting, mapping, and converting to DataFrames or RDDs, enabling Java users to perform time series analysis directly within the Spark ecosystem.
src/main/scala/com/cloudera/sparkts/api · high confidence
New time series models: ARIMA, ARIMAX, Autoregression, GARCH, EWMA, Holt-Winters, and Regression-ARIMA
The models package now includes implementations for several time series forecasting methods. ARIMA and ARIMAX models support autoregressive integrated moving average modeling with optional exogenous variables. Autoregression and AutoregressionX provide simpler autoregressive modeling with and without exogenous regressors. GARCH and ARGARCH models handle volatility modeling. EWMA implements exponential smoothing, while Holt-Winters adds seasonal components. RegressionARIMA fits linear regression with ARIMA errors using the Cochrane-Orcutt iterative method. All models implement the TimeSeriesModel interface, enabling consistent usage for fitting and forecasting.
src/main/scala/com/cloudera/sparkts/models · high confidence
Test coverage
Added Java API tests for time series and index factories; Added Python test suite for TimeSeriesRDD and DateTimeIndex; Added TimeSeriesStatisticalTests for Augmented Dickey-Fuller analysis; Added comprehensive test suites for time series components; Added test coverage for Python time-series models; Added test coverage for Yahoo time-series parser; Added test data files for time series and regression models; Added unit tests for time series models.
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Score
- CAI 50 → 57 (+6.7)
- Rubric changed (rubric-2026.08.15 → rubric-2026.09.15) — scores are not directly comparable.
Lenses
- Code Health 100 → 94 (-5.5)
- Architecture 94 → 96 (+1.6)
- Maturity 55 → 48 (-7.2)
- Readiness 25 → 46 (+20.5)
- Security 85 → 78 (-7.2)
Resolved (7)
- Coverage not measured — test suite did not build
- Dimension evaluation failed
- No exposed public API
- No tests found
- Rotate the exposed credentials — git history can't be un-committed
- Test reliability not included
- The README notes that the library is no longer under active development by the author but does not state who can contribute or how to submit patches. (README.md)
New (45)
- ARIMA.findBestARMAModel (cognitive 20) (src/main/scala/com/cloudera/sparkts/models/ARIMA.scala)
- ARIMAXModel.iterateARMAX (cognitive 20) (src/main/scala/com/cloudera/sparkts/models/ARIMAX.scala)
- Dependency hygiene PARTLY measured — Maven/Gradle declarations read, no dependency graph resolved
- Dormant codebase
- Duplicated block (10 lines × 2) (src/main/scala/com/cloudera/sparkts/models/ARIMA.scala)
- Duplicated block (15 lines × 2) (src/main/scala/com/cloudera/sparkts/models/ARIMA.scala)
- Duplicated block (32 lines × 2) (src/main/scala/com/cloudera/sparkts/models/ARIMA.scala)
- Duplicated block (5 lines × 2) (src/main/scala/com/cloudera/sparkts/models/GARCH.scala)
- Duplicated block (6 lines × 2) (src/main/scala/com/cloudera/sparkts/TimeSeriesRDD.scala)
- Duplicated block (6 lines × 2) (src/main/scala/com/cloudera/sparkts/models/ARIMA.scala)
- Duplicated block (62 lines × 2) (src/main/scala/com/cloudera/sparkts/models/ARIMA.scala)
- Duplicated block (7 lines × 2) (src/main/scala/com/cloudera/sparkts/models/GARCH.scala)
- Hotspot: src/main/scala/com/cloudera/sparkts/TimeSeriesRDD.scala (src/main/scala/com/cloudera/sparkts/TimeSeriesRDD.scala)
- Most significant orphaned file (src/main/scala/com/cloudera/sparkts/DateTimeIndex.scala)
- Most significant orphaned file (src/main/scala/com/cloudera/sparkts/TimeSeriesRDD.scala)
- Most significant orphaned file (src/main/scala/com/cloudera/sparkts/models/ARIMA.scala)
- No SBOM
- No build provenance
- No dependency advisory monitoring
- Resample.resample (cognitive 22) (src/main/scala/com/cloudera/sparkts/Resample.scala)
- …and 25 more
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
sryza/spark-timeseries was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 27 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 280aa887dc08ab114411245268f230fdabb76eec — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-d00c643c3f66.