Skip to content
CAI
Software that uses CAICheck a score

YotpoLtd/metorikku

56.5

Adequate · 28 September 2026

3.9k

lines of production code

Scala

primary language

2

measurements over time

CAI band scale
CAI trend line
CAI lens gauges

What this system is

Metorikku is a configuration-driven data processing engine built on Apache Spark that executes metric pipelines defined via YAML or JSON. It ingests data from diverse sources including Kafka, JDBC, HDFS, and NoSQL databases, then applies SQL transformations, custom Scala code, and data quality validations using Amazon Deequ. The system writes results to various destinations such as Hudi, Elasticsearch, and InfluxDB, supporting both batch and structured streaming modes with built-in observability and lineage tracking.

How it got here

2017 — Architecture refactoring and streaming support

17 changes.

This period focused on a comprehensive refactoring of Metorikku's core architecture, replacing legacy session and configuration models with a new Job abstraction and unified writer interface. The changes introduced native support for structured streaming, data quality validation via Deequ, and remote configuration loading, while simultaneously upgrading dependencies to Spark 3.2 and addressing security vulnerabilities in Log4j.

2018–2019 — Streaming support and integration expansion

30 changes.

This period focused on extending Metorikku's capabilities to support structured streaming and adding native readers and writers for a wide range of data sources, including Kafka, JDBC, Hudi, Elasticsearch, MongoDB, and Cassandra. The work involved refactoring the configuration model to handle streaming and periodic jobs, implementing new input/output infrastructure, and establishing comprehensive end-to-end test environments for these integrations.

2020–2026 — CI automation and data quality integration

16 changes.

This period focused on establishing robust CI/CD pipelines through new automation scripts and expanding the project's data processing capabilities with Amazon Deequ-based quality validation. Significant efforts were also directed toward modernizing the Docker infrastructure to support newer Spark versions, integrating metadata lineage with Apache Atlas, and enhancing observability via Prometheus metrics.

Features

Add Cassandra input reader and InfluxDB instrumentation configuration

Users can now read data from Cassandra tables using the new CassandraInput reader, which supports host, keyspace, table, and optional authentication credentials. Additionally, an InfluxDBConfig case class has been introduced to support instrumentation configuration for InfluxDB.

src/main/scala/com/yotpo/metorikku/configuration/job/instrumentation, src/main/scala/com/yotpo/metorikku/input/readers/cassandra · high confidence

Add JDBC output writers for table and query-based data export

Users can now write Spark DataFrames to JDBC databases using two new output writers. The JDBCOutputWriter supports standard table writes with configurable save modes, connection properties (user, password, driver), and optional JDBC-specific options like truncate, cascadeTruncate, createTableColumnTypes, createTableOptions, and sessionInitStatement, including an optional post-query execution. The JDBCQueryWriter allows executing custom SQL queries against the database, supporting batched inserts with configurable maxBatchSize, minPartitions, and maxPartitions, and handles various data types including primitives, dates, timestamps, binary data, and complex types (arrays, maps, structs) via JSON serialization.

src/main/scala/com/yotpo/metorikku/output/writers/jdbc · high confidence

Add Kafka end-to-end test environment

Users can now run an end-to-end test suite for Kafka integration. This change introduces a Docker Compose configuration that spins up a Spark cluster, Zookeeper, and Kafka services, along with helper scripts to seed data and verify consumption, enabling automated validation of Kafka-based data pipelines.

e2e/hive, e2e/kafka · high confidence

Add MongoDB input reader

Introduces a new MongoDB input reader that allows users to load data from MongoDB collections into Spark DataFrames. The reader supports configuration via URI, database, collection, and partitioning options, and includes logic to sanitize BSON data types (such as regexes and nested structures) into standard Spark-compatible types like strings and arrays.

src/main/scala/com/yotpo/metorikku/input/readers/elasticsearch, src/main/scala/com/yotpo/metorikku/input/readers/mongodb · high confidence

Add UDF example demonstrating custom Scala code integration

The examples/udf directory now includes a complete example for using User Defined Functions (UDFs) with custom Scala code. This addition provides a Scala object (TestUDF) that registers a UDF, a corresponding YAML metric configuration (udf\_metric.yaml) that invokes this UDF within a SQL step, sample data (employees.jsonl), and a test definition (udf\_test.yaml) to verify the output. This allows users to see how to integrate custom logic into their metric pipelines.

examples/udf · high confidence

Add end-to-end test environment for Apache Hudi integration

Added a Docker Compose configuration and test script to run end-to-end tests for the Apache Hudi writer. The environment spins up Spark, Hive (with NOSASL authentication), and a MySQL metastore to validate Hudi operations, including manual Hive sync scenarios.

e2e/hudi · high confidence

Add file stream input example with JSONL streaming and Hudi output

A new example demonstrating file stream input has been added, featuring a job configuration that reads JSONL data as a stream and writes results to a Hudi table. The example includes sample input data, a job definition specifying streaming triggers and checkpointing, and a metric configuration for processing the data.

_examples/file\_input\stream · high confidence

Added Spark configuration files for Atlas integration and JSON logging

The Spark configuration directory now includes \atlas-application.properties\ to define Kafka connectivity settings for Atlas (using environment variables for Zookeeper and bootstrap servers) and \log4j.json.properties\ to configure JSON-formatted console logging with specific log levels for Spark, Jetty, Parquet, and Hive components.

docker/spark/k8s/conf · high confidence

Added Spark metrics configuration and Prometheus scraping rules

This change introduces dedicated configuration files for Spark 2 and Spark 3 metrics collection. For Spark 2, a new \metrics\_spark2.properties\ file configures JMX sinks for driver and executor JVM metrics. For Spark 3, \metrics\_spark3.properties\ switches to PrometheusServlet sinks, exposing metrics at specific paths for the driver, executor, master, and applications. Additionally, a new \prometheus.yaml\ file defines the scraping rules and metric name mappings for master, worker, driver, and executor metrics, including support for DAGScheduler, CodeGenerator, LiveListenerBus, and streaming metrics.

docker/spark/k8s/metrics · high confidence

Added sample data for JDBC file date-range inputs

The examples directory now includes sample CSV files (movies.csv) organized by date (2017/09/01, 02, 03) under examples/file\_date\_range\_inputs to support JDBC file input features, alongside a reorganization of existing example files into the examples/file\_inputs directory.

_examples/file\_date\_range\_inputs, examples/file\inputs · high confidence

Initial release of Apache Hive Docker container with Atlas, Hudi, and Prometheus support

This change introduces the complete Docker infrastructure for Apache Hive (version 2.3.7), including the Dockerfile, initialization scripts, and service start-up logic. The container is built on OpenJDK 8 and bundles Apache Atlas (version 2.0.0) for metadata management, Apache Hudi (version 0.10.0) for lakehouse capabilities, and the Apiary Kafka Metastore Listener for event streaming. It also integrates Prometheus JMX monitoring for observability and supports JSON logging. The setup includes a health check mechanism, AWS S3 access configuration, and a dedicated script to import Hive metadata into Atlas.

docker/hive · high confidence

Introduce Amazon Deequ-based data quality validation with failed data dumping

Users can now define and run data quality checks on Spark DataFrames using the Amazon Deequ library. This change adds new operators for verifying completeness, uniqueness, size, and value containment. When validation fails, the system logs the specific constraint violations and, if configured, automatically dumps the failed DataFrame to a specified S3 path in Parquet format for further inspection.

src/main/scala/com/yotpo/metorikku/metric/stepActions/dataQuality · high confidence

Introduce Apache Hudi support and refactor file output writers

Metorikku now supports writing to Apache Hudi tables via the new HudiOutputWriter, enabling features such as schema alignment, nullable field support, and manual Hive synchronization. The file output infrastructure has been refactored to use a shared FileOutputWriter base class, with dedicated writers for CSV, JSON, and Parquet formats. Additionally, a new CatalogWriter allows updating Hive catalog metadata from single-row DataFrames, and the system now supports protecting external tables from being updated with empty output folders.

src/main/scala/com/yotpo/metorikku/output/writers/file · high confidence

Introduce Kafka output writer with streaming support and extra options

Added a new KafkaOutputWriter that enables writing Spark DataFrames to Kafka topics. The writer supports both batch and structured streaming modes, allowing users to specify the topic, value column, optional key column, output mode, and trigger settings. It also introduces support for extra Kafka options via the \extraOptions\ configuration parameter and handles compression types for streaming outputs.

src/main/scala/com/yotpo/metorikku/output/writers/kafka · high confidence

Introduce file-based input readers with streaming support

This change adds new file input readers (FileInput, FilesInput, FileStreamInput) to the Metorikku engine, enabling users to read data from local or remote file paths in both batch and streaming modes. The implementation includes automatic format detection for CSV, JSON, JSONL, and Parquet files, along with support for custom schemas and read options. A key behavioral addition is the FileStreamInput class, which leverages Spark Structured Streaming to process files as they arrive, allowing for real-time data ingestion pipelines.

src/main/scala/com/yotpo/metorikku/input/readers/file · high confidence

New CI/CD automation scripts for building, testing, and publishing Docker images

This change introduces a suite of new shell scripts in the \scripts/\ directory to automate the build, test, and deployment pipeline. The \build.sh\ script handles SBT compilation and assembly, while \test.sh\ executes unit tests and validates Metorikku functionality against specific Scala and Spark versions. The \docker.sh\ script orchestrates the creation of base Spark images and final Metorikku Docker images for both K8s and standalone modes, supporting Spark 2 and Spark 3. Deployment is handled by \docker\_publish.sh\ and \docker\_publish\_dev.sh\ for Docker Hub, and \ecr\_publish.sh\ for AWS ECR, ensuring versioned tags are pushed correctly. Supporting scripts manage CI caching (\before\_cache.sh\, \load\_from\_cache.sh\, \save\_docker\_to\_cache.sh\) and Travis CI integration (\travis\_build.sh\).

scripts · high confidence

New Elasticsearch output writer with authentication support

A new Elasticsearch output writer has been added, allowing users to configure write operations such as save mode, index resource, mapping ID, and write operation type. The writer now supports HTTP basic authentication by accepting user and password configurations, which are passed to the underlying Elasticsearch Spark connector.

src/main/scala/com/yotpo/metorikku/output/writers/elasticsearch · high confidence

New Kafka streaming examples and CDC integration

Added a set of YAML-based example configurations in the \examples/kafka\ directory to demonstrate various streaming patterns. These include basic Kafka-to-Kafka pipelines with append and complete output modes, Kafka-to-Parquet writing, and a specific Change Data Capture (CDC) example that ingests from Kafka, processes events using SQL transformations, and writes to a Hudi table. The examples also include a test configuration with mock data to validate the aggregation logic.

examples/kafka · high confidence

New data quality operators for size, uniqueness, and containment checks

Users can now apply additional data quality checks via new operator classes: HasSize validates DataFrame row counts against a threshold using configurable comparison operators (==, !=, \>=, \>, \<=, \<); HasUniqueness verifies uniqueness across single or combined key columns with a configurable fraction and default equality operator; IsComplete ensures no nulls in specified columns; IsContainedIn validates that column values fall within a defined set of allowed strings; and IsUnique checks for uniqueness in a single column. These operators leverage the Deequ library and are integrated into the existing data quality step-action framework.

src/main/scala/com/yotpo/metorikku/metric/stepActions/dataQuality/operators · high confidence

New data transformation steps and UDFs for schema alignment, merging, and security

This update introduces several new processing steps and user-defined functions to the codebase. Users can now align table schemas using the new AlignTables step, convert column names to camelCase, and drop specific columns. Data integrity is enhanced with a RemoveDuplicates step and a SelectiveMerge step that merges two dataframes while prioritizing values from the second source. For data security, a new ObfuscateColumns step allows masking sensitive data using MD5, SHA-256, or custom values. Additionally, the system now supports loading tables only if they exist (LoadIfExists), setting watermarks for streaming, and converting data to Avro format with Schema Registry integration. A new UDF is also available to convert epoch milliseconds to timestamps.

src/main/scala/com/yotpo/metorikku/code · high confidence

New end-to-end CDC test environment with Debezium, Kafka, and Hudi

Added a new end-to-end test suite in the e2e/cdc directory that validates the full pipeline of Change Data Capture (CDC) from MySQL through Kafka and Schema Registry to Hudi and Hive. The environment is defined by a docker-compose.yml file orchestrating Debezium MySQL connector, Kafka, Schema Registry, and Spark/Hudi services. Supporting scripts handle service readiness checks, connector registration, and database seeding, while test metrics verify that CDC events (including updates and deletes) are correctly processed and persisted in the Hive metastore.

e2e/cdc · high confidence

New input sources and configuration refactoring

This change introduces configuration classes for new input sources: Cassandra, Elasticsearch, MongoDB, and JDBC, allowing users to read data from these systems. It also refactors the existing File input to support streaming via a new \isStream\ option and moves the \FileDateRange\ configuration into the input package structure.

src/main/scala/com/yotpo/metorikku/configuration/job/input · high confidence

New metric configuration model and parser

This change introduces the core data structures and parsing logic for metric configurations in Metorikku. It adds a \Configuration\ case class to define the top-level structure containing a list of \Step\ objects and optional \Output\ definitions. The \Step\ class now supports specific properties such as \ignoreOnFailures\ (allowing steps to continue despite errors), \checkpoint\ toggles, and embedded \DataQualityCheckList\ definitions. The \Output\ class defines supported output types (including Parquet, JDBC, Elasticsearch, Hudi, etc.) and options like repartitioning and empty-output protection. A new \ConfigurationParser\ handles reading these configurations from JSON or YAML files, validating extensions, and resolving file paths for both local and Hadoop-compatible storage.

src/main/scala/com/yotpo/metorikku/configuration/metric · high confidence

New self-contained Metorikku image for Spark 3.5.8 and Delta 3.1.0

A new Dockerfile introduces a self-contained Metorikku runtime image based on Apache Spark 3.5.8 and Delta Lake 3.1.0. This image builds the Metorikku application jar internally and bundles necessary dependencies for S3A access, JSON structured logging, and Prometheus JMX metrics. It also modifies the entrypoint to run as root, resolving access issues with hostPath scratch volumes.

docker/metorikku-spark35 · high confidence

New standalone Spark container with Atlas and lineage integration

This change introduces a new standalone Spark Docker image (docker/spark/standalone) that includes configuration and scripts for integrating with Apache Atlas for metadata management and a lineage reporting system. The container supports configurable logging (JSON or standard), S3 storage settings, and optional Hive metastore connections, allowing users to deploy a Spark cluster with built-in data governance and observability capabilities.

docker/spark/standalone · high confidence

Support for custom code steps and SQL query checkpointing

Users can now execute custom Scala code via the new Code step action, allowing for flexible, programmatic data transformations outside of standard SQL. Additionally, the SQL step action now supports optional DataFrame checkpointing, which helps manage memory usage and improve performance for large datasets by persisting intermediate results to storage.

src/main/scala/com/yotpo/metorikku/metric/stepActions · high confidence

Removals

Removal of legacy Parquet output writer implementation

The legacy Parquet output writer class has been removed from the codebase. This change eliminates the previous implementation that handled writing DataFrames to Parquet files using a specific configuration structure for save modes and output paths, indicating a shift in how Parquet outputs are managed within the application.

src/main/scala/com/yotpo/metorikku/output/writers/csv, src/main/scala/com/yotpo/metorikku/output/writers/parquet · high confidence

Removal of legacy YAML configuration system

The legacy YAML-based configuration infrastructure has been removed from the application. This change deletes the core configuration traits and classes (\Configuration\, \DefaultConfiguration\, \YAMLConfiguration\), the input/output data models (\Input\, \Output\), and the YAML-specific parsing logic (\YAMLConfigurationParser\, \ConfigurationParser\). Users relying on the previous YAML configuration format will no longer be able to load settings through this specific code path.

src/main/scala/com/yotpo/metorikku/configuration · high confidence

Removal of legacy output configuration classes

The configuration classes for Cassandra, File, Redis, and Segment outputs have been removed from the codebase. This change eliminates the legacy configuration structures that previously defined connection parameters (such as hosts, ports, and API keys) for these specific output destinations, indicating a shift away from these direct configuration models in favor of a new approach.

src/main/scala/com/yotpo/metorikku/configuration/outputs · high confidence

Behavioural changes

Add lineage reporting configuration and update logging defaults

Users can now configure lineage reporting via the new lineage.properties file, which defaults to a Kafka reporter using environment variables for bootstrap servers and topic names. Additionally, the logging configuration has been updated to set the default log level for the Spark REPL (org.apache.spark.repl.Main) to WARN, replacing the previous SparkContext setting, and includes INFO-level logging for Apache Hudi.

src/main/resources · high confidence

Instrumentation writer refactored to support tags and time columns

The instrumentation output writer has been refactored to replace the previous key-column-based logic with a more flexible tagging system. Users can now explicitly define a 'valueColumn' and a 'timeColumn' in the configuration; if the value column is omitted, the last column is used by default. All other columns are automatically converted into tags for the metric, and the writer now supports emitting timestamps for each data point, improving the granularity and context of the recorded metrics.

src/main/scala/com/yotpo/metorikku/output/writers/instrumentation · high confidence

Introduces structured job configuration model with streaming and periodic support

The job configuration system has been refactored to use a new set of case classes (Configuration, Input, Output, Streaming, Periodic, etc.) that define the schema for job definitions. This change adds native support for configuring streaming jobs (with trigger modes, checkpoint locations, and output modes) and periodic jobs (with trigger durations) directly in the configuration file. It also expands the supported input sources to include Elasticsearch and MongoDB, and output destinations to include Hudi, while introducing configuration options for quoting Spark variables, ignoring Deequ validations, and specifying failed DataFrame locations.

src/main/scala/com/yotpo/metorikku/configuration/job · high confidence

JDBC input reader now supports automatic table partitioning

The JDBC input reader has been refactored to automatically partition data reads based on a partition column (defaulting to 'id'). It calculates the maximum ID in the table to determine the upper bound for partitioning, allowing users to specify the number of partitions or let the system calculate it automatically based on table size. This change improves performance for large table reads by enabling parallel processing.

src/main/scala/com/yotpo/metorikku/input/readers/jdbc · high confidence

Kafka input now supports regex-based topic subscription and Confluent Schema Registry deserialization

The Kafka input reader has been refactored to allow subscribing to multiple topics via a regex pattern using the new \topicPattern\ configuration option, in addition to the existing single-topic subscription. It also adds native support for Avro data serialized with the Confluent Schema Registry; when \schemaRegistryUrl\ is provided, the reader automatically fetches the schema and deserializes the Kafka message values into structured columns. A new \KafkaLagWriter\ component ensures that consumer group offsets are committed correctly during Spark streaming progress events.

src/main/scala/com/yotpo/metorikku/input/readers/kafka · high confidence

Metric execution and output handling refactored to support streaming, data quality, and configurable step failures

The metric processing engine has been restructured to replace the previous calculator-based execution model with a new step-action factory and integrated data quality checks. Users can now configure metrics to ignore failed steps via the \ignoreOnFailures\ flag or global \continueOnFailedStep\ setting, allowing pipelines to proceed despite individual step errors. The system introduces Deequ integration for data quality validations, with the ability to dump failing dataframes to S3 and skip validations entirely. Additionally, the metric writer now supports structured streaming outputs, including a new batch-mode streaming write capability, and improves lag time reporting by handling empty dataframes and supporting multiple time units.

src/main/scala/com/yotpo/metorikku/metric · high confidence

Redshift writer supports pre/post actions, extra options, and configurable string size

The Redshift output writer now allows users to configure pre-actions and post-actions via the \preActions\ and \postActions\ properties, enabling custom SQL execution before and after data loads. It also introduces an \extraOptions\ property to pass arbitrary key-value pairs to the underlying Redshift connector, and a \maxStringSize\ property to override the automatic detection of VARCHAR column lengths. Additionally, the writer has migrated from the Databricks Redshift connector to the \io.github.spark\_redshift\_community.spark.redshift\ library.

src/main/scala/com/yotpo/metorikku/output/writers/redshift · high confidence

Refactor configuration loading to support remote files and YAML

The utility layer for file and configuration handling has been refactored to support reading configuration files from remote storage systems (S3, HDFS) and YAML formats. FileUtils now uses Hadoop FileSystem APIs to read remote paths and Jackson YAML support for parsing, replacing the previous local-only JSON approach. Additionally, the codebase introduces HudiUtils to manage pending compactions for Hudi tables and TableUtils to parse table names, while removing the legacy TableType and TestUtils utilities.

src/main/scala/com/yotpo/metorikku/utils · high confidence

Refactored input reading and session management architecture

The input reading logic has been restructured to use a new \Reader\ trait in the \input\ package, replacing the previous \InputTableReader\ object that handled JSON, CSV, and Parquet formats directly. Concurrently, the \Session\ object has been removed, eliminating the previous singleton-based approach for managing the Spark session, configuration, and dataframe registration. This change shifts how input data is read and how the Spark context is initialized and accessed within the application.

src/main/scala/com/yotpo/metorikku/input, src/main/scala/com/yotpo/metorikku/session · high confidence

Refactored job execution model and added periodic job support

Metorikku now uses a new Job abstraction to manage Spark sessions, configuration, and instrumentation, replacing the previous Session-based approach. This change introduces support for periodic jobs, allowing users to configure tasks that run on a recurring schedule with cache clearing between executions. The execution flow has been updated to handle these periodic runs alongside standard metric execution, and output writers like Cassandra have been refactored to align with the new job context.

src/main/scala/com/yotpo/metorikku · high confidence

Refactored output writer architecture and added streaming support

The output subsystem has been restructured to support a unified writer interface and streaming capabilities. The \MetricOutputWriter\ trait has been renamed to \Writer\ and now includes a \writeStream\ method, enabling batch and streaming output modes. A new \WriterFactory\ replaces the previous \MetricOutputWriterFactory\ and \MetricOutputHandler\, centralizing the creation of output writers (including JDBC, Kafka, Hudi, and Elasticsearch) and improving configuration handling. Additionally, session registration logic has been moved into a dedicated \WriterSessionRegistration\ trait, as seen in the updated Redis writer, to better manage Spark session configurations.

src/main/scala/com/yotpo/metorikku/output · high confidence

Replace Groupon metrics with InfluxDB-based streaming instrumentation

The instrumentation subsystem has been refactored to support structured streaming metrics by introducing a new provider architecture. The previous Groupon-based metrics utility has been removed and replaced with an InfluxDB integration that writes counters and gauges to an InfluxDB instance. A new streaming query listener now automatically tracks input event counts, processed events per second, and query termination exceptions, exposing these metrics via the new InfluxDB provider.

src/main/scala/com/yotpo/metorikku/instrumentation · high confidence

Spark 3.5-compatible Redis output and stream mock components added to overlay

The Spark 3.5 Docker overlay now includes a Redis output writer that writes DataFrames to Redis using a manual JSON conversion helper, replacing the deprecated \JSONObject\ class from \scala.util.parsing\ to ensure compatibility with Scala 2.13+. Additionally, a stream mock input reader has been added that uses \ExpressionEncoder.apply\ to handle the \RowEncoder\ API changes introduced in Spark 3.5.0, allowing streaming tests to function correctly within the Spark 3.5 environment.

docker/metorikku-spark35/spark35-overlay · high confidence

Standardized output configuration models and relocated Redshift config

The output configuration models for Cassandra, Elasticsearch, File, Hudi, JDBC, Kafka, Redis, and Segment have been introduced as new case classes, establishing mandatory fields (such as host, nodes, or connection URL) and optional parameters for each destination. Additionally, the Redshift configuration class has been moved from the general outputs package to the job output package, and its JSON serialization annotations have been removed to align with the new standard configuration structure.

src/main/scala/com/yotpo/metorikku/configuration/job/output · high confidence

Support for Hive table properties and refined schema updates for partitioned tables

The catalog output now allows users to set custom table properties on Hive tables via a new \setTableMetadata\ method. Additionally, when overwriting existing external tables, the schema update logic has been refined to exclude partition-by columns from the schema alteration, ensuring that partitioned tables are handled correctly during metadata updates.

src/main/scala/com/yotpo/metorikku/output/catalog · high confidence

Updated Spark container with Hadoop 2.9.2, Hive 2.3.3, and Hudi 0.10.0

The custom Spark Docker image in docker/spark/custom-hadoop has been updated to include specific versions of key dependencies: Hadoop 2.9.2, Hive 2.3.3, and Hudi 0.10.0. The Dockerfile now explicitly removes old Hadoop 2.7 JARs, installs the new Hadoop and Hive versions, and adds Hudi Hive sync and Hadoop MR bundles to the Hive library path. Additionally, a new spark-env.sh script is introduced to set the SPARK\_DIST\_CLASSPATH using the hadoop classpath command, ensuring proper integration with the installed Hadoop distribution.

docker/spark/custom-hadoop · high confidence

Updated movie analysis example to use YAML configuration and JSONL mock data

The movie analysis example in the \examples\ directory has been updated to use the newer YAML-based metric configuration (\movies\_metric.yaml\) instead of the previous JSON format. The example now includes dedicated test fixtures (\movies\_test.yaml\) and mock data files in JSONL format (\movies.jsonl\, \ratings.jsonl\) to replace the old JSON mocks. Additionally, the configuration demonstrates new capabilities such as file date range inputs, checkpointing, and multiple output formats (Parquet, CSV, JSON, XML).

examples · high confidence

Fixes

Segment output writer refactored for batching, instrumentation, and robust user ID handling

The Segment output writer now supports configurable batch sizes and optional sleep intervals between batches, allowing users to control throughput and avoid rate limits. User IDs are handled as strings rather than being cast to integers, preventing data loss or errors for large identifiers. Additionally, the writer now uses the new instrumentation framework for metrics instead of the legacy Spark counter system, and the class signature has been updated to accept an instrumentation factory.

src/main/scala/com/yotpo/metorikku/output/writers/segment · high confidence

Test coverage

Add InfluxDB end-to-end test environment; Add Kafka end-to-end test scripts; Add test configuration model and parser; Added end-to-end tests for Elasticsearch integration with Spark 2 and Spark 3; Added mock data fixtures for test configurations; Added sample movie data for Kafka end-to-end testing; Added test coverage for new data processing steps and UDFs; Added test tag for unsupported Spark versions; Added tests for MetricReporting time unit validation; Added tests for data quality check operators and failure handling; Added tests for the ObfuscateColumns UDF; Added tests for the new epoch-to-timestamp conversion function; Expanded test coverage for Metorikku validation and error handling; Initial implementation of the Metorikku test framework.

Dependencies

Build infrastructure and signing setup

The build environment has been updated to SBT 1.3.12 and sbt-assembly 0.14.10. Several SBT plugins were upgraded (sbt-scoverage to 1.6.0, sbt-release to 1.0.10) and new plugins were added for dependency graphing (sbt-dependency-graph), release management (sbt-release), Sonatype publishing (sbt-sonatype), and PGP signing (sbt-pgp). Additionally, PGP public and private keys for Metorikku were added to the project to support artifact signing.

project · high confidence

Major dependency upgrade and build restructuring for Spark 3.2+ and Log4j 2.17

The main build.sbt has been significantly refactored to support Spark 3.2.1 (default) and 2.4.8, upgrading core dependencies including Jackson to 2.10.0, Spark-Redshift to 4.2.0, Deequ to 2.0.1, and Hudi to 0.10.0. Crucially, Log4j has been upgraded from 2.10.0 to 2.17.1 to address security vulnerabilities, and the build now uses environment variables for version management. A new Spark 3.5 overlay build.sbt is added for Docker image construction, and assembly output names now include the Scala binary version.

(dependencies) · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

Score

  • CAI 57 → 56 (-0.4)
  • Rubric changed (rubric-2026.09.8 → rubric-2026.09.16) — scores are not directly comparable.

Lenses

  • Code Health 97 → 97 (+0.0)
  • Architecture 100 → 78 (-21.9)
  • Maturity 55 → 55 (+0.0)
  • Readiness 45 → 48 (+2.6)
  • Security 64 → 64 (+0.0)

Resolved (3)

  • Documentation: no architecture or design documentation (README.md)
  • Documentation: no installation or build instructions (README.md)
  • Documentation: no usage examples (README.md)

New (7)

  • No ADRs found
  • Outdated: org.apache.commons:commons-text
  • Outdated: org.apache.logging.log4j:log4j-api
  • Outdated: org.apache.logging.log4j:log4j-core
  • Outdated: org.apache.logging.log4j:log4j-slf4j-impl
  • Outdated: org.influxdb:influxdb-java
  • Projects may be oversized for their cohesion

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

YotpoLtd/metorikku was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 28 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit b15d77c46e760c30a4f2ac70921c7f0c539a56c3 — the exact code this score is about.
  • Scored under rubric-2026.09.16 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-d46da229e3fd.