YotpoLtd/metorikku
56.5
Adequate · 28 September 2026
3.9k
lines of production code
Scala
primary language
2
measurements over time
What this system is
Metorikku is a configuration-driven data processing engine built on Apache Spark that executes metric pipelines defined via YAML or JSON. It ingests data from diverse sources including Kafka, JDBC, HDFS, and NoSQL databases, then applies SQL transformations, custom Scala code, and data quality validations using Amazon Deequ. The system writes results to various destinations such as Hudi, Elasticsearch, and InfluxDB, supporting both batch and structured streaming modes with built-in observability and lineage tracking.
How it got here
2017 — Architecture refactoring and streaming support
17 changes.
This period focused on a comprehensive refactoring of Metorikku's core architecture, replacing legacy session and configuration models with a new Job abstraction and unified writer interface. The changes introduced native support for structured streaming, data quality validation via Deequ, and remote configuration loading, while simultaneously upgrading dependencies to Spark 3.2 and addressing security vulnerabilities in Log4j.
2018–2019 — Streaming support and integration expansion
30 changes.
This period focused on extending Metorikku's capabilities to support structured streaming and adding native readers and writers for a wide range of data sources, including Kafka, JDBC, Hudi, Elasticsearch, MongoDB, and Cassandra. The work involved refactoring the configuration model to handle streaming and periodic jobs, implementing new input/output infrastructure, and establishing comprehensive end-to-end test environments for these integrations.
2020–2026 — CI automation and data quality integration
16 changes.
This period focused on establishing robust CI/CD pipelines through new automation scripts and expanding the project's data processing capabilities with Amazon Deequ-based quality validation. Significant efforts were also directed toward modernizing the Docker infrastructure to support newer Spark versions, integrating metadata lineage with Apache Atlas, and enhancing observability via Prometheus metrics.
Features
Add Cassandra input reader and InfluxDB instrumentation configuration
Users can now read data from Cassandra tables using the new CassandraInput reader, which supports host, keyspace, table, and optional authentication credentials. Additionally, an InfluxDBConfig case class has been introduced to support instrumentation configuration for InfluxDB.
src/main/scala/com/yotpo/metorikku/configuration/job/instrumentation, src/main/scala/com/yotpo/metorikku/input/readers/cassandra · high confidence
Add JDBC output writers for table and query-based data export
Users can now write Spark DataFrames to JDBC databases using two new output writers. The JDBCOutputWriter supports standard table writes with configurable save modes, connection properties (user, password, driver), and optional JDBC-specific options like truncate, cascadeTruncate, createTableColumnTypes, createTableOptions, and sessionInitStatement, including an optional post-query execution. The JDBCQueryWriter allows executing custom SQL queries against the database, supporting batched inserts with configurable maxBatchSize, minPartitions, and maxPartitions, and handles various data types including primitives, dates, timestamps, binary data, and complex types (arrays, maps, structs) via JSON serialization.
src/main/scala/com/yotpo/metorikku/output/writers/jdbc · high confidence
Add Kafka end-to-end test environment
Users can now run an end-to-end test suite for Kafka integration. This change introduces a Docker Compose configuration that spins up a Spark cluster, Zookeeper, and Kafka services, along with helper scripts to seed data and verify consumption, enabling automated validation of Kafka-based data pipelines.
e2e/hive, e2e/kafka · high confidence
Add MongoDB input reader
Introduces a new MongoDB input reader that allows users to load data from MongoDB collections into Spark DataFrames. The reader supports configuration via URI, database, collection, and partitioning options, and includes logic to sanitize BSON data types (such as regexes and nested structures) into standard Spark-compatible types like strings and arrays.
src/main/scala/com/yotpo/metorikku/input/readers/elasticsearch, src/main/scala/com/yotpo/metorikku/input/readers/mongodb · high confidence
Add UDF example demonstrating custom Scala code integration
The examples/udf directory now includes a complete example for using User Defined Functions (UDFs) with custom Scala code. This addition provides a Scala object (TestUDF) that registers a UDF, a corresponding YAML metric configuration (udf\_metric.yaml) that invokes this UDF within a SQL step, sample data (employees.jsonl), and a test definition (udf\_test.yaml) to verify the output. This allows users to see how to integrate custom logic into their metric pipelines.
examples/udf · high confidence
Add end-to-end test environment for Apache Hudi integration
Added a Docker Compose configuration and test script to run end-to-end tests for the Apache Hudi writer. The environment spins up Spark, Hive (with NOSASL authentication), and a MySQL metastore to validate Hudi operations, including manual Hive sync scenarios.
e2e/hudi · high confidence
Add file stream input example with JSONL streaming and Hudi output
A new example demonstrating file stream input has been added, featuring a job configuration that reads JSONL data as a stream and writes results to a Hudi table. The example includes sample input data, a job definition specifying streaming triggers and checkpointing, and a metric configuration for processing the data.
_examples/file\_input\stream · high confidence
Added Spark configuration files for Atlas integration and JSON logging
The Spark configuration directory now includes \atlas-application.properties\ to define Kafka connectivity settings for Atlas (using environment variables for Zookeeper and bootstrap servers) and \log4j.json.properties\ to configure JSON-formatted console logging with specific log levels for Spark, Jetty, Parquet, and Hive components.
docker/spark/k8s/conf · high confidence
Added Spark metrics configuration and Prometheus scraping rules
This change introduces dedicated configuration files for Spark 2 and Spark 3 metrics collection. For Spark 2, a new \metrics\_spark2.properties\ file configures JMX sinks for driver and executor JVM metrics. For Spark 3, \metrics\_spark3.properties\ switches to PrometheusServlet sinks, exposing metrics at specific paths for the driver, executor, master, and applications. Additionally, a new \prometheus.yaml\ file defines the scraping rules and metric name mappings for master, worker, driver, and executor metrics, including support for DAGScheduler, CodeGenerator, LiveListenerBus, and streaming metrics.
docker/spark/k8s/metrics · high confidence
Added sample data for JDBC file date-range inputs
The examples directory now includes sample CSV files (movies.csv) organized by date (2017/09/01, 02, 03) under examples/file\_date\_range\_inputs to support JDBC file input features, alongside a reorganization of existing example files into the examples/file\_inputs directory.
_examples/file\_date\_range\_inputs, examples/file\inputs · high confidence
Initial release of Apache Hive Docker container with Atlas, Hudi, and Prometheus support
This change introduces the complete Docker infrastructure for Apache Hive (version 2.3.7), including the Dockerfile, initialization scripts, and service start-up logic. The container is built on OpenJDK 8 and bundles Apache Atlas (version 2.0.0) for metadata management, Apache Hudi (version 0.10.0) for lakehouse capabilities, and the Apiary Kafka Metastore Listener for event streaming. It also integrates Prometheus JMX monitoring for observability and supports JSON logging. The setup includes a health check mechanism, AWS S3 access configuration, and a dedicated script to import Hive metadata into Atlas.
docker/hive · high confidence
Introduce Amazon Deequ-based data quality validation with failed data dumping
Users can now define and run data quality checks on Spark DataFrames using the Amazon Deequ library. This change adds new operators for verifying completeness, uniqueness, size, and value containment. When validation fails, the system logs the specific constraint violations and, if configured, automatically dumps the failed DataFrame to a specified S3 path in Parquet format for further inspection.
src/main/scala/com/yotpo/metorikku/metric/stepActions/dataQuality · high confidence
Introduce Apache Hudi support and refactor file output writers
Metorikku now supports writing to Apache Hudi tables via the new HudiOutputWriter, enabling features such as schema alignment, nullable field support, and manual Hive synchronization. The file output infrastructure has been refactored to use a shared FileOutputWriter base class, with dedicated writers for CSV, JSON, and Parquet formats. Additionally, a new CatalogWriter allows updating Hive catalog metadata from single-row DataFrames, and the system now supports protecting external tables from being updated with empty output folders.
src/main/scala/com/yotpo/metorikku/output/writers/file · high confidence
Introduce Kafka output writer with streaming support and extra options
Added a new KafkaOutputWriter that enables writing Spark DataFrames to Kafka topics. The writer supports both batch and structured streaming modes, allowing users to specify the topic, value column, optional key column, output mode, and trigger settings. It also introduces support for extra Kafka options via the \extraOptions\ configuration parameter and handles compression types for streaming outputs.
src/main/scala/com/yotpo/metorikku/output/writers/kafka · high confidence
Introduce file-based input readers with streaming support
This change adds new file input readers (FileInput, FilesInput, FileStreamInput) to the Metorikku engine, enabling users to read data from local or remote file paths in both batch and streaming modes. The implementation includes automatic format detection for CSV, JSON, JSONL, and Parquet files, along with support for custom schemas and read options. A key behavioral addition is the FileStreamInput class, which leverages Spark Structured Streaming to process files as they arrive, allowing for real-time data ingestion pipelines.
src/main/scala/com/yotpo/metorikku/input/readers/file · high confidence
New CI/CD automation scripts for building, testing, and publishing Docker images
This change introduces a suite of new shell scripts in the \scripts/\ directory to automate the build, test, and deployment pipeline. The \build.sh\ script handles SBT compilation and assembly, while \test.sh\ executes unit tests and validates Metorikku functionality against specific Scala and Spark versions. The \docker.sh\ script orchestrates the creation of base Spark images and final Metorikku Docker images for both K8s and standalone modes, supporting Spark 2 and Spark 3. Deployment is handled by \docker\_publish.sh\ and \docker\_publish\_dev.sh\ for Docker Hub, and \ecr\_publish.sh\ for AWS ECR, ensuring versioned tags are pushed correctly. Supporting scripts manage CI caching (\before\_cache.sh\, \load\_from\_cache.sh\, \save\_docker\_to\_cache.sh\) and Travis CI integration (\travis\_build.sh\).
scripts · high confidence
New Elasticsearch output writer with authentication support
A new Elasticsearch output writer has been added, allowing users to configure write operations such as save mode, index resource, mapping ID, and write operation type. The writer now supports HTTP basic authentication by accepting user and password configurations, which are passed to the underlying Elasticsearch Spark connector.
src/main/scala/com/yotpo/metorikku/output/writers/elasticsearch · high confidence
New Kafka streaming examples and CDC integration
Added a set of YAML-based example configurations in the \examples/kafka\ directory to demonstrate various streaming patterns. These include basic Kafka-to-Kafka pipelines with append and complete output modes, Kafka-to-Parquet writing, and a specific Change Data Capture (CDC) example that ingests from Kafka, processes events using SQL transformations, and writes to a Hudi table. The examples also include a test configuration with mock data to validate the aggregation logic.
examples/kafka · high confidence
New data quality operators for size, uniqueness, and containment checks
Users can now apply additional data quality checks via new operator classes: HasSize validates DataFrame row counts against a threshold using configurable comparison operators (==, !=, \>=, \>, \<=, \<); HasUniqueness verifies uniqueness across single or combined key columns with a configurable fraction and default equality operator; IsComplete ensures no nulls in specified columns; IsContainedIn validates that column values fall within a defined set of allowed strings; and IsUnique checks for uniqueness in a single column. These operators leverage the Deequ library and are integrated into the existing data quality step-action framework.
src/main/scala/com/yotpo/metorikku/metric/stepActions/dataQuality/operators · high confidence
New data transformation steps and UDFs for schema alignment, merging, and security
This update introduces several new processing steps and user-defined functions to the codebase. Users can now align table schemas using the new AlignTables step, convert column names to camelCase, and drop specific columns. Data integrity is enhanced with a RemoveDuplicates step and a SelectiveMerge step that merges two dataframes while prioritizing values from the second source. For data security, a new ObfuscateColumns step allows masking sensitive data using MD5, SHA-256, or custom values. Additionally, the system now supports loading tables only if they exist (LoadIfExists), setting watermarks for streaming, and converting data to Avro format with Schema Registry integration. A new UDF is also available to convert epoch milliseconds to timestamps.
src/main/scala/com/yotpo/metorikku/code · high confidence
New end-to-end CDC test environment with Debezium, Kafka, and Hudi
Added a new end-to-end test suite in the e2e/cdc directory that validates the full pipeline of Change Data Capture (CDC) from MySQL through Kafka and Schema Registry to Hudi and Hive. The environment is defined by a docker-compose.yml file orchestrating Debezium MySQL connector, Kafka, Schema Registry, and Spark/Hudi services. Supporting scripts handle service readiness checks, connector registration, and database seeding, while test metrics verify that CDC events (including updates and deletes) are correctly processed and persisted in the Hive metastore.
e2e/cdc · high confidence
New input sources and configuration refactoring
This change introduces configuration classes for new input sources: Cassandra, Elasticsearch, MongoDB, and JDBC, allowing users to read data from these systems. It also refactors the existing File input to support streaming via a new \isStream\ option and moves the \FileDateRange\ configuration into the input package structure.
src/main/scala/com/yotpo/metorikku/configuration/job/input · high confidence
New metric configuration model and parser
This change introduces the core data structures and parsing logic for metric configurations in Metorikku. It adds a \Configuration\ case class to define the top-level structure containing a list of \Step\ objects and optional \Output\ definitions. The \Step\ class now supports specific properties such as \ignoreOnFailures\ (allowing steps to continue despite errors), \checkpoint\ toggles, and embedded \DataQualityCheckList\ definitions. The \Output\ class defines supported output types (including Parquet, JDBC, Elasticsearch, Hudi, etc.) and options like repartitioning and empty-output protection. A new \ConfigurationParser\ handles reading these configurations from JSON or YAML files, validating extensions, and resolving file paths for both local and Hadoop-compatible storage.
src/main/scala/com/yotpo/metorikku/configuration/metric · high confidence
New self-contained Metorikku image for Spark 3.5.8 and Delta 3.1.0
A new Dockerfile introduces a self-contained Metorikku runtime image based on Apache Spark 3.5.8 and Delta Lake 3.1.0. This image builds the Metorikku application jar internally and bundles necessary dependencies for S3A access, JSON structured logging, and Prometheus JMX metrics. It also modifies the entrypoint to run as root, resolving access issues with hostPath scratch volumes.
docker/metorikku-spark35 · high confidence
New standalone Spark container with Atlas and lineage integration
This change introduces a new standalone Spark Docker image (docker/spark/standalone) that includes configuration and scripts for integrating with Apache Atlas for metadata management and a lineage reporting system. The container supports configurable logging (JSON or standard), S3 storage settings, and optional Hive metastore connections, allowing users to deploy a Spark cluster with built-in data governance and observability capabilities.
docker/spark/standalone · high confidence
Support for custom code steps and SQL query checkpointing
Users can now execute custom Scala code via the new Code step action, allowing for flexible, programmatic data transformations outside of standard SQL. Additionally, the SQL step action now supports optional DataFrame checkpointing, which helps manage memory usage and improve performance for large datasets by persisting intermediate results to storage.
src/main/scala/com/yotpo/metorikku/metric/stepActions · high confidence
Removals
Removal of legacy Parquet output writer implementation
The legacy Parquet output writer class has been removed from the codebase. This change eliminates the previous implementation that handled writing DataFrames to Parquet files using a specific configuration structure for save modes and output paths, indicating a shift in how Parquet outputs are managed within the application.
src/main/scala/com/yotpo/metorikku/output/writers/csv, src/main/scala/com/yotpo/metorikku/output/writers/parquet · high confidence
Removal of legacy YAML configuration system
The legacy YAML-based configuration infrastructure has been removed from the application. This change deletes the core configuration traits and classes (\Configuration\, \DefaultConfiguration\, \YAMLConfiguration\), the input/output data models (\Input\, \Output\), and the YAML-specific parsing logic (\YAMLConfigurationParser\, \ConfigurationParser\). Users relying on the previous YAML configuration format will no longer be able to load settings through this specific code path.
src/main/scala/com/yotpo/metorikku/configuration · high confidence
Removal of legacy output configuration classes
The configuration classes for Cassandra, File, Redis, and Segment outputs have been removed from the codebase. This change eliminates the legacy configuration structures that previously defined connection parameters (such as hosts, ports, and API keys) for these specific output destinations, indicating a shift away from these direct configuration models in favor of a new approach.
src/main/scala/com/yotpo/metorikku/configuration/outputs · high confidence
Behavioural changes
Add lineage reporting configuration and update logging defaults
Users can now configure lineage reporting via the new lineage.properties file, which defaults to a Kafka reporter using environment variables for bootstrap servers and topic names. Additionally, the logging configuration has been updated to set the default log level for the Spark REPL (org.apache.spark.repl.Main) to WARN, replacing the previous SparkContext setting, and includes INFO-level logging for Apache Hudi.
src/main/resources · high confidence
Instrumentation writer refactored to support tags and time columns
The instrumentation output writer has been refactored to replace the previous key-column-based logic with a more flexible tagging system. Users can now explicitly define a 'valueColumn' and a 'timeColumn' in the configuration; if the value column is omitted, the last column is used by default. All other columns are automatically converted into tags for the metric, and the writer now supports emitting timestamps for each data point, improving the granularity and context of the recorded metrics.
src/main/scala/com/yotpo/metorikku/output/writers/instrumentation · high confidence
Introduces structured job configuration model with streaming and periodic support
The job configuration system has been refactored to use a new set of case classes (Configuration, Input, Output, Streaming, Periodic, etc.) that define the schema for job definitions. This change adds native support for configuring streaming jobs (with trigger modes, checkpoint locations, and output modes) and periodic jobs (with trigger durations) directly in the configuration file. It also expands the supported input sources to include Elasticsearch and MongoDB, and output destinations to include Hudi, while introducing configuration options for quoting Spark variables, ignoring Deequ validations, and specifying failed DataFrame locations.
src/main/scala/com/yotpo/metorikku/configuration/job · high confidence
JDBC input reader now supports automatic table partitioning
The JDBC input reader has been refactored to automatically partition data reads based on a partition column (defaulting to 'id'). It calculates the maximum ID in the table to determine the upper bound for partitioning, allowing users to specify the number of partitions or let the system calculate it automatically based on table size. This change improves performance for large table reads by enabling parallel processing.
src/main/scala/com/yotpo/metorikku/input/readers/jdbc · high confidence
Kafka input now supports regex-based topic subscription and Confluent Schema Registry deserialization
The Kafka input reader has been refactored to allow subscribing to multiple topics via a regex pattern using the new \topicPattern\ configuration option, in addition to the existing single-topic subscription. It also adds native support for Avro data serialized with the Confluent Schema Registry; when \schemaRegistryUrl\ is provided, the reader automatically fetches the schema and deserializes the Kafka message values into structured columns. A new \KafkaLagWriter\ component ensures that consumer group offsets are committed correctly during Spark streaming progress events.
src/main/scala/com/yotpo/metorikku/input/readers/kafka · high confidence
Metric execution and output handling refactored to support streaming, data quality, and configurable step failures
The metric processing engine has been restructured to replace the previous calculator-based execution model with a new step-action factory and integrated data quality checks. Users can now configure metrics to ignore failed steps via the \ignoreOnFailures\ flag or global \continueOnFailedStep\ setting, allowing pipelines to proceed despite individual step errors. The system introduces Deequ integration for data quality validations, with the ability to dump failing dataframes to S3 and skip validations entirely. Additionally, the metric writer now supports structured streaming outputs, including a new batch-mode streaming write capability, and improves lag time reporting by handling empty dataframes and supporting multiple time units.
src/main/scala/com/yotpo/metorikku/metric · high confidence
Redshift writer supports pre/post actions, extra options, and configurable string size
The Redshift output writer now allows users to configure pre-actions and post-actions via the \preActions\ and \postActions\ properties, enabling custom SQL execution before and after data loads. It also introduces an \extraOptions\ property to pass arbitrary key-value pairs to the underlying Redshift connector, and a \maxStringSize\ property to override the automatic detection of VARCHAR column lengths. Additionally, the writer has migrated from the Databricks Redshift connector to the \io.github.spark\_redshift\_community.spark.redshift\ library.
src/main/scala/com/yotpo/metorikku/output/writers/redshift · high confidence
Refactor configuration loading to support remote files and YAML
The utility layer for file and configuration handling has been refactored to support reading configuration files from remote storage systems (S3, HDFS) and YAML formats. FileUtils now uses Hadoop FileSystem APIs to read remote paths and Jackson YAML support for parsing, replacing the previous local-only JSON approach. Additionally, the codebase introduces HudiUtils to manage pending compactions for Hudi tables and TableUtils to parse table names, while removing the legacy TableType and TestUtils utilities.
src/main/scala/com/yotpo/metorikku/utils · high confidence
Refactored input reading and session management architecture
The input reading logic has been restructured to use a new \Reader\ trait in the \input\ package, replacing the previous \InputTableReader\ object that handled JSON, CSV, and Parquet formats directly. Concurrently, the \Session\ object has been removed, eliminating the previous singleton-based approach for managing the Spark session, configuration, and dataframe registration. This change shifts how input data is read and how the Spark context is initialized and accessed within the application.
src/main/scala/com/yotpo/metorikku/input, src/main/scala/com/yotpo/metorikku/session · high confidence
Refactored job execution model and added periodic job support
Metorikku now uses a new Job abstraction to manage Spark sessions, configuration, and instrumentation, replacing the previous Session-based approach. This change introduces support for periodic jobs, allowing users to configure tasks that run on a recurring schedule with cache clearing between executions. The execution flow has been updated to handle these periodic runs alongside standard metric execution, and output writers like Cassandra have been refactored to align with the new job context.
src/main/scala/com/yotpo/metorikku · high confidence
Refactored output writer architecture and added streaming support
The output subsystem has been restructured to support a unified writer interface and streaming capabilities. The \MetricOutputWriter\ trait has been renamed to \Writer\ and now includes a \writeStream\ method, enabling batch and streaming output modes. A new \WriterFactory\ replaces the previous \MetricOutputWriterFactory\ and \MetricOutputHandler\, centralizing the creation of output writers (including JDBC, Kafka, Hudi, and Elasticsearch) and improving configuration handling. Additionally, session registration logic has been moved into a dedicated \WriterSessionRegistration\ trait, as seen in the updated Redis writer, to better manage Spark session configurations.
src/main/scala/com/yotpo/metorikku/output · high confidence
Replace Groupon metrics with InfluxDB-based streaming instrumentation
The instrumentation subsystem has been refactored to support structured streaming metrics by introducing a new provider architecture. The previous Groupon-based metrics utility has been removed and replaced with an InfluxDB integration that writes counters and gauges to an InfluxDB instance. A new streaming query listener now automatically tracks input event counts, processed events per second, and query termination exceptions, exposing these metrics via the new InfluxDB provider.
src/main/scala/com/yotpo/metorikku/instrumentation · high confidence
Spark 3.5-compatible Redis output and stream mock components added to overlay
The Spark 3.5 Docker overlay now includes a Redis output writer that writes DataFrames to Redis using a manual JSON conversion helper, replacing the deprecated \JSONObject\ class from \scala.util.parsing\ to ensure compatibility with Scala 2.13+. Additionally, a stream mock input reader has been added that uses \ExpressionEncoder.apply\ to handle the \RowEncoder\ API changes introduced in Spark 3.5.0, allowing streaming tests to function correctly within the Spark 3.5 environment.
docker/metorikku-spark35/spark35-overlay · high confidence
Standardized output configuration models and relocated Redshift config
The output configuration models for Cassandra, Elasticsearch, File, Hudi, JDBC, Kafka, Redis, and Segment have been introduced as new case classes, establishing mandatory fields (such as host, nodes, or connection URL) and optional parameters for each destination. Additionally, the Redshift configuration class has been moved from the general outputs package to the job output package, and its JSON serialization annotations have been removed to align with the new standard configuration structure.
src/main/scala/com/yotpo/metorikku/configuration/job/output · high confidence
Support for Hive table properties and refined schema updates for partitioned tables
The catalog output now allows users to set custom table properties on Hive tables via a new \setTableMetadata\ method. Additionally, when overwriting existing external tables, the schema update logic has been refined to exclude partition-by columns from the schema alteration, ensuring that partitioned tables are handled correctly during metadata updates.
src/main/scala/com/yotpo/metorikku/output/catalog · high confidence
Updated Spark container with Hadoop 2.9.2, Hive 2.3.3, and Hudi 0.10.0
The custom Spark Docker image in docker/spark/custom-hadoop has been updated to include specific versions of key dependencies: Hadoop 2.9.2, Hive 2.3.3, and Hudi 0.10.0. The Dockerfile now explicitly removes old Hadoop 2.7 JARs, installs the new Hadoop and Hive versions, and adds Hudi Hive sync and Hadoop MR bundles to the Hive library path. Additionally, a new spark-env.sh script is introduced to set the SPARK\_DIST\_CLASSPATH using the hadoop classpath command, ensuring proper integration with the installed Hadoop distribution.
docker/spark/custom-hadoop · high confidence
Updated movie analysis example to use YAML configuration and JSONL mock data
The movie analysis example in the \examples\ directory has been updated to use the newer YAML-based metric configuration (\movies\_metric.yaml\) instead of the previous JSON format. The example now includes dedicated test fixtures (\movies\_test.yaml\) and mock data files in JSONL format (\movies.jsonl\, \ratings.jsonl\) to replace the old JSON mocks. Additionally, the configuration demonstrates new capabilities such as file date range inputs, checkpointing, and multiple output formats (Parquet, CSV, JSON, XML).
examples · high confidence
Fixes
Segment output writer refactored for batching, instrumentation, and robust user ID handling
The Segment output writer now supports configurable batch sizes and optional sleep intervals between batches, allowing users to control throughput and avoid rate limits. User IDs are handled as strings rather than being cast to integers, preventing data loss or errors for large identifiers. Additionally, the writer now uses the new instrumentation framework for metrics instead of the legacy Spark counter system, and the class signature has been updated to accept an instrumentation factory.
src/main/scala/com/yotpo/metorikku/output/writers/segment · high confidence
Test coverage
Add InfluxDB end-to-end test environment; Add Kafka end-to-end test scripts; Add test configuration model and parser; Added end-to-end tests for Elasticsearch integration with Spark 2 and Spark 3; Added mock data fixtures for test configurations; Added sample movie data for Kafka end-to-end testing; Added test coverage for new data processing steps and UDFs; Added test tag for unsupported Spark versions; Added tests for MetricReporting time unit validation; Added tests for data quality check operators and failure handling; Added tests for the ObfuscateColumns UDF; Added tests for the new epoch-to-timestamp conversion function; Expanded test coverage for Metorikku validation and error handling; Initial implementation of the Metorikku test framework.
Dependencies
Build infrastructure and signing setup
The build environment has been updated to SBT 1.3.12 and sbt-assembly 0.14.10. Several SBT plugins were upgraded (sbt-scoverage to 1.6.0, sbt-release to 1.0.10) and new plugins were added for dependency graphing (sbt-dependency-graph), release management (sbt-release), Sonatype publishing (sbt-sonatype), and PGP signing (sbt-pgp). Additionally, PGP public and private keys for Metorikku were added to the project to support artifact signing.
project · high confidence
Major dependency upgrade and build restructuring for Spark 3.2+ and Log4j 2.17
The main build.sbt has been significantly refactored to support Spark 3.2.1 (default) and 2.4.8, upgrading core dependencies including Jackson to 2.10.0, Spark-Redshift to 4.2.0, Deequ to 2.0.1, and Hudi to 0.10.0. Crucially, Log4j has been upgraded from 2.10.0 to 2.17.1 to address security vulnerabilities, and the build now uses environment variables for version management. A new Spark 3.5 overlay build.sbt is added for Docker image construction, and assembly output names now include the Scala binary version.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Score
- CAI 57 → 56 (-0.4)
- Rubric changed (rubric-2026.09.8 → rubric-2026.09.16) — scores are not directly comparable.
Lenses
- Code Health 97 → 97 (+0.0)
- Architecture 100 → 78 (-21.9)
- Maturity 55 → 55 (+0.0)
- Readiness 45 → 48 (+2.6)
- Security 64 → 64 (+0.0)
Resolved (3)
- Documentation: no architecture or design documentation (README.md)
- Documentation: no installation or build instructions (README.md)
- Documentation: no usage examples (README.md)
New (7)
- No ADRs found
- Outdated: org.apache.commons:commons-text
- Outdated: org.apache.logging.log4j:log4j-api
- Outdated: org.apache.logging.log4j:log4j-core
- Outdated: org.apache.logging.log4j:log4j-slf4j-impl
- Outdated: org.influxdb:influxdb-java
- Projects may be oversized for their cohesion
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
YotpoLtd/metorikku was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 28 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit b15d77c46e760c30a4f2ac70921c7f0c539a56c3 — the exact code this score is about.
- Scored under rubric-2026.09.16 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-d46da229e3fd.