awslabs/deequ
58.0
Adequate · 27 September 2026
22k
lines of production code
Scala
primary language
4
measurements over time
What this system is
This system is a data quality verification and analysis library for Apache Spark, designed to profile datasets, enforce quality constraints, and compare data integrity. It provides a fluent API for running checks, computing metrics, and persisting results to memory or file systems, while also supporting rule definitions via the DQDL syntax. The library enables users to detect anomalies, suggest constraints, and perform row-level evaluations to identify specific data violations.
How it got here
2018 — Initial Deequ library development
12 changes.
This period marks the initial import and release of the Deequ data quality library for Apache Spark, establishing the core verification and analysis engine. It introduces foundational components including fluent APIs for constraint checking, column profiling, and constraint suggestion, alongside infrastructure for persisting metrics via in-memory and file-system repositories. The work also includes essential compatibility shims for Spark 2.x and comprehensive test coverage to validate the new data quality capabilities.
2019–2026 — DQDL integration and dataset comparison
8 changes.
This period focused on integrating the Data Quality Definition Language (DQDL) to enable row-level data quality evaluation and composite rule execution. It also introduced experimental utilities for dataset comparison, including column and schema matching, alongside fixes for KLL sketch accuracy and column name escaping.
Features
Added script to generate a size-bounded knowledge base of source code
A new Python script, \src/scripts/generate\_kb.py\, has been added to automatically generate a comprehensive knowledge base of the repository's Scala source code. The script implements a two-tier strategy to ensure the output stays within a 500,000-character budget: small files (100 lines or fewer) are included in full, while larger files are summarized by extracting signatures, implementation-significant lines (such as method declarations and specific SQL patterns), and the first line of documentation blocks. The generated output also includes the README, relevant build configuration from \pom.xml\, and a list of test files.
src/scripts · high confidence
Initial import of Deequ analyzer runners and fluent analysis API
This change introduces the core execution engine for Deequ's data quality analysis. It adds a fluent \AnalysisRunBuilder\ API that allows users to construct analysis runs by adding analyzers, configuring a metrics repository for result reuse, and setting up Spark sessions for JSON output. It also includes the \KLLRunner\ for computing approximate quantile sketches using KLL algorithms across various numeric data types, and a set of specific \MetricCalculationException\ classes to handle errors during metric computation.
src/main/scala/com/amazon/deequ/analyzers/runners · high confidence
Initial import of Deequ verification and analysis core
This change introduces the core data quality verification and analysis engine for the repository. It adds the VerificationSuite and VerificationRunBuilder to provide a fluent API for running data checks and collecting metrics, along with the VerificationResult class to expose check statuses, constraint details, and row-level results. The diff also establishes the foundational analyzer framework (Analyzer, State, Analysis) and includes initial implementations for key analyzers such as ApproxCountDistinct, ApproxQuantile, ColumnCount, ColumnExists, Completeness, and Compliance, enabling users to profile data and enforce quality constraints.
repository · high confidence
Initial import of analysis result serialization and multi-result loading components
This change introduces the core infrastructure for persisting and retrieving Deequ analysis results. It adds \AnalysisResultSerde.scala\, which implements JSON serialization and deserialization for analysis results, metrics, analyzers, and distribution data using Gson, ensuring that complex metric structures can be stored and restored accurately. It also adds \MetricsRepositoryMultipleResultsLoader.scala\, a trait and companion object that enables loading multiple analysis results from a repository, filtering them by tags or date ranges, and combining them into unified DataFrames or JSON strings with consistent column schemas.
src/main/scala/com/amazon/deequ/repository · high confidence
Initial release of Deequ data quality library
This entry introduces the initial codebase for Deequ, a library for defining and verifying data quality in Apache Spark. It adds a \ConstrainableDataTypes\ enumeration to categorize column types (e.g., String, Numeric, Boolean) for constraint definitions. It includes \DfsUtils\ to handle reading and writing binary and text files to distributed file systems like S3 or HDFS, ensuring correct path qualification. Additionally, it provides an \InMemoryMetricsRepository\ implementation backed by a concurrent hash map, allowing users to save, load, and filter analysis results (metrics) by tags, analyzers, and date ranges in memory.
src/main/scala/com/amazon/deequ/constraints, src/main/scala/com/amazon/deequ/io, src/main/scala/com/amazon/deequ/repository/memory · high confidence
Initial release of Deequ example applications and documentation
This change introduces the initial set of executable Scala examples and accompanying Markdown documentation for the Deequ library. The examples demonstrate core data quality capabilities, including basic verification checks, automatic constraint suggestion using heuristic rules, anomaly detection via rate-of-change strategies, and stateful metrics computation for incremental and partitioned data updates. Supporting utility classes and configuration helpers are also included to enable these demonstrations.
src/main/scala/com/amazon/deequ/examples · high confidence
Introduce FileSystemMetricsRepository for persisting analysis results
Users can now store and retrieve Deequ analysis metrics using a file-system-based repository. This new component allows saving analysis results to a specified path (supporting DFS/S3), filtering loaded results by date ranges, tags, or specific analyzers, and ensures atomic writes via temporary files. It provides a concrete implementation of the MetricsRepository interface for scenarios where metrics need to be persisted to local or distributed file systems.
src/main/scala/com/amazon/deequ/repository/fs · high confidence
New DQDL rule execution engine with row-level results and composite rule support
The DQDL execution layer has been replaced with a new executor architecture that supports a broader set of data quality rules and provides row-level evaluation results. Users can now define composite rules using AND/OR operators to combine multiple checks, and the system returns detailed pass/fail/skip outcomes for individual rows via the \DataQualityEvaluationResult\ column. The engine natively executes rules including DatasetMatch, SchemaMatch, RowCountMatch, AggregateMatch, ReferentialIntegrity, DataFreshness, ColumnDataType, ColumnNamesMatchPattern, and CustomSql, routing each to its specific executor while handling unsupported rules gracefully.
src/main/scala/com/amazon/deequ/dqdl/execution · high confidence
New experimental dataset comparison utilities
This location introduces a new experimental API for comparing datasets, centered around the \DataSynchronization\ object which provides \columnMatch\ and \columnMatchRowLevel\ methods to verify column equality and annotate rows with match outcomes. The package also adds \ReferentialIntegrity\ for subset checks and row-level annotation, alongside new rule objects \RowCountMatch\ and \SchemaMatch\ that allow users to assert ratios for row counts and schema column matches respectively.
src/main/scala/com/amazon/deequ/comparison · high confidence
New fluent builder API for Column Profiler runs
A new \ColumnProfilerRunBuilder\ class has been introduced to provide a fluent, chainable API for configuring and executing column profiling runs. This builder allows users to easily set options such as caching inputs, restricting analysis to specific columns, enabling KLL sketch profiling, and defining predefined data types before triggering the analysis via the \run()\ method. It also supports integration with metrics repositories and Spark sessions for saving results.
src/main/scala/com/amazon/deequ/profiles · high confidence
Support for DQDL rule translation and execution
Users can now define data quality rules using the Data Quality Definition Language (DQDL) and have them automatically translated into executable Deequ checks. This change introduces a comprehensive rule translator that supports a wide range of rule types, including statistical aggregations (Mean, StandardDeviation, Variance, Skewness, Kurtosis, Entropy), completeness and uniqueness checks, referential integrity, dataset and schema matching, and composite rules with AND/OR logic. The system also handles outcome translation, converting Deequ check results back into a standardized format for reporting.
src/main/scala/com/amazon/deequ/dqdl/translation · high confidence
Support for row-level data quality evaluation via DQDL
The \EvaluateDataQuality\ entry point now includes a \processRows\ method that evaluates DQDL rulesets and returns per-row pass/fail/skip results alongside aggregated outcomes. This allows users to identify exactly which rows violate specific quality constraints, supporting detailed debugging and filtering, in addition to the existing \process\ method which returns only aggregated rule results.
src/main/scala/com/amazon/deequ/dqdl · high confidence
Behavioural changes
Compatibility layer for Spark 2.x internal API changes
The library now includes internal compatibility shims to support Spark 2.2, 2.3, and 2.4. A new Java utility (\AttributeReferenceCreation\) uses reflection to invoke the \AttributeReference\ constructor with the correct parameter signature for each specific Spark minor version, resolving API incompatibilities. Additionally, constants required by the HyperLogLogPlus implementation are copied into the codebase (\HLLConstants\) because their locations differ across Spark 2.x versions, ensuring stable approximate distinct-count and quantile calculations.
src/main/scala/com/amazon/deequ/analyzers/catalyst · high confidence
Introduce fluent ConstraintSuggestionRunBuilder and extensible ConstraintRule API
The constraint suggestion workflow now uses a new fluent builder class, ConstraintSuggestionRunBuilder, which allows users to configure runs with options such as train-test splitting, column restrictions, KLL sketch parameters, predefined data types, and output file overwriting. The underlying suggestion logic has been refactored to use an abstract ConstraintRule base class, enabling extensible, rule-based constraint generation that can be selectively applied to specific columns based on their profiles.
src/main/scala/com/amazon/deequ/suggestions · high confidence
Support for multi-column completeness checks with conditional filtering
Users can now evaluate data completeness across combinations of columns rather than just single columns, enabling checks that verify if any or all columns in a set contain non-null values. Additionally, the check builder now supports applying a filter (via a \where\ clause) to the last constraint, allowing completeness validations to be scoped to specific subsets of data before evaluation.
src/main/scala/com/amazon/deequ/checks · high confidence
Fixes
Added column name escaping and unescaping utilities
A new \ColumnUtil\ object has been added to the utilities package, providing \escapeColumn\ and \removeEscapeColumn\ methods. These functions handle the safe quoting of column names with backticks, ensuring that special characters are properly escaped (e.g., replacing internal backticks with double backticks) and that already-quoted names are preserved, which helps prevent syntax errors when referencing columns with non-standard names in SQL-like contexts.
src/main/scala/com/amazon/deequ/utilities · high confidence
Fixes under-counting of repeated data points in KLL sketch rank and CDF calculations
The KLL sketch implementation in QuantileNonSample now correctly handles repeated data points when computing rank maps and cumulative distribution functions. Previously, aggregating items into a map by key discarded duplicate entries, leading to under-counting; the new logic folds over sorted pairs to accumulate weights, ensuring accurate rank and CDF outputs for datasets with duplicate values.
src/main/scala/com/amazon/deequ/analyzers · high confidence
Test coverage
Added Titanic dataset for column profiling tests; Added test infrastructure and coverage for KLL sketches, verification results, and Spark session management.
Dependencies
Deequ 2.1.0 with Spark 3.5 and Glue DQDL support
This release updates the build manifest to target Spark 3.5.7 and Scala 2.12.10, raising the minimum Java requirement to 11. It introduces a new dependency on the Glue Data Quality Definition Language (DQDL) parser (version 1.0.7), enabling users to define data quality checks using DQDL syntax, and upgrades the Breeze library to version 2.1.0.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
This is the PUBLIC form of this artifact. Findings are listed in full, but the details of SECURITY findings — which rule fired, in which file, on which line, and how to fix it — are deliberately withheld, and any secret-scanner results are excluded entirely. Where detail is absent here it was REMOVED FOR PUBLICATION; it is not missing from the analysis. The complete artifact is available from the repository owner.
Score
- CAI 42 → 58 (+15.7)
- Rubric changed (rubric-2026.08.15 → rubric-2026.09.15) — scores are not directly comparable.
Lenses
- Code Health 100 → 92 (-8.3)
- Architecture 69 → 95 (+26.2)
- Maturity 57 → 47 (-9.6)
- Readiness 24 → 60 (+36.5)
- Security 50 → 60 (+9.7)
Resolved (8)
- Coverage not measured — test suite did not build
- Dimension evaluation failed
- High: security finding (details withheld)
- High: security finding (details withheld)
- No artifact signing
- No exposed public API
- No tests found
- Test reliability not included
New (82)
- AnalyzerDeserializer.deserialize (cyclomatic 40) (src/main/scala/com/amazon/deequ/repository/AnalysisResultSerde.scala)
- AnalyzerSerializer.serialize (cyclomatic 42) (src/main/scala/com/amazon/deequ/repository/AnalysisResultSerde.scala)
- Change coupling: ColumnProfiler.scala ↔ ConstraintSuggestionRunner.scala (src/main/scala/com/amazon/deequ/profiles/ColumnProfiler.scala)
- ColumnDataTypeRule.castColumnToSparkType (cognitive 20) (src/main/scala/com/amazon/deequ/dqdl/translation/rules/ColumnDataTypeRule.scala)
- ColumnDataTypeRule.toExecutableRule (cognitive 21) (src/main/scala/com/amazon/deequ/dqdl/translation/rules/ColumnDataTypeRule.scala)
- ColumnLengthRule.mkColumnLengthCheck (cognitive 16) (src/main/scala/com/amazon/deequ/dqdl/translation/rules/ColumnLengthRule.scala)
- ColumnLengthRule.mkColumnLengthCheck (cyclomatic 17) (src/main/scala/com/amazon/deequ/dqdl/translation/rules/ColumnLengthRule.scala)
- ColumnProfiler.extractNumericStatistics (cognitive 52) (src/main/scala/com/amazon/deequ/profiles/ColumnProfiler.scala)
- ColumnProfiler.extractNumericStatistics (cyclomatic 27) (src/main/scala/com/amazon/deequ/profiles/ColumnProfiler.scala)
- ColumnValuesRule.constructComplianceCondition (cognitive 26) (src/main/scala/com/amazon/deequ/dqdl/translation/rules/ColumnValuesRule.scala)
- ColumnValuesRule.mkNumericCheck (cognitive 41) (src/main/scala/com/amazon/deequ/dqdl/translation/rules/ColumnValuesRule.scala)
- ColumnValuesRule.mkNumericCheck (cyclomatic 31) (src/main/scala/com/amazon/deequ/dqdl/translation/rules/ColumnValuesRule.scala)
- ColumnValuesRule.validateDateOperandCount (cognitive 17) (src/main/scala/com/amazon/deequ/dqdl/translation/rules/ColumnValuesRule.scala)
- Concentrated knowledge decay
- DQDLRuleTranslator.toExecutableRule (cognitive 22) (src/main/scala/com/amazon/deequ/dqdl/translation/DQDLRuleTranslator.scala)
- DQDLRuleTranslator.toExecutableRule (cyclomatic 26) (src/main/scala/com/amazon/deequ/dqdl/translation/DQDLRuleTranslator.scala)
- DataSynchronization.columnMatchRowLevel (cognitive 21) (src/main/scala/com/amazon/deequ/comparison/DataSynchronization.scala)
- DataTypeHistogram.determineType (cognitive 17) (src/main/scala/com/amazon/deequ/analyzers/DataType.scala)
- Dependency hygiene PARTLY measured — Maven/Gradle declarations read, no dependency graph resolved
- Documentation: no installation or build instructions (README.md)
- …and 62 more
Changes since last survey
- 4 commits — 3 feature/other, 1 fixes
By area
- (root) — 3 commits
- src/main — 1 commit
Notable commits
- fix: Fix KLL rank map and CDF under-counting repeated data points (#765)
- change: Require Java 11 and upgrade dqdl to 1.0.7 (#687)
- change: Update version in pom.xml to 2.1.0-spark-3.5 (#766)
- change: docs: add Spark and Deequ compatibility table (#763)
Architecture
- Containers 0 added · 0 removed · contexts 1 added · 0 removed · edges 0 added · 0 removed
Added bounded contexts (1)
- repository
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
awslabs/deequ was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 27 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit d7f778310cadcfaf7ad641bdfe07368416cc7c64 — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-d00c643c3f66.