bigdatagenomics/adam
60.2
Adequate · 28 September 2026
39.2k
lines of production code
Scala
primary language
2
measurements over time
What this system is
ADAM is a Spark-based genomic data analysis platform that provides tools for loading, transforming, and analyzing large-scale genomic datasets. It supports a wide range of bioinformatics operations, including read alignment, variant calling, k-mer counting, and coverage calculation, while handling standard formats like FASTQ, BAM, SAM, VCF, and GTF. The system offers programmatic APIs for Java, Python, and R, enabling users to integrate genomic data processing into diverse analytical workflows.
How it got here
2013–2014 — Project scaffolding and core architecture
23 changes.
This period established the ADAM project through initial repository setup, documentation, and governance files. It focused on restructuring the codebase into a multi-module Maven architecture with updated Spark and Scala dependencies, while introducing core genomic data models, CLI commands, and robust FASTQ input handling.
2015–2021 — Dataset API migration and language bindings
18 changes.
This period focused on migrating the core architecture from RDD-based abstractions to a new GenomicDataset layer, accompanied by extensive refactoring of the Java, Python, and R APIs to align with this change. The work also involved establishing standalone package structures for Python and R, enhancing code generation capabilities, and significantly expanding test coverage across all supported data formats and operations.
Features
Add --version command to display build metadata
The CLI now supports a --version option that displays detailed build information, including the ADAM version, Apache Spark version, Scala version, Git commit ID, and build timestamp. This is implemented via a new About.java template class that injects these values from Maven properties during the build process.
adam-cli/src/main/java-templates · high confidence
Initial project scaffolding and documentation
The repository has been initialized with core project documentation and configuration files. This includes a comprehensive README.md describing ADAM as a Spark-based genomic data analysis tool, along with installation instructions for Conda, Homebrew, and Docker. The project also introduces standard open-source governance files: a Code of Conduct (Contributor Covenant v1.4), Contributor Guidelines, Support information, and an Apache 2.0 License. Additionally, a CHANGES.md file is added to track release history starting from version 1.0.1, and a .gitignore file is configured to exclude build artifacts and common temporary files.
(repo-wide) · high confidence
Initial release of the standalone adam-python package
This change introduces the \adam-python\ directory, establishing a standalone Python package structure for ADAM. It adds the necessary build infrastructure, including a \setup.py\ for packaging, a \Makefile\ for development and testing workflows, and a \version.py\ file (currently set to 1.1a0). This enables users to install and manage the Python bindings independently via pip, separate from the core Scala/Java build.
adam-python · high confidence
Introduce command-grouped CLI and new genomic data structures
The ADAM command-line interface is reorganized into logical groups (ADAM ACTIONS, CONVERSION OPERATIONS, PRINT) to improve discoverability, and new commands for counting k-mers and slices are added. Internally, the codebase introduces the GenomicDataset abstraction to wrap genomic RDDs and Datasets with metadata, along with supporting structures like GenomeBins for partitioning and GenomicPositionPartitioner for spatially-coherent data distribution.
repository · high confidence
Introduces Smith-Waterman alignment algorithm with configurable scoring
Adds the Smith-Waterman pairwise alignment implementation to the indel realigner, providing a new \SmithWaterman\ trait that handles core logic like scoring matrix construction and traceback, along with a \SmithWatermanGapScoringFromFn\ implementation that allows custom residue-specific scoring functions. This enables more accurate local alignments by supporting flexible gap and match/mismatch scoring strategies within the indel realignment process.
adam-core/src/main/scala/org/bdgenomics/adam/algorithms/smithwaterman · high confidence
New CLI commands for k-mer counting, coverage, and data inspection
The CLI now includes dedicated commands for counting k-mers from reads and slices (\countKmers\, \countSliceKmers\), calculating genomic coverage with optional strand-specific filtering and collapsing (\coverage\), and inspecting raw ADAM/Parquet files as JSON (\print\). Additionally, the \flagstat\ command now supports writing its output to a file via the \-o\ flag, and the \adam2fastq\ command exposes options for handling paired-end FASTQ output, validation stringency, and quality score preservation.
adam-cli/src/main/scala · high confidence
New genomic data models and lookup tables for variant processing
This change introduces several new core classes in the \org.bdgenomics.adam.models\ package to support genomic data handling. \ReferencePosition\ provides a concrete representation of a single genomic locus, including serialization support. \VariantContext\ aggregates a variant, its position, and associated genotypes into a single unit. Additionally, \IndelTable\ and \SnpTable\ are added to efficiently store and query known indel and SNP variants by reference region, while \NonoverlappingRegions\ provides infrastructure for computing non-overlapping genomic partitions. \SAMFileHeaderWritable\ is also included to handle serialization of SAM file headers.
adam-core/src/main/scala/org/bdgenomics/adam/models · high confidence
New utility and migration scripts for build and data processing
This change introduces several new shell and Python scripts to the project. The \move\_to\_scala\_2.11.sh\ and \move\_to\_scala\_2.12.sh\ scripts allow developers to switch the project's Scala version between 2.11 and 2.12 by updating POM files. Similarly, \move\_to\_spark\_2.sh\ and \move\_to\_spark\_3.sh\ facilitate switching between Spark 2 and Spark 3 dependencies. A new \release.sh\ script automates the release process, including changelog generation, version bumping, and performing releases for Spark 3 with Scala 2.12. Additionally, \fastq-interleaver.py\ provides a utility to interleave paired-end FASTQ files, and \make-flag-values-sam.py\ generates test SAM files covering all valid SAM flag values.
scripts · high confidence
New utility classes for shell output, attribute parsing, and file handling
This change introduces a suite of new utility objects in the \org.bdgenomics.adam.util\ package to support the \adam-shell\ and improve file management. \ADAMShell\ provides new \printAlignmentAttributes\, \printFeatureAttributes\, and \printFormatFields\ methods for displaying data in the shell. \AttributeUtils\ adds support for parsing all numeric array (TYPE=B) tags and fixes SAM attribute conversion bugs. \FileExtensions\ is made visible for downstream use, and \FileMerger\ along with \ParallelFileMerger\ handles merging sharded files, including a workaround for BAM files hosted on the S3 archive. Additional utilities include \GenomeFileReader\ for Bedtools genome files, \ParquetFileTraversable\ for iterating Parquet records, \PhredUtils\ for log-space probability conversions, and \TextRddWriter\ for writing RDDs to text.
adam-core/src/main/scala/org/bdgenomics/adam/util · high confidence
Release of bdgenomics.adam R package version 1.0.1
The bdgenomics.adam R package is now available at version 1.0.1, providing an R frontend for the ADAM genomics analysis platform. This release introduces a comprehensive API for loading, saving, and transforming genomic data types—including alignments, variants, genotypes, coverage, and sequences—via SparkR integration. Users can now perform operations such as region joins, sorting, filtering, and conversion between R data frames and ADAM datasets directly from R.
adam-r/bdgenomics.adam · high confidence
Architecture
Refactor RDD package to dataset package
The \org.bdgenomics.adam.rdd\ package has been refactored and moved to \org.bdgenomics.adam.ds\. This change introduces new foundational components for the dataset layer, including \ADAMParquetInputFormat\ to ensure deterministic read ordering, \InFormatter\ and \OutFormatter\ traits for handling data streaming to and from external processes, and specific formatters for fragment data such as \InterleavedFASTQInFormatter\, \Tab5InFormatter\, and \Tab6InFormatter\. Additionally, the refactor includes new output formats for BAM, CRAM, and SAM files (\ADAMBAMOutputFormat\, \ADAMCRAMOutputFormat\, \ADAMSAMOutputFormat\) that wrap Hadoop-BAM to handle header writing, as well as utilities for parsing GTF/GFF features and writing VCF/SAM headers.
adam-core/src/main/scala/org/bdgenomics/adam/ds · high confidence
Behavioural changes
Added empty Assembly object to support Maven artifact generation
An empty Assembly object has been added to the adam-assembly module. This change enables the Maven build process to successfully generate sources and Javadoc artifacts for this module, resolving build configuration issues without introducing new product functionality.
adam-assembly · high confidence
Code generator now produces projection enums and improved Scala product classes
The adam-codegen module has been updated to automatically generate projection enums for schema fields (DumpSchemasToProjectionEnums) and to produce Scala case classes that implement Product with robust Avro conversion logic (DumpSchemasToProduct). This change enhances the generated Scala code in the sql package by adding field enumeration support and improving how records are mapped to and from Avro schemas, supporting more complex field types like unions, arrays, and maps.
adam-codegen · high confidence
Java API refactors GenomicRDD to GenomicDataset and adds conversion wrappers
The Java API in the \org.bdgenomics.adam.api.java\ package has been updated to align with the internal shift from \GenomicRDD\ to \GenomicDataset\. This change introduces new converter classes (e.g., \GenomicDatasetConverters\, \GenomicRDDConverters\) and updates \JavaADAMContext\ to expose Java-friendly methods for loading and converting genomic data types (such as Coverage, Features, Fragments, Alignments, etc.) using the new \GenomicDataset\ abstraction. Additionally, a \DataFrameConversionWrapper\ is added to the Python API to facilitate DataFrame transformations.
adam-apis/src/main · high confidence
Python API restructured to use namespace packages and Dataset-based APIs
The Python package has been reorganized into a namespace package structure (bdgenomics.adam) and migrated from RDD-based to Dataset-based APIs. This change introduces new dataset wrappers (AlignmentDataset, CoverageDataset, etc.) and updates the ADAMContext to initialize via SparkSession, while also adding support for validation stringency and indexed BAM loading.
adam-python/bdgenomics · high confidence
R API refactored to use GenomicDataset and renamed AlignmentRecord to Alignment
The R package has been refactored to replace the legacy GenomicRDD abstraction with a new GenomicDataset class hierarchy, and renamed the AlignmentRecord type to Alignment. Users will now interact with datasets such as AlignmentDataset, ReadDataset, and SequenceDataset, and should update any code referencing the previous GenomicRDD or AlignmentRecord names to use the new dataset classes and alignment terminology.
adam-r/bdgenomics.adam/R · high confidence
Refactored FASTQ input handling with splittable compressed support
The FASTQ input format in adam-core has been refactored to improve robustness and performance when reading FASTQ files. The new implementation introduces a configurable maximum read length (defaulting to 10,000 bp) to prevent issues with unusually long reads. It also adds support for BGZF-compressed FASTQ files, allowing them to be split and processed in parallel, which significantly speeds up large-scale data ingestion. Additionally, the reader now correctly handles interleaved FASTQ formats and ensures that multi-line records at split boundaries are read completely, preventing data loss or corruption during distributed processing.
adam-core/src/main/java · high confidence
Refactored INDEL realignment consensus generation and added read-based INDEL normalization
The consensus generation logic for INDEL realignment has been refactored into a modular system with three distinct strategies: generating consensus from known variants, from read alignments, and from Smith-Waterman local realignments. A key behavioral improvement is the introduction of left normalization for INDELs found in read alignments, which standardizes indel positions to improve realignment accuracy. Additionally, the system now supports combining multiple consensus strategies via a union operator, allowing users to leverage both known variant databases and observed read evidence simultaneously.
adam-core/src/main/scala/org/bdgenomics/adam/algorithms/consensus · high confidence
Refactored bin scripts with new launcher utilities and improved argument handling
The bin directory has been restructured to improve reliability and usability. New helper scripts (find-adam-assembly.sh, find-adam-egg.sh, find-adam-home, find-spark.sh) centralize path resolution for ADAM JARs, Python eggs, and Spark binaries, supporting both standard and pip-installed environments. The adam-submit script now correctly splits Spark and ADAM arguments using a double-dash separator and allows the main class to be overridden via the ADAM\_MAIN environment variable. Additionally, adam-shell, adamR, and pyadam have been updated to use these helpers, ensuring consistent Kryo serialization configuration and proper classpath setup.
bin · high confidence
Refactored conversion classes and added configurable variant annotation nesting
The converters package has been reorganized to improve modularity and usability. AlignmentRecord conversion is now handled by a new AlignmentConverter class, and FASTQ parsing is managed by FastqRecordConverter. Variant and genotype conversions are centralized in VariantContextConverter, which now supports a configurable property (org.bdgenomics.adam.converters.VariantContextConverter.NEST\_ANN\_IN\_GENOTYPES) to optionally populate the variant.annotation field within Genotype records. Additionally, DefaultHeaderLines provides a standardized set of VCF INFO and FORMAT header definitions, and TranscriptEffectConverter handles parsing of Ensembl VEP ANN attributes into TranscriptEffect objects.
adam-core/src/main/scala/org/bdgenomics/adam/converters · high confidence
Refactored schema projection to embed Avro schemas in field enumerations
The projection system in the ADAM core library has been restructured to support more robust schema handling. A new \FieldEnumeration\ base class and \FieldValue\ trait allow record-specific enumerations to embed their associated Avro schemas directly within each field value, replacing the previous approach that relied on static schema fields. The \Projection\ object now extracts the schema from the first provided \FieldValue\ to generate the projected schema, and a new \Filter\ helper object has been added to simplify the creation of projections that exclude specific fields.
adam-core/src/main/scala/org/bdgenomics/adam/projections · high confidence
Test coverage
1 commit adding/updating tests in adam-core/src/test/resources/bams; 2 commits adding/updating tests in adam-core/src/test/resources/indexed\_bams; Added R package test suite for ADAM data operations; Added VCF test fixtures for directory loading and position 0 support; Added integration tests for Java ADAM Context data I/O; Added invalid VCF test resources for validation stringency and INFO flag handling; Added multi-allelic gVCF test fixtures; Added multiallelic gVCF test resource; Added test coverage for ADAM CLI commands; Added test coverage for consensus generation and indel normalization; Added test resources for BQSR and k-mer counting; Added test resources for GIAB VCF, HLA reference, and Ensembl GTF annotations; Added test suites for ADAM utility components; Added test suites for alignment, FASTQ, fragment, transcript effect, and variant context converters; Added test suites for the refactored ds package; Added tests for FASTQ input format readers; Added tests for alignment and CIGAR string handling; Added unit tests for Smith-Waterman alignment algorithm; Added unit tests for core genomic models.
Dependencies
ADAM 1.1-SNAPSHOT: New Spark 3.5.5 and Scala 2.12 build with multi-module structure
This release introduces a new build configuration targeting Apache Spark 3.5.5 and Scala 2.12, replacing previous Spark 2.x and Scala 2.10/2.11 builds. The project has been restructured into a multi-module Maven setup (adam-core, adam-cli, adam-apis, adam-codegen, adam-assembly, adam-distribution, adam-python, and adam-r) to better organize the codebase. Key dependency updates include Avro 1.11.4, Parquet 1.13.1, Hadoop-BAM 7.10.0, and bdg-formats 1.0.3. The build now uses maven-shade-plugin for assembly and includes dedicated modules for Python and R API bindings, which are built via exec-maven-plugin. The parent POM enforces Java 1.8+ and Maven 3.6.3+ requirements.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Score
- CAI 60 → 60 (-0.1)
- Rubric changed (rubric-2026.09.8 → rubric-2026.09.16) — scores are not directly comparable.
Lenses
- Code Health 91 → 91 (+0.0)
- Architecture 98 → 95 (-3.4)
- Maturity 53 → 53 (+0.0)
- Readiness 52 → 52 (+0.2)
- Security 72 → 72 (+0.0)
New (2)
- No ADRs found
- Projects may be oversized for their cohesion
Architecture
- Unchanged — 0 containers · 1 contexts · 0 edges
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
bigdatagenomics/adam was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 28 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 6f2aec83d2289f48104801a88c326b894ec7df2c — the exact code this score is about.
- Scored under rubric-2026.09.16 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-2d9048c36d26.