Skip to content
CAI
Software that uses CAICheck a score

fulcrumgenomics/fgbio

61.0

Adequate · 20 September 2026

28k

lines of production code

Scala

primary language

1

measurement over time

CAI band scale
CAI lens gauges

What this system is

This system is a bioinformatics toolkit for processing and analyzing genomic sequencing data, primarily written in Scala. It provides command-line tools and libraries for handling core file formats like SAM/BAM, FASTQ, and VCF, including operations such as alignment, filtering, normalization, and variant calling. The system also supports specialized molecular indexing, UMI consensus generation, and quality control metrics for RNA-seq and duplex sequencing workflows.

How it got here

2015–2016 — Initial release and core tool development

24 changes.

This period established the project's foundational infrastructure, including build workflows, release processes, and comprehensive test coverage. It focused on developing a broad suite of core bioinformatics tools for processing BAM, FASTQ, FASTA, and VCF files, alongside utilities for UMI annotation and molecular index generation.

2017 — New tools and API development

17 changes.

This period focused on expanding the toolkit with new features for Illumina data parsing, somatic variant filtering, and high-precision statistical calculations. Significant infrastructure work included the initial release of a native SAM/BAM/CRAM API and optimizations to the alignment engine. The work was complemented by comprehensive test coverage and the addition of R-based quality control plotting scripts for various sequencing metrics.

2018–2022 — Scala VCF API and Pileup Infrastructure

8 changes.

This period focused on developing core genomic data processing infrastructure, notably introducing a native Scala VCF API to replace Java-centric approaches and implementing flexible random-access and streaming pileup builders for BAM files. The work also included adding coordinate-based ordering for genomic locatables and expanding test coverage for these new components to ensure robust handling of variant data and alignment pileups.

Features

Added DNA/RNA secondary structure prediction utilities

New utility classes have been added to the \com.fulcrumgenomics.util\ package to support DNA and RNA folding analysis. \DnaFoldPrediction\ serves as a data container for sequence, structure, and free energy (deltaG) results, while \DnaFoldPredictor\ provides a Java interface to the ViennaRNA RNAFold executable. The predictor manages a background process to calculate minimum free energy secondary structures for input sequences, allowing users to integrate external RNA folding tools directly into their bioinformatics workflows.

src/main/java · high confidence

Added ERCC QC plotting script and standard mixture data

The \CollectErccMetrics\ tool now includes an R script (\CollectErccMetrics.R\) that generates quality control plots for ERCC spike-in controls. This script takes metrics data and outputs PDF plots showing the correlation between expected and observed counts (both raw and normalized) using log2 scaling and linear regression. It also adds the \ErccStandardMixtures.txt\ resource file, which contains the reference concentrations and expected fold-change ratios for the ERCC Standard Mixtures (Mixes 1-4), enabling accurate comparison of observed sequencing data against known standards.

src/main/resources/com/fulcrumgenomics/rnaseq · high confidence

Added R script for generating QC plots in CollectDuplexSeqMetrics

The \CollectDuplexSeqMetrics\ tool now includes an R script (\CollectDuplexSeqMetrics.R\) that generates a comprehensive PDF report of quality control plots for duplex sequencing data. This script visualizes family size distributions, duplex yield metrics (including actual vs. ideal duplex ratios), UMI representation, and read distribution among tag families, allowing users to visually assess the quality and characteristics of their duplex sequencing runs.

src/main/resources/com/fulcrumgenomics/umi · high confidence

Added coordinate-based ordering for genomic locatable objects

Users can now sort genomic intervals (Locatable instances) consistently using a new ordering mechanism that respects the sequence dictionary. This ordering prioritizes contig index, followed by start position, and then end position, ensuring that variants and intervals are arranged in a biologically meaningful order suitable for downstream processing like pileup building.

src/main/scala/com/fulcrumgenomics/coord · high confidence

Initial project setup and release infrastructure

The repository has been initialized with the core project structure, including the MIT license, contributing guidelines, and a version definition file (version.sbt) set to 4.1.2-SNAPSHOT. This change establishes the build and release workflow, configuring Git LFS for large genomic data files (BAM, VCF), integrating Codecov for test coverage reporting, and setting up git-cliff for automated changelog generation.

(repo-wide) · high confidence

Initial release of the Scala VCF API

Introduces a new Scala-native API for reading and writing VCF and BCF files, replacing the previous Java-centric approach. The new \VcfSource\ and \VcfWriter\ classes provide efficient iteration and writing capabilities, while \ArrayAttr\ and \GenotypeMap\ offer specialized data structures for handling multivalued attributes and genotype samples. Internal conversion logic in \VcfConversions\ bridges the Scala models with HTSJDK, ensuring correct handling of header metadata, counts, and field types.

src/main/scala/com/fulcrumgenomics/vcf/api · high confidence

Initial release of the fgbio SAM/BAM/CRAM API

This change introduces the core \com.fulcrumgenomics.bam.api\ package, providing a Scala-native wrapper around HTSJDK for reading and writing SAM, BAM, and CRAM files. The new API includes \SamSource\ for streaming and querying records, \SamWriter\ for output with optional sorting and indexing, and \SamRecord\ for convenient access to read attributes (such as flags, CIGAR strings, and mate information) without exposing the underlying HTSJDK types. It also adds \HeaderHelper\ for accessing header metadata and \SamIterator\ for efficient iteration, including support for CRAM conversion.

src/main/scala/com/fulcrumgenomics/bam/api · high confidence

Introduce SampleSheet parser with strict uniqueness validation

Adds a new \SampleSheet\ parser for Illumina Experiment Manager sample sheets (typically MiSeq) that enforces strict uniqueness constraints: sample identifiers must be unique, and the combination of sample name and library identifier must also be unique across samples. The parser automatically sets the library ID to the sample ID if the library ID field is missing, handles case-insensitive header lookups, and skips empty lines at the end of the data section. It also supports filtering samples by a specific lane number.

src/main/scala/com/fulcrumgenomics/illumina · high confidence

Introduce core FASTQ I/O tools and utilities

Adds a new set of tools and library components for reading, writing, sorting, trimming, and converting FASTQ files. This includes the FastqSource and FastqWriter classes for efficient FASTQ I/O, a SortFastq tool for lexicographic sorting of FASTQ records, a TrimFastq tool to trim reads to a specified length (with optional exclusion of short reads), and a FastqToBam tool to convert FASTQ files into unmapped BAM/SAM/CRAM formats while supporting read structures for extracting UMIs, cell barcodes, and sample indices.

src/main/scala/com/fulcrumgenomics/fastq · high confidence

Introduction of structured command-line tool grouping

The command-line interface now organizes tools into distinct categories (such as SAM/BAM, FASTA, FASTQ, RNA-Seq, UMIs, and Utilities) via the new \ClpGroups\ object. This change improves the user experience by providing a structured and navigable help menu, allowing users to easily locate tools based on the data type they are working with rather than searching through a flat list of commands.

src/main/scala/com/fulcrumgenomics/cmdline · high confidence

New ExtractIlluminaRunInfo tool to parse Illumina RunInfo.xml

A new command-line tool, ExtractIlluminaRunInfo, has been added to the basecalling module. It reads an Illumina RunInfo.xml file and outputs a metric file containing key run details, including the run barcode, flowcell barcode, instrument name, run date, read structure, and number of lanes.

src/main/scala/com/fulcrumgenomics/basecalling · high confidence

New FASTA manipulation tools: HardMaskFasta, SortSequenceDictionary, and UpdateFastaContigNames

This release introduces three new command-line tools for managing FASTA files and sequence dictionaries. HardMaskFasta converts soft-masked (lowercase) bases to hard-masked 'N' characters and standardizes line lengths for compatibility with tools like samtools faidx. SortSequenceDictionary reorders the contigs in a sequence dictionary file to match the order defined in another reference dictionary, with an option to skip or append missing contigs. UpdateFastaContigNames renames contigs in a FASTA file based on a provided sequence dictionary, supports sorting the output by dictionary order, and can append missing contigs from a default FASTA file.

src/main/scala/com/fulcrumgenomics/fasta · high confidence

New R scripts for ErrorRateByReadPosition and FindSwitchbackReads QC plots

Added R scripts to generate quality control plots for the ErrorRateByReadPosition and FindSwitchbackReads tools. The ErrorRateByReadPosition script now supports plotting individual substitution types (e.g., A\>C, G\>T) when the 'collapsed' column is false, allowing for higher sensitivity in error analysis. The FindSwitchbackReads script generates distribution plots for switchback length, offset, and tandem gap metrics.

src/main/resources/com/fulcrumgenomics/bam · high confidence

New VCF manipulation tools and iterators

This update introduces several new capabilities for working with VCF files. The \MakeTwoSampleMixtureVcf\ tool allows users to create simulated tumor or tumor/normal VCFs by in-silico mixing genotypes from two samples at a specified proportion, including support for tumor-only mode and interval restriction. New iterator utilities, \ByIntervalListVariantContextIterator\ and \JointVariantContextIterator\, enable efficient querying of variants overlapping specific intervals and merging multiple variant context streams, respectively. Additionally, the \UpdateVcfContigNames\ tool provides a way to remap contig names in a VCF based on a provided sequence dictionary, and \VariantMask\ offers a compact representation for rapidly querying whether positions are overlapped by variants.

src/main/scala/com/fulcrumgenomics/vcf · high confidence

New high-precision binomial and Fisher's exact test math utilities

Added BinomialDistribution and FishersExactTest classes to the math library. BinomialDistribution provides high-precision binomial probability calculations using unlimited precision BigDecimal to avoid underflow issues found in standard libraries, while FishersExactTest implements Fisher's Exact Test for 2x2 contingency tables, supporting two-sided, greater, and less alternative hypotheses.

src/main/scala/com/fulcrumgenomics/math · high confidence

New internal tool for generating per-tool Markdown documentation

An internal command-line tool (\BuildToolDocs\) has been added to automatically generate Markdown documentation for all available tools. This tool scans specified packages (defaulting to \com.fulcrumgenomics\), groups tools by category, and produces an index file along with individual documentation pages for each tool, including details on arguments, types, and default values. The output can optionally exclude the version string and is written to a specified output directory.

src/main/scala/com/fulcrumgenomics/internal · high confidence

New personal tools for region generation and tag splitting

Added two new personal tools: \GenerateRegionsFromFasta\, which creates freebayes/bamtools-style region specifiers from a FASTA file by dividing sequences into configurable chunk sizes, and \SplitTag\, which splits a single optional SAM/BAM tag into multiple distinct tags based on a specified delimiter. Both tools are registered under the Personal command-line group.

src/main/scala/com/fulcrumgenomics/personal/nhomer · high confidence

New personal tools for stripping Fastq read numbers and summarizing GFF biotypes

Added two new command-line tools in the personal tools section: StripFastqReadNumbers, which removes trailing slash-number suffixes (e.g., /1, /2) from FASTQ read names by clearing the read number field, and SummarizeGff, which processes RefSeq GFF files to produce summary statistics of single-exon versus multi-exon gene counts grouped by gene biotype.

src/main/scala/com/fulcrumgenomics/personal/tfenne · high confidence

New random-access and streaming pileup builders for BAM processing

Added two new pileup builder implementations in the BAM pileup module: RandomAccessPileupBuilder, which uses index-based random access on indexed BAM files, and StreamingPileupBuilder, which uses coordinate-sorted lazy streaming with an internal cache. Both builders support configurable filtering options (minimum mapping/base quality, duplicate/secondary/supplementary alignment inclusion, mapped pairs only) and provide a consistent API for generating pileups at specific genomic loci.

src/main/scala/com/fulcrumgenomics/bam/pileup · high confidence

New somatic variant filtering tool with A-tailing and End Repair artifact detection

The \FilterSomaticVcf\ tool is introduced to apply statistical filters to somatic variant calls by analyzing read pileups. It now supports two distinct artifact filters: the A-tailing Artifact Filter (previously 'End Repair Artifact Filter') and the new End Repair Fill-in Artifact Filter. These filters evaluate whether mismatches at read ends are likely library preparation artifacts rather than true mutations, adding corresponding p-value annotations (ATAP, ERFAP) and FILTER tags to the output VCF. The tool also introduces configurable BAM access patterns (RandomAccess vs. Streaming) to optimize performance based on VCF density and sorting.

src/main/scala/com/fulcrumgenomics/vcf/filtration · high confidence

New tools for BAM normalization, filtering, and technical read detection

This release introduces several new command-line tools for processing BAM files. DownsampleAndNormalizeBam downsamples BAMs to a uniform coverage across specified regions, ensuring every base meets a target depth. FilterBam now supports writing rejected reads to a separate file via the --rejects option, allowing users to inspect filtered-out data. FindTechnicalReads identifies and extracts reads derived from technical sequences (like adapter dimers), with options to tag matches or output all reads. Additional new tools include RandomizeBam for shuffling read order, RemoveSamTags for stripping specific SAM tags, SetMateInformation for fixing paired-end metadata, SortBam for various sort orders including TemplateCoordinate, TrimPrimers for post-alignment primer trimming with optional tag recalculation, and UpdateReadGroups for mapping old read group IDs to new ones.

src/main/scala/com/fulcrumgenomics/bam · high confidence

New tools for CODEC duplex consensus and UMI annotation

This release introduces \CallCodecConsensusReads\, a new tool for generating duplex consensus sequences from CODEC sequencing data, which assembles single-strand consensuses from R1 and R2 reads into a final duplex consensus. It also adds \AnnotateBamWithUmis\, a new tool that annotates existing BAM files with UMIs extracted from separate FASTQ files, supporting multiple FASTQ inputs and optional UMI quality storage. These additions expand the UMI toolkit to support specific sequencing protocols and pre-annotated data workflows.

src/main/scala/com/fulcrumgenomics/umi · high confidence

New tools for picking molecular indices and generating read groups

Added the \PickLongIndices\ tool for efficiently generating large sets of molecular indices with customizable constraints (length, edit distance, GC content, homopolymers, and secondary structure via ViennaRNA), and the \PickIlluminaIndicesCommand\ utility for selecting optimal index sets. Also added \AutoGenerateReadGroupsByName\, which parses Illumina-style read names to automatically create and assign read groups to BAM files.

repository · high confidence

New utility classes for amplicon detection, genomic spans, and interval list I/O

This change introduces several new utility components in the \com.fulcrumgenomics.util\ package. The \AmpliconDetector\ class allows users to identify which amplicon a sequencing read (or read pair) originated from by matching read coordinates against a set of known amplicons, with support for adjusting for clipped bases. New \GenomicSpan\ and \IntervalListSource\/\IntervalListWriter\ classes provide standardized ways to represent genomic intervals and stream-read/write interval list files, including proper handling of sequence dictionaries. Additionally, \Io.scala\ is updated to automatically handle \.bgz\ and \.bgzip\ compression when writing files, and \Metric.scala\ gains support for collection fields and improved serialization of enum entries.

src/main/scala/com/fulcrumgenomics/util · high confidence

Behavioural changes

Optimized alignment performance via specialized integer matrix

The alignment engine now uses a specialized \SimpleIntMatrix\ implementation for storing integer values during sequence alignment. This change replaces the previous generic matrix approach with a structure optimized for \Int\ types, resulting in a 30-50% speedup for aligning short sequences. This optimization is part of the broader Needleman-Wunsch alignment implementation that supports affine gap penalties and both global and glocal alignment modes.

src/main/scala/com/fulcrumgenomics/alignment · high confidence

Test coverage

Added test coverage for BAM processing tools; Added test coverage for FASTA and sequence dictionary tools; Added test coverage for UMI and consensus calling tools; Added test coverage for alignment components; Added test data for VCF phasing validation; Added test resources for BAM analysis tools; Added test resources for FASTA assembly reports and soft-masked sequences; Added test resources for Illumina Sample Sheet parsing and RefSeq annotation; Added test resources for UMI annotation; Added tests for BinomialDistribution and FishersExactTest; Added tests for CollectErccMetrics and EstimateRnaSeqInsertSize; Added tests for FgBioDef utilities; Added tests for Illumina RunInfo and SampleSheet parsing; Added tests for ReferenceSetBuilder and VariantContextSetBuilder; Added tests for SAM/BAM I/O, sorting, and record handling; Added tests for basecalling parameter extraction and Illumina run info parsing; Added tests for command-line parsing, global argument handling, and tool metadata; Added tests for utility classes in the com.fulcrumgenomics.util package; Added unit tests for LocatableOrdering; Added unit tests for PileupBuilder and related components; Added unit tests for somatic VCF filtration and read-end artifact filters; Added unit tests for the Scala VCF API core components; New testing utilities for generating reference sets, SAM records, and VCF variants.

Dependencies

Upgrade to HTSJDK 5.0.0 and Scala 2.13.16

The build configuration has been updated to use Scala 2.13.16 as the primary version and upgraded the core HTSJDK library to version 5.0.0. Additionally, the project now depends on commons 1.9.0, sopt 1.2.1, and scala-xml 2.4.0, while removing support for older Scala cross-builds.

(dependencies) · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

Baseline

  • First survey — no prior run to compare against. CAI 61.

Lenses

  • Code Health 91
  • Architecture 96
  • Maturity 53
  • Readiness 55
  • Security 70

Changes since last survey

  • 300 commits — 243 feature/other, 57 fixes

By area

  • src/main — 132 commits
  • (root) — 116 commits
  • .github/workflows — 32 commits
  • src/test — 9 commits
  • .github/CODEOWNERS — 2 commits
  • docs/best-practice-consensus-pipeline.md — 2 commits
  • project/build.properties — 2 commits
  • project/plugins.sbt — 2 commits
  • (repo) — 1 commit
  • .github/logos — 1 commit
  • docs/FastqToConsensus-RnD.smk — 1 commit

Notable commits

  • fix: Alternative bugfix for "ConsensusCallingIterator could fail when no consensus reads are called" (#780)
  • fix: Fix Alignment test (#913)
  • fix: Fix GroupReadsByUmi to treat reads with differing pair-orientations separately (except for strategy=paired) (#1074)
  • fix: Fix a bug where consensus reads are produced with zero depth (#859)
  • fix: Fix a regression where VcfWriter would not write a regular file index (#816)
  • fix: Fix complement of W and S iupac codes. (#912)
  • fix: Fix deprecations and warnings emitted by sbt about it's own build file. (#801)
  • fix: Fix link in DemuxFastqs.scala (#938)
  • fix: Fix most scala 2.13 deprecation warnings. (#806)
  • fix: Fix overly aggressive overlap-clipping regressions in ClipBam (#850)
  • fix: Fix pass-QC in output FASTQ read names (#923)
  • fix: Fix reference to transient MI tag in DuplexConsensusCaller (#946)
  • fix: Fix snapshot releases (#1004)
  • fix: Fixed a couple of bugs in the best practice pipeline doc and added an example snakemake implementation. (#815)
  • fix: [bugfix] SamRecordClipper.clipOverlappingReads now accounts for soft-clipped bases starting before the ends (#842)
  • fix: [bugfix] fix issue #858 in CallDuplexConsensusReads (#864)
  • fix: [bugfix] metrics are now updated and logged in the overlapping bases (#825)
  • fix: [bugfix] overlapping consensus caller should examine only (#824)
  • fix: bugfix: ReviewConsensusVariants should not require grouped raw reads to overlap each variant.
  • fix: bugfix: ZipperBams should fail if there any remaining mapped reads (#929)
  • …and 280 more

Architecture

  • 0 containers · 1 bounded contexts · 0 dependency edges (baseline)

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

fulcrumgenomics/fgbio was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 20 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit e51a661b0e3b6a2ba2f1c98e389450d6b9fa3e97 — the exact code this score is about.
  • Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-b51f968c9b10.