fulcrumgenomics/fgbio
61.0
Adequate · 20 September 2026
28k
lines of production code
Scala
primary language
1
measurement over time
What this system is
This system is a bioinformatics toolkit for processing and analyzing genomic sequencing data, primarily written in Scala. It provides command-line tools and libraries for handling core file formats like SAM/BAM, FASTQ, and VCF, including operations such as alignment, filtering, normalization, and variant calling. The system also supports specialized molecular indexing, UMI consensus generation, and quality control metrics for RNA-seq and duplex sequencing workflows.
How it got here
2015–2016 — Initial release and core tool development
24 changes.
This period established the project's foundational infrastructure, including build workflows, release processes, and comprehensive test coverage. It focused on developing a broad suite of core bioinformatics tools for processing BAM, FASTQ, FASTA, and VCF files, alongside utilities for UMI annotation and molecular index generation.
2017 — New tools and API development
17 changes.
This period focused on expanding the toolkit with new features for Illumina data parsing, somatic variant filtering, and high-precision statistical calculations. Significant infrastructure work included the initial release of a native SAM/BAM/CRAM API and optimizations to the alignment engine. The work was complemented by comprehensive test coverage and the addition of R-based quality control plotting scripts for various sequencing metrics.
2018–2022 — Scala VCF API and Pileup Infrastructure
8 changes.
This period focused on developing core genomic data processing infrastructure, notably introducing a native Scala VCF API to replace Java-centric approaches and implementing flexible random-access and streaming pileup builders for BAM files. The work also included adding coordinate-based ordering for genomic locatables and expanding test coverage for these new components to ensure robust handling of variant data and alignment pileups.
Features
Added DNA/RNA secondary structure prediction utilities
New utility classes have been added to the \com.fulcrumgenomics.util\ package to support DNA and RNA folding analysis. \DnaFoldPrediction\ serves as a data container for sequence, structure, and free energy (deltaG) results, while \DnaFoldPredictor\ provides a Java interface to the ViennaRNA RNAFold executable. The predictor manages a background process to calculate minimum free energy secondary structures for input sequences, allowing users to integrate external RNA folding tools directly into their bioinformatics workflows.
src/main/java · high confidence
Added ERCC QC plotting script and standard mixture data
The \CollectErccMetrics\ tool now includes an R script (\CollectErccMetrics.R\) that generates quality control plots for ERCC spike-in controls. This script takes metrics data and outputs PDF plots showing the correlation between expected and observed counts (both raw and normalized) using log2 scaling and linear regression. It also adds the \ErccStandardMixtures.txt\ resource file, which contains the reference concentrations and expected fold-change ratios for the ERCC Standard Mixtures (Mixes 1-4), enabling accurate comparison of observed sequencing data against known standards.
src/main/resources/com/fulcrumgenomics/rnaseq · high confidence
Added R script for generating QC plots in CollectDuplexSeqMetrics
The \CollectDuplexSeqMetrics\ tool now includes an R script (\CollectDuplexSeqMetrics.R\) that generates a comprehensive PDF report of quality control plots for duplex sequencing data. This script visualizes family size distributions, duplex yield metrics (including actual vs. ideal duplex ratios), UMI representation, and read distribution among tag families, allowing users to visually assess the quality and characteristics of their duplex sequencing runs.
src/main/resources/com/fulcrumgenomics/umi · high confidence
Added coordinate-based ordering for genomic locatable objects
Users can now sort genomic intervals (Locatable instances) consistently using a new ordering mechanism that respects the sequence dictionary. This ordering prioritizes contig index, followed by start position, and then end position, ensuring that variants and intervals are arranged in a biologically meaningful order suitable for downstream processing like pileup building.
src/main/scala/com/fulcrumgenomics/coord · high confidence
Initial project setup and release infrastructure
The repository has been initialized with the core project structure, including the MIT license, contributing guidelines, and a version definition file (version.sbt) set to 4.1.2-SNAPSHOT. This change establishes the build and release workflow, configuring Git LFS for large genomic data files (BAM, VCF), integrating Codecov for test coverage reporting, and setting up git-cliff for automated changelog generation.
(repo-wide) · high confidence
Initial release of the Scala VCF API
Introduces a new Scala-native API for reading and writing VCF and BCF files, replacing the previous Java-centric approach. The new \VcfSource\ and \VcfWriter\ classes provide efficient iteration and writing capabilities, while \ArrayAttr\ and \GenotypeMap\ offer specialized data structures for handling multivalued attributes and genotype samples. Internal conversion logic in \VcfConversions\ bridges the Scala models with HTSJDK, ensuring correct handling of header metadata, counts, and field types.
src/main/scala/com/fulcrumgenomics/vcf/api · high confidence
Initial release of the fgbio SAM/BAM/CRAM API
This change introduces the core \com.fulcrumgenomics.bam.api\ package, providing a Scala-native wrapper around HTSJDK for reading and writing SAM, BAM, and CRAM files. The new API includes \SamSource\ for streaming and querying records, \SamWriter\ for output with optional sorting and indexing, and \SamRecord\ for convenient access to read attributes (such as flags, CIGAR strings, and mate information) without exposing the underlying HTSJDK types. It also adds \HeaderHelper\ for accessing header metadata and \SamIterator\ for efficient iteration, including support for CRAM conversion.
src/main/scala/com/fulcrumgenomics/bam/api · high confidence
Introduce SampleSheet parser with strict uniqueness validation
Adds a new \SampleSheet\ parser for Illumina Experiment Manager sample sheets (typically MiSeq) that enforces strict uniqueness constraints: sample identifiers must be unique, and the combination of sample name and library identifier must also be unique across samples. The parser automatically sets the library ID to the sample ID if the library ID field is missing, handles case-insensitive header lookups, and skips empty lines at the end of the data section. It also supports filtering samples by a specific lane number.
src/main/scala/com/fulcrumgenomics/illumina · high confidence
Introduce core FASTQ I/O tools and utilities
Adds a new set of tools and library components for reading, writing, sorting, trimming, and converting FASTQ files. This includes the FastqSource and FastqWriter classes for efficient FASTQ I/O, a SortFastq tool for lexicographic sorting of FASTQ records, a TrimFastq tool to trim reads to a specified length (with optional exclusion of short reads), and a FastqToBam tool to convert FASTQ files into unmapped BAM/SAM/CRAM formats while supporting read structures for extracting UMIs, cell barcodes, and sample indices.
src/main/scala/com/fulcrumgenomics/fastq · high confidence
Introduction of structured command-line tool grouping
The command-line interface now organizes tools into distinct categories (such as SAM/BAM, FASTA, FASTQ, RNA-Seq, UMIs, and Utilities) via the new \ClpGroups\ object. This change improves the user experience by providing a structured and navigable help menu, allowing users to easily locate tools based on the data type they are working with rather than searching through a flat list of commands.
src/main/scala/com/fulcrumgenomics/cmdline · high confidence
New ExtractIlluminaRunInfo tool to parse Illumina RunInfo.xml
A new command-line tool, ExtractIlluminaRunInfo, has been added to the basecalling module. It reads an Illumina RunInfo.xml file and outputs a metric file containing key run details, including the run barcode, flowcell barcode, instrument name, run date, read structure, and number of lanes.
src/main/scala/com/fulcrumgenomics/basecalling · high confidence
New FASTA manipulation tools: HardMaskFasta, SortSequenceDictionary, and UpdateFastaContigNames
This release introduces three new command-line tools for managing FASTA files and sequence dictionaries. HardMaskFasta converts soft-masked (lowercase) bases to hard-masked 'N' characters and standardizes line lengths for compatibility with tools like samtools faidx. SortSequenceDictionary reorders the contigs in a sequence dictionary file to match the order defined in another reference dictionary, with an option to skip or append missing contigs. UpdateFastaContigNames renames contigs in a FASTA file based on a provided sequence dictionary, supports sorting the output by dictionary order, and can append missing contigs from a default FASTA file.
src/main/scala/com/fulcrumgenomics/fasta · high confidence
New R scripts for ErrorRateByReadPosition and FindSwitchbackReads QC plots
Added R scripts to generate quality control plots for the ErrorRateByReadPosition and FindSwitchbackReads tools. The ErrorRateByReadPosition script now supports plotting individual substitution types (e.g., A\>C, G\>T) when the 'collapsed' column is false, allowing for higher sensitivity in error analysis. The FindSwitchbackReads script generates distribution plots for switchback length, offset, and tandem gap metrics.
src/main/resources/com/fulcrumgenomics/bam · high confidence
New VCF manipulation tools and iterators
This update introduces several new capabilities for working with VCF files. The \MakeTwoSampleMixtureVcf\ tool allows users to create simulated tumor or tumor/normal VCFs by in-silico mixing genotypes from two samples at a specified proportion, including support for tumor-only mode and interval restriction. New iterator utilities, \ByIntervalListVariantContextIterator\ and \JointVariantContextIterator\, enable efficient querying of variants overlapping specific intervals and merging multiple variant context streams, respectively. Additionally, the \UpdateVcfContigNames\ tool provides a way to remap contig names in a VCF based on a provided sequence dictionary, and \VariantMask\ offers a compact representation for rapidly querying whether positions are overlapped by variants.
src/main/scala/com/fulcrumgenomics/vcf · high confidence
New high-precision binomial and Fisher's exact test math utilities
Added BinomialDistribution and FishersExactTest classes to the math library. BinomialDistribution provides high-precision binomial probability calculations using unlimited precision BigDecimal to avoid underflow issues found in standard libraries, while FishersExactTest implements Fisher's Exact Test for 2x2 contingency tables, supporting two-sided, greater, and less alternative hypotheses.
src/main/scala/com/fulcrumgenomics/math · high confidence
New internal tool for generating per-tool Markdown documentation
An internal command-line tool (\BuildToolDocs\) has been added to automatically generate Markdown documentation for all available tools. This tool scans specified packages (defaulting to \com.fulcrumgenomics\), groups tools by category, and produces an index file along with individual documentation pages for each tool, including details on arguments, types, and default values. The output can optionally exclude the version string and is written to a specified output directory.
src/main/scala/com/fulcrumgenomics/internal · high confidence
New personal tools for region generation and tag splitting
Added two new personal tools: \GenerateRegionsFromFasta\, which creates freebayes/bamtools-style region specifiers from a FASTA file by dividing sequences into configurable chunk sizes, and \SplitTag\, which splits a single optional SAM/BAM tag into multiple distinct tags based on a specified delimiter. Both tools are registered under the Personal command-line group.
src/main/scala/com/fulcrumgenomics/personal/nhomer · high confidence
New personal tools for stripping Fastq read numbers and summarizing GFF biotypes
Added two new command-line tools in the personal tools section: StripFastqReadNumbers, which removes trailing slash-number suffixes (e.g., /1, /2) from FASTQ read names by clearing the read number field, and SummarizeGff, which processes RefSeq GFF files to produce summary statistics of single-exon versus multi-exon gene counts grouped by gene biotype.
src/main/scala/com/fulcrumgenomics/personal/tfenne · high confidence
New random-access and streaming pileup builders for BAM processing
Added two new pileup builder implementations in the BAM pileup module: RandomAccessPileupBuilder, which uses index-based random access on indexed BAM files, and StreamingPileupBuilder, which uses coordinate-sorted lazy streaming with an internal cache. Both builders support configurable filtering options (minimum mapping/base quality, duplicate/secondary/supplementary alignment inclusion, mapped pairs only) and provide a consistent API for generating pileups at specific genomic loci.
src/main/scala/com/fulcrumgenomics/bam/pileup · high confidence
New somatic variant filtering tool with A-tailing and End Repair artifact detection
The \FilterSomaticVcf\ tool is introduced to apply statistical filters to somatic variant calls by analyzing read pileups. It now supports two distinct artifact filters: the A-tailing Artifact Filter (previously 'End Repair Artifact Filter') and the new End Repair Fill-in Artifact Filter. These filters evaluate whether mismatches at read ends are likely library preparation artifacts rather than true mutations, adding corresponding p-value annotations (ATAP, ERFAP) and FILTER tags to the output VCF. The tool also introduces configurable BAM access patterns (RandomAccess vs. Streaming) to optimize performance based on VCF density and sorting.
src/main/scala/com/fulcrumgenomics/vcf/filtration · high confidence
New tools for BAM normalization, filtering, and technical read detection
This release introduces several new command-line tools for processing BAM files. DownsampleAndNormalizeBam downsamples BAMs to a uniform coverage across specified regions, ensuring every base meets a target depth. FilterBam now supports writing rejected reads to a separate file via the --rejects option, allowing users to inspect filtered-out data. FindTechnicalReads identifies and extracts reads derived from technical sequences (like adapter dimers), with options to tag matches or output all reads. Additional new tools include RandomizeBam for shuffling read order, RemoveSamTags for stripping specific SAM tags, SetMateInformation for fixing paired-end metadata, SortBam for various sort orders including TemplateCoordinate, TrimPrimers for post-alignment primer trimming with optional tag recalculation, and UpdateReadGroups for mapping old read group IDs to new ones.
src/main/scala/com/fulcrumgenomics/bam · high confidence
New tools for CODEC duplex consensus and UMI annotation
This release introduces \CallCodecConsensusReads\, a new tool for generating duplex consensus sequences from CODEC sequencing data, which assembles single-strand consensuses from R1 and R2 reads into a final duplex consensus. It also adds \AnnotateBamWithUmis\, a new tool that annotates existing BAM files with UMIs extracted from separate FASTQ files, supporting multiple FASTQ inputs and optional UMI quality storage. These additions expand the UMI toolkit to support specific sequencing protocols and pre-annotated data workflows.
src/main/scala/com/fulcrumgenomics/umi · high confidence
New tools for picking molecular indices and generating read groups
Added the \PickLongIndices\ tool for efficiently generating large sets of molecular indices with customizable constraints (length, edit distance, GC content, homopolymers, and secondary structure via ViennaRNA), and the \PickIlluminaIndicesCommand\ utility for selecting optimal index sets. Also added \AutoGenerateReadGroupsByName\, which parses Illumina-style read names to automatically create and assign read groups to BAM files.
repository · high confidence
New utility classes for amplicon detection, genomic spans, and interval list I/O
This change introduces several new utility components in the \com.fulcrumgenomics.util\ package. The \AmpliconDetector\ class allows users to identify which amplicon a sequencing read (or read pair) originated from by matching read coordinates against a set of known amplicons, with support for adjusting for clipped bases. New \GenomicSpan\ and \IntervalListSource\/\IntervalListWriter\ classes provide standardized ways to represent genomic intervals and stream-read/write interval list files, including proper handling of sequence dictionaries. Additionally, \Io.scala\ is updated to automatically handle \.bgz\ and \.bgzip\ compression when writing files, and \Metric.scala\ gains support for collection fields and improved serialization of enum entries.
src/main/scala/com/fulcrumgenomics/util · high confidence
Behavioural changes
Optimized alignment performance via specialized integer matrix
The alignment engine now uses a specialized \SimpleIntMatrix\ implementation for storing integer values during sequence alignment. This change replaces the previous generic matrix approach with a structure optimized for \Int\ types, resulting in a 30-50% speedup for aligning short sequences. This optimization is part of the broader Needleman-Wunsch alignment implementation that supports affine gap penalties and both global and glocal alignment modes.
src/main/scala/com/fulcrumgenomics/alignment · high confidence
Test coverage
Added test coverage for BAM processing tools; Added test coverage for FASTA and sequence dictionary tools; Added test coverage for UMI and consensus calling tools; Added test coverage for alignment components; Added test data for VCF phasing validation; Added test resources for BAM analysis tools; Added test resources for FASTA assembly reports and soft-masked sequences; Added test resources for Illumina Sample Sheet parsing and RefSeq annotation; Added test resources for UMI annotation; Added tests for BinomialDistribution and FishersExactTest; Added tests for CollectErccMetrics and EstimateRnaSeqInsertSize; Added tests for FgBioDef utilities; Added tests for Illumina RunInfo and SampleSheet parsing; Added tests for ReferenceSetBuilder and VariantContextSetBuilder; Added tests for SAM/BAM I/O, sorting, and record handling; Added tests for basecalling parameter extraction and Illumina run info parsing; Added tests for command-line parsing, global argument handling, and tool metadata; Added tests for utility classes in the com.fulcrumgenomics.util package; Added unit tests for LocatableOrdering; Added unit tests for PileupBuilder and related components; Added unit tests for somatic VCF filtration and read-end artifact filters; Added unit tests for the Scala VCF API core components; New testing utilities for generating reference sets, SAM records, and VCF variants.
Dependencies
Upgrade to HTSJDK 5.0.0 and Scala 2.13.16
The build configuration has been updated to use Scala 2.13.16 as the primary version and upgraded the core HTSJDK library to version 5.0.0. Additionally, the project now depends on commons 1.9.0, sopt 1.2.1, and scala-xml 2.4.0, while removing support for older Scala cross-builds.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Baseline
- First survey — no prior run to compare against. CAI 61.
Lenses
- Code Health 91
- Architecture 96
- Maturity 53
- Readiness 55
- Security 70
Changes since last survey
- 300 commits — 243 feature/other, 57 fixes
By area
- src/main — 132 commits
- (root) — 116 commits
- .github/workflows — 32 commits
- src/test — 9 commits
- .github/CODEOWNERS — 2 commits
- docs/best-practice-consensus-pipeline.md — 2 commits
- project/build.properties — 2 commits
- project/plugins.sbt — 2 commits
- (repo) — 1 commit
- .github/logos — 1 commit
- docs/FastqToConsensus-RnD.smk — 1 commit
Notable commits
- fix: Alternative bugfix for "ConsensusCallingIterator could fail when no consensus reads are called" (#780)
- fix: Fix Alignment test (#913)
- fix: Fix GroupReadsByUmi to treat reads with differing pair-orientations separately (except for strategy=paired) (#1074)
- fix: Fix a bug where consensus reads are produced with zero depth (#859)
- fix: Fix a regression where VcfWriter would not write a regular file index (#816)
- fix: Fix complement of W and S iupac codes. (#912)
- fix: Fix deprecations and warnings emitted by sbt about it's own build file. (#801)
- fix: Fix link in DemuxFastqs.scala (#938)
- fix: Fix most scala 2.13 deprecation warnings. (#806)
- fix: Fix overly aggressive overlap-clipping regressions in ClipBam (#850)
- fix: Fix pass-QC in output FASTQ read names (#923)
- fix: Fix reference to transient MI tag in DuplexConsensusCaller (#946)
- fix: Fix snapshot releases (#1004)
- fix: Fixed a couple of bugs in the best practice pipeline doc and added an example snakemake implementation. (#815)
- fix: [bugfix] SamRecordClipper.clipOverlappingReads now accounts for soft-clipped bases starting before the ends (#842)
- fix: [bugfix] fix issue #858 in CallDuplexConsensusReads (#864)
- fix: [bugfix] metrics are now updated and logged in the overlapping bases (#825)
- fix: [bugfix] overlapping consensus caller should examine only (#824)
- fix: bugfix: ReviewConsensusVariants should not require grouped raw reads to overlap each variant.
- fix: bugfix: ZipperBams should fail if there any remaining mapped reads (#929)
- …and 280 more
Architecture
- 0 containers · 1 bounded contexts · 0 dependency edges (baseline)
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
fulcrumgenomics/fgbio was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 20 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit e51a661b0e3b6a2ba2f1c98e389450d6b9fa3e97 — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-b51f968c9b10.