Skip to content
CAI
Software that uses CAICheck a score

broadinstitute/cromwell

61.0

Adequate · 28 September 2026

96.3k

lines of production code

Scala

primary language

2

measurements over time

CAI band scale
CAI trend line
CAI lens gauges

What this system is

This system is a workflow engine and associated infrastructure designed to execute, manage, and monitor computational workflows defined in the Workflow Description Language (WDL). It provides a robust execution engine that handles job scheduling, backend abstraction for diverse cloud and local environments, and comprehensive state management including call caching and workflow restart capabilities. The system also includes a REST API for workflow submission and monitoring, a standardized testing framework for validation, and utilities for managing cloud storage access and Docker image resolution.

How it got here

2015–2016 — Engine architecture and backend standardization

54 changes.

This period focused on rebuilding the Cromwell workflow engine with a modular actor-based architecture, introducing standardized backend interfaces and pluggable filesystem abstractions. It established core infrastructure for workflow execution, job storage, and call caching, while migrating the web service to Akka HTTP and restructuring the database access layer for improved modularity and reliability.

2017 — Engine architecture and I/O refactoring

70 changes.

This period focused on a major refactoring of the workflow engine and I/O subsystem, introducing asynchronous Akka-based actors for file operations with robust backpressure and retry logic. It also established new core abstractions for labels, execution keys, and state management, while adding support for the TES backend and WDL 1.1 features. Comprehensive testing infrastructure improvements were made to the Centaur suite, alongside the release of a dedicated API client library.

2018 — WDL version expansion and cloud infrastructure

60 changes.

The project significantly expanded its WDL language support by implementing parsers and execution models for Draft 2, Draft 3, and the Biscayne (1.1/2.0) specifications. Concurrently, it established a robust AWS backend by introducing S3 filesystem integration, AWS Batch job execution, and configurable authentication, while also launching the CromIAM service for centralized API management and authorization.

2019–2026 — backend modernization and observability

29 changes.

This period focused on modernizing backend integrations, notably migrating the S3 provider to AWS SDK v2 and restructuring the GCP Batch backend with a new request manager architecture. Significant effort was also directed toward enhanced observability and performance analysis, introducing tools for metadata comparison, I/O backpressure reporting, and real-time resource monitoring. These changes were supported by expanded test coverage and new utilities for workflow execution, such as the standalone DRS Localizer and automated Java client generation.

Features

Add Centaur configuration templates for JES and local Cromwell backends

New configuration templates have been added to the Centaur test suite to support running Cromwell against the JES (Google Life Sciences) and local backends. The \jes-compose.yml.ctmpl\ and \local-compose.yml.ctmpl\ files define Docker Compose services for MySQL and Cromwell, including specific JVM options and backend settings (e.g., \backend.defaultBackend=JES\ vs \local\). Additionally, \cromwell-account.pem.ctmpl\ provides a template for injecting the service account private key via environment variables, enabling authenticated access to Google Cloud services during testing.

centaur/src/main/config · high confidence

Add Docker registry flows for AWS ECR, Docker Hub, Google, and Quay

New registry flow implementations have been added to handle Docker image resolution for AWS ECR (including public ECR), Docker Hub, Google Container/Artifact Registry, and Quay. These changes enable Cromwell to authenticate and fetch manifests from these specific registries using their respective protocols and SDKs.

dockerHashing/src/main/scala/cromwell/docker/registryv2/flows · high confidence

Add FTP filesystem support to Cloud NIO

Introduces a new FTP implementation for the Cloud NIO abstraction, enabling users to read, write, list, and manage files on FTP servers. This change adds the \FtpCloudNioFileSystemProvider\ and supporting classes to handle connection pooling, authentication (including anonymous and account-based login), and file operations like copying and directory creation. It also includes configuration options for connection modes (active/passive) and caching strategies to manage resources efficiently.

cloud-nio/cloud-nio-impl-ftp, filesystems/ftp · high confidence

Add GA4GH Workflow Execution Service (WES) API endpoints

Introduces a new set of REST endpoints under the \/ga4gh/wes/v1\ path to expose Cromwell's workflow capabilities via the GA4GH WES standard. This includes a \service-info\ endpoint for discovering supported workflow types, filesystems, and engine versions; a \runs\ endpoint for submitting new workflows (both single and batch) and listing existing runs; a \runs/{id}\ endpoint to retrieve detailed run logs and metadata; a \runs/{id}/status\ endpoint for checking workflow state; and a \runs/{id}/cancel\ endpoint for aborting workflows. The implementation includes necessary data models for WES responses, state mapping from Cromwell's internal states to WES states, and JSON serialization support.

engine/src/main/scala/cromwell/webservice/routes/wes · high confidence

Add Google Cloud PAPIv2 resource monitoring and file transfer utilities

This change introduces a new Docker-based Stackdriver task monitor for Google PAPIv2 workflows, allowing users to track real-time CPU, memory, disk utilization, and IOPS metrics via the \monitoring\_image\ workflow option. It also adds a \gcs\_transfer.sh\ script to handle file localization and delocalization with support for requester-pays buckets and parallel composite uploads.

supportedBackends/google/pipelines/common · high confidence

Add Simpletons conversion utility for database entries

A new \Simpletons\ object has been added to the \cromwell\ package to handle the conversion of database simpleton entries (\CallCachingSimpletonEntry\ and \JobStoreSimpletonEntry\) into \WomValueSimpleton\ instances. This utility explicitly maps WDL/WOM types (such as String, Int, Float, Long, Boolean, File, and Directory) to their corresponding WomValue implementations, ensuring that File and Directory types are correctly instantiated as \WomSingleFile\ and \WomUnlistedDirectory\ respectively during the conversion process.

engine/src/main/scala/cromwell · high confidence

Add WDL 1.0 language support via Draft 3 factory

Introduces the WdlDraft3LanguageFactory, enabling the Cromwell engine to parse, validate, and execute workflows written in WDL version 1.0. This new factory handles the full lifecycle of WDL 1.0 files, including import resolution, AST parsing, and conversion to the WomExecutable, while leveraging the existing parser cache for performance.

languageFactories/wdl-draft3/src/main · high confidence

Add WDL 1.1 language factory

The WDL 1.1 language factory (WdlBiscayneLanguageFactory) is now available, enabling the system to parse and execute workflows declaring 'version 1.1'. This factory handles validation, bundle creation, and executable generation for this specific WDL version, distinct from other supported versions like 'Cascades' or 'development'.

languageFactories/wdl-biscayne · high confidence

Add WDL transformation utilities for database migration

A new \WdlTransformation\ object has been added to the database migration module to handle WDL-specific data transformations. This includes an \inflate\ method that decompresses and decodes Base64-encoded GZIP strings (falling back to the raw value if decompression fails) and a \coerceStringToWdl\ method that converts string representations into appropriate WDL/WOM types based on the target type.

database/migration/src/main/scala/cromwell/database/migration · high confidence

Add support for NIH Sequence Read Archive (SRA) paths

Users can now reference files stored in the NIH Sequence Read Archive using the \sra://\ scheme (e.g., \sra://\<accession\>/\<path\>\). This change introduces the \SraPathBuilder\ and \SraPathBuilderFactory\ components, which parse and validate these URIs to extract the accession ID and internal file path, enabling workflows to interact with SRA data sources directly.

filesystems/sra · high confidence

Added Cromwell server management script

A new shell script, \scripts/server/cromwell.server\, has been added to provide a convenient interface for managing the Cromwell server lifecycle. Users can now easily start, stop, check the status of, view logs for, or edit the configuration of the Cromwell server using simple command-line arguments (e.g., \cromwell.server start\).

scripts/server · high confidence

Added JSON serialization support for WomValue types

A new WomValueJsonFormatter has been introduced to enable the serialization and deserialization of WDL workflow values (such as strings, integers, files, arrays, maps, and objects) to and from JSON. This allows WDL values to be easily converted for storage or transmission, with specific handling for optional values and nested structures.

core/src/main/scala/cromwell/util/JsonFormatting · high confidence

Added WDL workflow definitions for process execution

New WDL (Workflow Description Language) files, specifically 3step.wdl and sample.wdl, have been added to the engine resources. These files define reusable workflow components: 3step.wdl outlines a three-step process involving process listing, pattern matching, and word counting, while sample.wdl provides a scatter-gather pattern for grepping and counting lines across multiple input files. These additions enable the engine to execute these specific workflow definitions.

engine/src/main/resources · high confidence

Added script to generate reference disk manifests

A new shell script, create\_images.sh, has been added to the scripts/reference\_disks directory. This tool automates the creation of HOCON manifest files for reference images by reading an input TSV, verifying existing resources (images, disks, mounts), and computing CRC32C hashes for cloud objects to ensure data integrity in the generated manifests.

_scripts/reference\disks · high confidence

Automated Java client generation and publishing workflow

This change introduces a new set of scripts to automate the build and distribution of the Cromwell Java client. The \gen\_java\_client.sh\ script uses the OpenAPI Generator to create the client code from the Swagger definition, specifically patching the API spec to include OAuth2 security definitions and replacing 'file' types with 'string' for better usability. It also configures the generated client to include a custom User-Agent header containing the current Cromwell commit hash. Additionally, \publish-client.sh\ handles publishing the generated library to Artifactory, distinguishing between snapshot and release builds based on the CI branch, while \make\_pr\_clone\_branch.sh\ provides a utility for creating isolated branches to test pull request changes in Travis CI.

scripts · high confidence

Centaur test runner configuration and AWS credential support

The Centaur test runner now includes explicit configuration files for managing test behavior and cloud credentials. A new \application.conf\ sets Akka HTTP connection pool limits (max 1024 open requests, 20 max connections, 10 min connections) to optimize network performance during testing. A new \reference.conf\ defines core Centaur settings, including Cromwell interaction modes (URL or JAR), timeouts for workflow progress and metadata consistency, and Google Cloud configuration defaults such as authentication schemes and endpoint URLs. Additionally, a new \centaur\_aws\_credentials.conf\ file allows AWS access keys, secret keys, and region to be configured via environment variables, enabling the test suite to run against AWS backends.

centaur/src/main/resources · high confidence

Configurable SFS backend for custom job submission scripts

The Shared File System (SFS) backend now supports a new 'config' mode that allows users to define custom job submission, kill, and health-check commands via WDL-like task definitions in the backend configuration. This enables flexible integration with various job schedulers or local execution environments without writing new backend code. The implementation includes a new actor hierarchy (ConfigAsyncJobExecutionActor, ConfigInitializationActor) and a pluggable file hashing strategy (ConfigHashingStrategy) that supports multiple algorithms (MD5, xxh64, fingerprint) and strategies (path, file content, modification time) for call caching.

supportedBackends/sfs · high confidence

Core infrastructure and configuration models

This change introduces a suite of foundational classes and utilities within the core module to support workflow execution and configuration management. It adds models for workflow identification (WorkflowId, RootWorkflowId), execution tracking (ExecutionStatus, ExecutionIndex, JobKey, CallKey), and source file handling (WorkflowSourceFilesCollection). It also introduces configuration parsers for Docker settings (BackendDockerConfiguration, DockerConfiguration), load control thresholds (LoadConfig), and workflow options with support for encrypted fields (WorkflowOptions). Additionally, it provides utility classes for configuration validation (ConfigUtil), encryption (Encryption), and actor monitoring (MonitoringCompanionActor).

core/src/main/scala/cromwell/core · high confidence

Draft 3 WDL model elements and graph linking infrastructure

The draft 3 WDL model now includes a comprehensive set of AST element case classes (Call, Scatter, If, Declaration, Task, Workflow, Struct, Expression, and Section elements) and a graph-linking layer (LinkedGraph, GeneratedValueHandle, UnlinkedConsumedValueHook) that connects workflow nodes and resolves value dependencies. This enables Cromwell to parse, type-check, and evaluate Draft 3 WDL workflows by building a linked execution graph from the parsed model.

wdl/model/draft3 · high confidence

Enable local Docker image hashing via CLI

Cromwell now supports looking up and pulling Docker images using the local Docker CLI instead of relying solely on the Docker REST API. This change introduces a new \DockerCliFlow\ and \DockerCliClient\ in the \dockerHashing\ module that execute \docker images\ and \docker pull\ commands to resolve image digests. The implementation includes a 5-second timeout for the initial lookup to mitigate known Docker CLI hanging issues and automatically pulls an image if it is not found locally before attempting to resolve its hash. This provides a fallback mechanism for environments where the Docker API might be unavailable or unreliable.

dockerHashing/src/main/scala/cromwell/docker/local · high confidence

Engine-level instrumentation for HTTP, I/O, jobs, and workflows

The engine now includes dedicated instrumentation modules (HttpInstrumentation, IoInstrumentation, JobInstrumentation, WorkflowInstrumentation) that expose detailed metrics to StatsD. HTTP requests are tracked by normalized path, method, and status code; I/O operations are measured by filesystem type (local vs. GCS) and outcome (success, failure, retry); job execution times are recorded per state (succeeded, aborted, failed); and workflow lifecycle metrics (queued, running, on-hold, aborting) are exposed as gauges and counters, enabling users to monitor engine performance and resource utilization.

engine/src/main/scala/cromwell/engine/instrumentation · high confidence

GCS filesystem introduces batched I/O command execution

The GCS filesystem implementation now uses a new batched I/O architecture for operations such as size, delete, copy, hash, touch, exists, and isDirectory. This change introduces \GcsBatchCommandBuilder\ and a suite of \GcsBatchIoCommand\ classes (e.g., \GcsBatchCopyCommand\, \GcsBatchSizeCommand\) that wrap Google Cloud Storage API calls into batchable units. This enables more efficient handling of GCS requests by allowing them to be grouped and executed in batches, improving performance for workflows involving many file operations.

filesystems/gcs/src/main/scala/cromwell/filesystems/gcs/batch · high confidence

Initial Docker entrypoint and build scripts

Added \docker/install.sh\ to build the Cromwell JAR using SBT and \docker/run.sh\ to start the server. The run script now executes the JAR using the \JAVA\_OPTS\ environment variable, allowing users to configure JVM options externally rather than relying on hard-coded settings.

docker · high confidence

Initial Draft 3 WDL parser and transformation pipeline

This change introduces the core infrastructure for parsing and transforming WDL Draft 3 workflows. It adds a generated Java parser (WdlParser) and Scala wrappers that convert raw WDL source code into an Abstract Syntax Tree (AST). The diff also implements the transformation chain from this AST to the WDL Object Model (WDLOM), and subsequently links and evaluates expressions (handling types, file dependencies, and values) before converting the WDLOM into the Workflow Object Model (WOM) used for execution. This establishes the foundational parsing and linking logic for Draft 3 workflows.

wdl/transforms/draft3 · high confidence

Initial S3 storage client implementation

Added S3Storage.scala, which provides factory methods to create AWS S3 clients. The new code supports configuring S3 acceleration, dual-stack networking, and path-style access via configuration, and allows specifying an optional AWS region and credentials provider for client instantiation.

cloudSupport/src/main/scala/cromwell/cloudsupport/aws/s3 · high confidence

Initial build configuration for Java code generator

The Java code generator module now includes its own build infrastructure, introducing sbt 1.5.8 and configuration files for versioning, publishing, and Artifactory integration. This enables the module to be built independently and published to the Broad Institute's Artifactory repositories, with version strings automatically appended with the git commit hash.

_codegen\java · high confidence

Initial implementation of the AWS Batch backend

Cromwell now supports executing workflows on AWS Batch. This change introduces the \AwsBatchAsyncBackendJobExecutionActor\ and its supporting lifecycle actors, configuration, and job definition logic, enabling users to submit and manage jobs via the AWS Batch service. The implementation includes support for S3 and local filesystems, EFS/FSx mounting, Docker mirroring, and private Docker token storage via AWS Secrets Manager.

supportedBackends/aws · high confidence

Initial implementation of the S3 filesystem backend

Adds the core components for the new AWS S3 backend, including \S3PathBuilder\ and \S3PathBuilderFactory\. This introduces the ability to parse, validate, and construct S3 URIs (e.g., \s3://bucket/key\) and integrates with the AWS SDK to handle authentication modes and region configuration, enabling Cromwell to read from and write to S3 storage locations.

filesystems/s3/src/main/scala/cromwell/filesystems/s3 · high confidence

Initial release of the Cromwell API client

Introduces the CromwellClient, a new Scala library for interacting with the Cromwell workflow engine. This client provides methods to submit workflows (including batch submissions), query workflow status, retrieve metadata, outputs, logs, and costs, manage workflow labels, and check the engine version. It supports HTTP authentication via optional credentials and handles JSON marshalling/unmarshalling for Cromwell API responses.

cromwellApiClient/src/main/scala/cromwell/api · high confidence

Initial support for DRS (DOS) filesystem: Size and File Hash

This change introduces the initial implementation of the DRS (Data Object Resolution Service) filesystem provider within Cromwell. It adds the core path handling classes (\DrsPath\, \DrsPathBuilder\) and configuration factories (\DrsPathBuilderFactory\) required to resolve DRS URIs. Specifically, it enables retrieving file hashes and sizes for DRS-backed inputs, supports pre-resolving DRS URIs to underlying GCS paths when possible, and handles reading data via access URLs or direct GCS access with requester-pays project support.

filesystems/drs/src/main · high confidence

Initial support for DRS (Data Repository Service) filesystem operations

This change introduces the core implementation for the DRS filesystem provider within the cloud-nio library, enabling Cromwell to interact with Data Repository Service endpoints. It adds the \DrsCloudNioFileSystemProvider\ and \DrsCloudNioFileProvider\ classes which handle path resolution, file existence checks, and reading file contents via HTTP requests to a DRS resolver. The implementation includes configuration management for resolver URLs, timeouts, and retry strategies (exponential backoff for 429/5xx errors), as well as credential strategies for obtaining access tokens (Google OAuth and Application Default Credentials). It also supports retrieving file attributes such as size, hashes, and timestamps, and handles non-standard DRS URIs (including Compact Identifier-based URIs) by treating them as opaque strings.

cloud-nio/cloud-nio-impl-drs/src/main · high confidence

Introduce Biscayne WDL parser and engine functions

This change introduces the Biscayne WDL parser (generated via Hermes) and its associated AST-to-model transformation pipeline, enabling support for WDL 2.0/development features. Users can now use new engine functions including \keys\, \as\_map\, \as\_pairs\, \collect\_by\_key\, \min\, \max\, \sep\, \sub\ (POSIX-flavored), \suffix\, \quote\, \squote\, and \unzip\. The parser also supports struct literals, task input passthrough syntax (e.g., \foo\ instead of \foo=foo\), and the \None\ literal. Additionally, the legacy \read\_object\, \read\_objects\, \write\_object\, and \write\_objects\ functions are removed in favor of \read\_json\ and \write\_json\.

wdl/transforms/biscayne · high confidence

Introduce Cloud NIO SPI abstraction for cloud storage access

Adds a new SPI module (\cloud-nio-spi\) that provides a Java NIO.2-compatible file system provider for cloud storage. This introduces core abstractions including \CloudNioFileSystemProvider\, \CloudNioPath\, and \CloudNioFileProvider\, along with supporting classes for retry logic (\CloudNioRetry\, \CloudNioBackoff\), file attributes (\CloudNioFileAttributes\), and I/O channels (\CloudNioReadChannel\, \CloudNioWriteChannel\). This allows applications to interact with cloud storage using standard \java.nio.file\ APIs, with built-in support for exponential backoff and transient error handling.

cloud-nio/cloud-nio-spi · high confidence

Introduce CromIAM health monitoring subsystem

Added a new health monitoring system for CromIAM that periodically checks the status of dependent services (Cromwell and Sam). The implementation includes a \StatusService\ which configures and schedules a \HealthMonitor\ actor to poll these subsystems at a configurable interval, and a \StatusCheckedSubsystem\ trait that provides a reusable method for performing HTTP status checks against specific endpoints. This allows users to query the overall health status of the CromIAM infrastructure via a dedicated status endpoint.

CromIAM/src/main/scala/cromiam/server/status · high confidence

Introduce SubWorkflowStore for tracking nested workflow executions

Added a new SubWorkflowStore component that enables the system to track and query relationships between parent workflows and their nested sub-workflows. This includes a database-backed implementation (SqlSubWorkflowStore) for persisting these relationships and an actor interface (SubWorkflowStoreActor) to handle registration, querying, and cleanup of sub-workflow entries, along with an empty implementation for scenarios where sub-workflow tracking is not required.

engine/src/main/scala/cromwell/subworkflowstore · high confidence

Introduce WDL Draft 2 model and parser implementation

Added the core model and parser components for WDL Draft 2, including the generated \WdlParser\ (Java) and a suite of Scala model classes (\AstTools\, \Declaration\, \Scope\, \WdlCall\, \WdlExpression\, etc.) that define the abstract syntax tree, type system, and execution graph for Draft 2 workflows.

wdl/model/draft2 · high confidence

Introduce WomValueSimpleton for flattening and rebuilding WDL values

Added WomValueSimpleton and WomValueBuilder to enable the serialization of complex WDL values (arrays, maps, objects, pairs) into a flat list of simpleton components for storage or transmission, and their subsequent reconstruction back into full WomValue objects. This mechanism supports features like call caching by allowing efficient comparison and storage of workflow outputs.

core/src/main/scala/cromwell/core/simpleton · high confidence

Introduce configurable AWS backend with flexible authentication modes

Users can now configure the AWS backend via a new \AwsConfiguration\ component that supports multiple authentication schemes: default, custom keys, and assume role. The configuration allows specifying an application name, an optional AWS region, and a list of named auth blocks, enabling flexible credential management for different AWS services or accounts.

cloudSupport/src/main/scala/cromwell/cloudsupport/aws · high confidence

Introduce configurable language factory registry and import resolution infrastructure

This change introduces the core infrastructure for registering and resolving WDL language versions via configuration, replacing hardcoded logic with a dynamic registry in CromwellLanguages. It adds a robust import resolution system (ImportResolver) supporting local filesystem, HTTP, and zipped imports, including host allowlisting and authentication providers. Additionally, it implements a parser cache (ParserCache) to optimize repeated workflow parsing by hashing source, URL, root, and import resolvers, and provides utility functions for validating WDL namespaces and resolving workflow sources.

languageFactories/language-factory-core/src/main/scala/cromwell/languages/util · high confidence

Introduce dedicated Docker hashing subsystem with GCR mirroring support

This change introduces a new \dockerHashing\ module containing the core infrastructure for resolving Docker image hashes, including actors for managing registry lookups, caching, and request throttling. It adds support for Docker image mirroring, specifically enabling GCR (Google Container Registry) mirroring of Docker Hub images via a configurable pull-through cache, and updates the image identifier parser to correctly handle tag and digest combinations (e.g., \image:tag@sha256:...\).

dockerHashing/src/main/scala/cromwell/docker · high confidence

Introduce new TES backend implementation

Adds a new Task Execution Service (TES) backend to Cromwell, enabling users to execute workflows on Azure TES infrastructure. This change introduces the core backend components, including the \TesAsyncBackendJobExecutionActor\ for job execution and status polling, \TesConfiguration\ for managing backend settings (such as endpoint URLs, backoff strategies, and bearer tokens), and \TesTask\ for constructing TES payloads. It also includes support for workflow execution and data access identities, runtime attribute handling (CPU, memory, disk, preemptible), and output modes (granular vs. root).

supportedBackends/tes · high confidence

Introduce new backend actor lifecycle and execution interfaces

Cromwell introduces a new set of core backend actor traits and data structures to standardize job execution and lifecycle management. This includes \BackendJobExecutionActor\ for handling job execution, recovery, and abortion, and \BackendWorkflowInitializationActor\ and \BackendWorkflowFinalizationActor\ for managing workflow-level setup and teardown. The \BackendLifecycleActorFactory\ now provides a unified interface for backends to define their initialization, execution, and finalization actors, along with configuration for job token limits and call caching. New data classes like \BackendJobDescriptor\, \BackendWorkflowDescriptor\, and \BackendInitializationData\ facilitate the passing of context and configuration between the engine and backend actors. Additionally, \OutputEvaluator\ and \RuntimeAttributeDefinition\ provide standardized mechanisms for evaluating workflow outputs and validating runtime attributes.

backend/src/main/scala/cromwell/backend · high confidence

Introduce parallel DRS downloads via Getm and retryable GCS downloads

The DRS localizer now supports downloading multiple files in parallel using the external 'getm' tool, which accepts a JSON manifest of URLs and checksums to accelerate bulk data localization. For Google Cloud Storage (GCS) downloads, the system now implements a retry mechanism with exponential backoff and automatically handles 'Requester Pays' buckets by retrying with the appropriate billing project flag if the initial attempt fails. Additionally, checksum validation is handled more robustly, including the conversion of CRC32C hashes to the base64 format required by getm.

cromwell-drs-localizer/src/main/scala/drs/localizer/downloaders · high confidence

Introduce standalone DRS Localizer CLI with manifest support

The DRS Localizer is now available as a standalone executable JAR that accepts command-line arguments for localization. Users can localize a single file by providing a DRS object ID and a container path, or batch-localize multiple files by supplying a CSV manifest file via the -m flag. The tool currently supports the Google access token strategy and includes a requester-pays project option for Google Cloud Storage downloads.

cromwell-drs-localizer/src/main/scala/drs/localizer · high confidence

Introduces batched S3 I/O command implementation

Adds new Scala classes in the S3 filesystem module to support batched I/O operations. The \S3BatchCommandBuilder\ constructs specific command objects for actions like size checks, deletions, copies, hashing, and existence checks, while \S3BatchIoCommand\ defines the command structures that map AWS SDK responses (such as \HeadObjectResponse\ for size/hash/exists and \CopyObjectResponse\ for copies) to Cromwell's internal I/O model. This enables more efficient, batched communication with S3 for these specific operations.

filesystems/s3/src/main/scala/cromwell/filesystems/s3/batch · high confidence

Introduction of EngineJobExecutionActor for workflow job lifecycle management

The workflow engine now utilizes a new EngineJobExecutionActor to manage the execution state of individual jobs. This actor implements a finite state machine that handles job initialization, token acquisition (both execution and restart tokens), interaction with the job store for restart checks, and coordination with call caching actors. It serves as the central coordinator for job execution events, ensuring proper state transitions and resource management during workflow runs.

engine/src/main/scala/cromwell/engine/workflow/lifecycle/execution/job · high confidence

Introduction of GCS storage configuration and client builder utilities

A new GcsStorage object has been added to the cloud support library to centralize Google Cloud Storage client initialization. This component provides factory methods for constructing Storage instances with configurable retry settings and credentials, and defines a default CloudStorageConfiguration that allows empty path components, strips prefix slashes, uses pseudo-directories, and respects a configurable upload buffer size.

cloudSupport/src/main/scala/cromwell/cloudsupport/gcp/gcs · high confidence

Introduction of MaterializeWorkflowDescriptorActor for workflow descriptor materialization

A new MaterializeWorkflowDescriptorActor has been added to the engine to handle the materialization of workflow descriptors. This actor manages the state machine for processing workflow source files, validating call caching modes and memory retry multipliers, and coordinating with backend services and import resolvers to produce a validated EngineWorkflowDescriptor or return specific failure responses.

engine/src/main/scala/cromwell/engine/workflow/lifecycle/materialization · high confidence

Introduction of WDL Draft 2 Language Factory with HTTP Import Resolution

A new WDL Draft 2 language factory has been added, enabling the execution of workflows using the Draft 2 specification. This implementation introduces support for HTTP-based import resolution, allowing workflows to import external dependencies via HTTP URLs. It also includes validation logic that enforces a maximum length of 100 characters for workflow names and integrates with the existing parser caching mechanism for improved performance.

languageFactories/wdl-draft2/src/main · high confidence

Introduction of engine core infrastructure and failure mode handling

The engine module now includes foundational components for workflow execution, specifically introducing the \WorkflowFailureMode\ trait with \ContinueWhilePossible\ and \NoNewCalls\ options to control how workflows proceed after call failures. This change adds the \EngineWorkflowDescriptor\ to manage workflow state and path builders, \EngineFilesystems\ for configuring local and pluggable filesystems, and \EngineIoFunctions\ for I/O operations. Additionally, a \CromwellTerminator\ trait is introduced to handle coordinated shutdowns, and utility extensions are added to the package object for handling fully qualified names and job output maps.

engine/src/main/scala/cromwell/engine · high confidence

Introduction of new workflow engine actor components

This change introduces the core actor infrastructure for the workflow engine, including the WorkflowManagerActor for coordinating workflow execution, the WorkflowActor for managing individual workflow lifecycles, and the SingleWorkflowRunnerActor for handling single workflow runs. It also adds supporting components like WorkflowDockerLookupActor for managing Docker image hashing and helper traits for metadata publishing. These new files establish the foundational actors and state machines that drive workflow submission, execution, and finalization.

engine/src/main/scala/cromwell/engine/workflow · high confidence

Introduction of standardized instrumentation keys and prefixes

Cromwell now provides a centralized set of constants for metric instrumentation, defining standard keys for workflow outcomes (success, aborted, failure, retry) and prefixes for different system components (API, backend, job, I/O, services, workflow, and CromIAM). This allows services to communicate metrics consistently and ensures that instrumentation data is properly categorized by source.

core/src/main/scala/cromwell/core/instrumentation · high confidence

Introduction of validated Label and Labels data structures

Cromwell now includes a core \Label\ and \Labels\ abstraction in the \cromwell.core.labels\ package to manage key-value metadata. These structures enforce validation rules, specifically ensuring that label keys and values do not exceed 255 characters and that keys are non-empty, providing a standardized way to handle label data within the workflow engine.

core/src/main/scala/cromwell/core/labels · high confidence

New API client model types and JSON serialization support

The Cromwell API client now includes a comprehensive set of model classes and JSON formatters for interacting with the Cromwell server. Users can now parse and construct data for workflow submissions (including single and batch modes with labels and zipped imports), workflow status tracking, query results, call cache diffs, workflow descriptions, and cost information. The client also standardizes time formatting to UTC with millisecond precision and provides utilities for handling HTTP responses via Cats Effect IO.

cromwellApiClient/src/main/scala/cromwell/api/model · high confidence

New AWS integration test cases for Mutect2 and CNV-Pair workflows

Added new integration test configurations for the Mutect2 and CNV-Pair somatic workflows to support execution on AWS infrastructure. This includes new WDL workflow definitions, input JSON files pointing to S3 data sources (e.g., \s3://cromwell-integration-tests/...\), and runtime options for Google Cloud Platform (for CNV-Pair). The Mutect2 test specifically configures the workflow to run with Oncotator annotation and scatter over 50 intervals, while the CNV-Pair test covers matched-pair analysis with optional Oncotator support.

centaur/src/main/resources/integrationTestCases/Somatic/Mutect2 · high confidence

New Centaur integration test reporting infrastructure

The Centaur integration test framework now includes a new reporting subsystem located in \centaur/src/it/scala/centaur/reporting\. This change introduces a pluggable error reporting architecture (\ErrorReporter\, \ErrorReporters\) that supports multiple destinations, including Google Cloud Storage (\GcsReporter\) for test metadata and SLF4J (\Slf4jReporter\) for logging. It also adds support for querying the Cromwell database directly from tests (\ErrorReporterCromwellDatabase\) to retrieve job key-value and metadata entries, and provides utilities for aggregating exceptions (\AggregatedIo\) and capturing CI environment details (\CiEnvironment\).

centaur/src/it/scala/centaur/reporting · high confidence

New Cromwell reference disk manifest creator tool

Added a new Java application (CromwellRefdiskManifestCreatorApp) that scans a directory of reference files and generates a JSON manifest. The tool calculates CRC32C checksums for each file, supports parallel processing via configurable thread counts, and includes the image identifier and disk size (diskSizeGb) in the output manifest. A unit test verifies the manifest creation logic against a reference JSON file.

CromwellRefdiskManifestCreator · high confidence

New GCP configuration and HTTP transport options

Added GoogleConfiguration to parse and validate GCP authentication settings (service account, user account, application default, and mock modes) from the config file, and introduced GoogleHttpTransportOptions to set default HTTP read timeouts for GCP API calls.

cloudSupport/src/main/scala/cromwell/cloudsupport/gcp · high confidence

New JobStore actor system for job completion tracking

The engine now includes a new JobStore subsystem (in engine/src/main/scala/cromwell/jobstore) that manages job completion data. This introduces an EmptyJobStoreActor (a no-op implementation), a JobStore trait with SQL-backed SqlJobStore for persisting results, and a set of Akka actors (JobStoreActor, JobStoreReaderActor, JobStoreWriterActor) that handle batching, throttling, and graceful shutdown of job completion writes and reads. A new JSON formatter (JobResultJsonFormatter) serializes job results, and core data types (JobStoreKey, JobResultSuccess, JobResultFailure) are defined to represent completed jobs and their outputs or failures.

engine/src/main/scala/cromwell/jobstore · high confidence

New PollResultMonitorActor for tracking VM start/end times and cost data

A new PollResultMonitorActor has been introduced in the standard backend poll monitoring package to process poll results from backends. This actor tracks the earliest job start time, the VM start and end times (updating them if later values are received), and initiates cost-per-hour lookups by extracting VM info. It reports these metrics to the Cromwell metadata service, ensuring that cost calculations and metadata reflect the most recent timing data observed during polling.

backend/src/main/scala/cromwell/backend/standard/pollmonitoring · high confidence

New WDL AST-to-model transformation components

Added a new set of Scala transformation classes in the \wdl/transforms/new-base\ module that convert parsed WDL Abstract Syntax Tree (AST) nodes into the internal WDL Model (WDLom) representation. These components handle the conversion of various WDL constructs, including expressions, command sections, task and workflow definitions, struct literals, imports, and metadata sections, providing the foundational logic for parsing WDL documents into the new model structure.

wdl/transforms/new-base · high confidence

New backpressure\_report script to analyze Cromwell runner I/O backpressure

A new Python tool has been added at scripts/backpressure\_report to measure the time Cromwell runner instances spend in a high I/O state that triggers internal backpressure, which otherwise slows job starts and restarts. The script parses JSON-formatted Google Logs Explorer logs (specifically looking for 'IoActor backpressure' start and end messages), calculates the duration of these backpressure events per pod, and outputs a CSV report showing total and per-pod backpressure durations in hourly intervals. This allows users to identify which pods are contributing most to backpressure and understand its impact on job execution throughput.

_scripts/backpressure\_report, scripts/backpressure\_report/backpressure\report · high confidence

New collection utilities and data structures

Added new common library components including a \Table\ data structure for row/column/value storage, a \WeightedQueue\ for managing items with associated weights, and \EnhancedCollections\ which provides safe, strict map value transformations (\safeMapValues\), type-based filtering, and weighted queue operations. Also introduced \IntegerUtil\ for integer ordinal formatting.

common/src/main/scala/common/collections · high confidence

New core actor utilities for batching, throttling, and stream integration

Added a new set of actor helpers in the core actor package to improve workflow execution reliability and performance. BatchActor provides an abstract finite state machine for queueing and batching commands based on configurable batch sizes and flush rates, supporting weighted queues. ThrottlerActor extends this to process commands one at a time, effectively serializing execution. StreamActorHelper integrates Akka Streams with actor logic, handling backpressure by sending BackPressure messages when stream enqueue operations fail or are rejected, and ensuring graceful shutdown. RobustClientHelper adds exponential backoff and request timeout management for inter-actor communication, preventing request loss during transient failures.

core/src/main/scala/cromwell/core/actor · high confidence

New example backend configuration files for Cromwell

The \cromwell.example.backends\ directory now includes example configuration files for a wide range of backend providers, allowing users to easily copy and customize settings for their environment. Supported backends include cloud providers (AWS Batch, Google Cloud Batch, PAPIv2, TES, TESK), local and container runtimes (Docker, Singularity, Singularity+Slurm, udocker, udocker+Slurm), and HPC workload managers (HtCondor, LSF, SGE, SLURM, TORQUE, Volcano). A comprehensive README and a central \cromwell.examples.conf\ file are also provided to guide users on how to integrate these providers into their Cromwell setup.

cromwell.example.backends · high confidence

New exception aggregation utilities and IO helper

The common module now includes new exception handling capabilities. ExceptionAggregation.scala introduces traits (MessageAggregation, ThrowableAggregation) and case classes (AggregatedException, AggregatedMessageException, CompositeException) that allow multiple error messages or throwables to be aggregated into a single exception, with improved formatting for file-not-found errors. Additionally, package.scala adds a toIO helper that converts an Option to a Cats Effect IO, failing with a RuntimeException if the option is empty.

common/src/main/scala/common/exception · high confidence

New metadata alert types for workflow monitoring

A new sealed trait \MetadataAlert\ and its implementations \HeavyMetadataAlert\ and \MaxMetadataAlert\ have been added to the core events module. These case classes provide a structured way to report metadata-related issues for specific workflows, including the workflow ID, the current count, and (in the case of \MaxMetadataAlert\) the configured limit, enabling downstream systems to detect and react to metadata thresholds.

core/src/main/scala/cromwell/core/events · high confidence

New metadata comparison scripts for PAPI performance analysis

Added a new \metadata\_comparison\ toolset to extract, digest, and compare performance metadata from Cromwell workflows running on PAPI. The \extractor\ script fetches workflow and operation metadata from Cromwell and Google Cloud Storage, including local Cromwell checkouts and configurations. The \digester\ processes this raw metadata into structured JSON digests, calculating timing metrics (startup, docker pull, localization, user command, delocalization) and cost estimates based on machine types and disk usage. The \comparer\ script then takes two digests and produces a CSV report comparing performance and costs between two workflow runs (e.g., PAPI v1 vs v2), with options to remove call prefixes for cleaner reporting.

_scripts/metadata\_comparison/metadata\comparison · high confidence

New retry and backoff utilities for handling transient failures

The core retry module now includes new \GoogleBackoff\ and \Retry\ components. \GoogleBackoff\ provides \InitialGapBackoff\ and \SimpleExponentialBackoff\ strategies that wrap Google's \ExponentialBackOff\ for configurable retry intervals. \Retry\ introduces \withRetry\, a generic function to retry asynchronous operations with customizable backoff, transient/fatal error handling, and retry limits, and \withRetryForTransactionRollback\, which specifically handles SQL deadlock exceptions (\SQLTransactionRollbackException\) by retrying up to a specified number of times.

core/src/main/scala/cromwell/core/retry · high confidence

New runtime attribute validation framework for Cromwell backends

This change introduces a comprehensive validation system for workflow runtime attributes (such as memory, CPU, GPU requirements, container images, and retry logic) within the Cromwell backend. It replaces ad-hoc validation with a structured, type-safe approach using traits like \RuntimeAttributesValidation\ and \ValidatedRuntimeAttributesBuilder\, ensuring that attributes like \memory\, \cpu\, \gpuRequired\, \continueOnReturnCode\, and \docker\/\container\ are correctly coerced, validated against WDL specifications (including WDL 1.1 features), and consistently handled for call caching and metadata reporting.

backend/src/main/scala/cromwell/backend/validation · high confidence

New shared transform utilities for WDL 1.0 and 1.1

A new \common/transforms\ package has been introduced to provide shared transformation logic for both WDL 1.0 and 1.1. This change adds a \CheckedAtoB\ type alias based on \cats.data.Kleisli\ and a companion object containing factory methods (\fromCheck\, \fromErrorOr\) that wrap validation logic with optional context for error reporting. It also includes a \firstSuccess\ helper to attempt multiple transformation options and return the first successful result, consolidating common transform patterns previously duplicated across WDL versions.

common/src/main/scala/common/transforms · high confidence

New utility libraries for actor management and resource handling

Added several new helper classes to the core utility package to improve actor lifecycle management and code safety. GracefulShutdownHelper enables actors to cleanly shut down dependent actors before stopping themselves. PromiseActor provides a timeout-free alternative to Akka's ask pattern for synchronous-style actor communication. DatabaseUtil introduces retry logic with exponential backoff for transient database errors. StopAndLogSupervisor simplifies actor supervision by automatically stopping failed children and logging the error. TryWithResource implements a Scala version of Java's try-with-resources for safe handling of AutoCloseable resources.

core/src/main/scala/cromwell/util · high confidence

New utility libraries for retry logic, URI masking, and time formatting

This change introduces a suite of new utility classes in the common package to support safer and more consistent operations across the application. IORetry provides a stateful retry mechanism for Cats Effect IO operations, allowing callers to accumulate state and handle failures with configurable backoff strategies. UriUtil and StringUtil now include methods to mask sensitive parts of URIs (such as credentials and signatures) before logging, addressing security concerns around sensitive data exposure. TimeUtil standardizes timestamp formatting to UTC with millisecond precision, ensuring consistent time representation in logs and outputs. Additional utilities include Backoff for defining retry intervals, IntrospectableLazy for managing lazy initialization with existence checks, TryUtil for aggregating multiple Scala Try results, TerminalUtil for formatted console output, and VersionUtil for reading project version metadata from SBT-generated config files.

common/src/main/scala/common/util · high confidence

New validation library with ErrorOr, Checked, and IOChecked types

The common module now includes a new validation library providing \ErrorOr\ (based on cats \Validated\), \Checked\ (based on \Either\), and \IOChecked\ (based on \EitherT\[IO\]\) types. These types offer consistent, composable error handling with support for accumulating multiple errors, parallel evaluation of IO operations, and convenient conversion between \Try\, \Future\, and \IO\ contexts, simplifying validation logic across the application.

common/src/main/scala/common/validation · high confidence

New workflow finalization actors for output copying, log management, and callbacks

Cromwell introduces dedicated actors in the workflow finalization lifecycle to handle post-execution tasks more robustly. The new CopyWorkflowOutputsActor manages the copying of workflow output files to a final destination, including validation to prevent destination path collisions. CopyWorkflowLogsActor handles the copying and cleanup of workflow logs, ensuring metadata is updated with the final log location. Additionally, WorkflowCallbackActor enables configurable HTTP callbacks upon workflow completion, allowing external systems to be notified of workflow status and results. These changes are orchestrated by the WorkflowFinalizationActor, which coordinates backend-specific finalization with these new engine-level finalization steps.

engine/src/main/scala/cromwell/engine/workflow/lifecycle/finalization · high confidence

New workflow timing diagram visualization

A new HTML-based timing diagram view is now available for workflows, rendering execution events as a Google Timeline chart. This view visualizes call start and end times, including nested subworkflows, retry attempts, and specific execution phases such as queued, starting, and overhead periods. Users can interact with the chart to expand or collapse subworkflow details, providing a clearer view of workflow execution history and performance bottlenecks.

engine/src/main/resources/workflowTimings · high confidence

Pluggable filesystem configuration system

Cromwell now supports a pluggable filesystem architecture, allowing users to configure and register custom path builder factories via the global configuration. The new CromwellFileSystems component validates filesystem entries, instantiates the specified classes, and manages singleton configurations, enabling extensibility beyond the default local filesystem without code changes.

core/src/main/scala/cromwell/core/filesystem · high confidence

Support for HTTP/HTTPS workflow inputs

Users can now provide HTTP or HTTPS URLs as inputs for workflows. The system automatically detects these URLs, downloads the content to a temporary local file, and makes it available for execution, while also supporting size checks via HEAD requests.

filesystems/http · high confidence

Support for runtime attribute overrides and new engine functions in WDL shared transforms

This change introduces new capabilities in the WDL shared transforms layer. It adds a \transpose\ function to the engine functions, allowing users to transpose two-dimensional arrays in their workflows. Additionally, it implements runtime attribute overrides, enabling users to specify task runtime requirements (such as CPU, memory, and disk) directly in the input JSON using dot-notation keys (e.g., \task.runtime.cpu\), which are now correctly parsed and grouped into runtime objects. The update also includes logic to automatically add default outputs for scatter and conditional blocks in the WOM graph, ensuring that outputs from these control structures are exposed by default.

wdl/transforms/shared · high confidence

WDL Draft 2 workflow and task transformation to WOM

This change introduces the core transformation logic for converting WDL Draft 2 source code into the Workflow Execution Model (WOM). The new files in \wdlom2wom\ implement the builders for workflows, tasks, call nodes, scatter nodes, and conditional nodes, enabling Cromwell to parse and execute Draft 2 compliant workflows. It also includes specific handling for Draft 2 features such as aliased calls, input validation, and file read limits.

wdl/transforms/draft2 · high confidence

Womtool introduces dedicated commands for listing workflow inputs and outputs

Womtool now includes \inputs\ and \outputs\ commands that generate JSON representations of a workflow's expected inputs and declared outputs, respectively. The \inputs\ command supports flags to include optional inputs (\--optional-inputs\) and runtime override inputs (\--runtime-override-inputs\), while the \outputs\ command provides a structured JSON list of output types. These additions complement existing validation, parsing, and graph visualization capabilities.

womtool · high confidence

Architecture

Database access layer refactored into domain-specific SQL traits

The database access layer in the SQL module has been restructured to separate concerns by domain. The previous monolithic database interface has been split into distinct traits—\MetadataSqlDatabase\, \EngineSqlDatabase\, \CallCachingSqlDatabase\, \WorkflowStoreSqlDatabase\, \JobStoreSqlDatabase\, \SubWorkflowStoreSqlDatabase\, \DockerHashStoreSqlDatabase\, \JobKeyValueSqlDatabase\, and \GroupMetricsSqlDatabase\—each exposing methods specific to its area of responsibility. \EngineSqlDatabase\ now aggregates the engine-related traits, while \MetadataSqlDatabase\ handles workflow metadata, summarization, and archival status. This change also introduces new join models (\CallCachingJoin\, \JobStoreJoin\, \CallCachingDiffJoin\) and a \MetadataJobQueryValue\ filter mechanism to support more granular metadata queries.

database/sql/src/main/scala/cromwell/database/sql · high confidence

Database access layer restructured into separated Engine and Metadata components

The database access layer has been split into two distinct Slick components—EngineDataAccessComponent and MetadataDataAccessComponent—to separate engine-related data (such as job store, call caching, and group metrics) from metadata-related data (such as workflow summaries, custom labels, and metadata entries). This separation allows for independent configuration and management of the engine and metadata database schemas, improving modularity and aligning with the system's architectural shift toward decoupled service stores.

database/sql/src/main/scala/cromwell/database/slick/tables · high confidence

Database access logic is split into separate Engine and Metadata services

The database access layer in the Slick package has been restructured to separate engine operations (workflow store, job store, call caching, group metrics) from metadata operations. This change introduces distinct \EngineSlickDatabase\ and \MetadataSlickDatabase\ implementations, each with their own data access components and configuration paths. The engine service now handles workflow state management, job key-value storage, and call caching lookups with specific timeout and isolation settings, while the metadata service manages workflow metadata entries and summarization queues. This separation allows for independent configuration and transaction management of these two distinct data domains.

database/sql/src/main/scala/cromwell/database/slick · high confidence

Introduce pluggable filesystem path abstraction

The core path handling logic has been refactored to support pluggable filesystems. A new \Path\ trait and \PathBuilder\ interface allow Cromwell to resolve file paths for different storage backends (such as local, GCS, or DRS) via a prioritized list of builders. This change replaces the previous direct dependency on \better.files\ and \java.nio.file\ with a unified abstraction, enabling the system to handle diverse URI schemes and backend-specific behaviors while maintaining a consistent API for file operations.

core/src/main/scala/cromwell/core/path · high confidence

Introduction of Standard Backend architecture and Dummy test backend

The backend execution model has been refactored to use a new 'Standard' trait-based architecture. This introduces a standardized lifecycle for asynchronous job execution, including new base actors for initialization, execution, and finalization that common backends (like JES and SFS) now inherit from. Additionally, a new 'Dummy' backend implementation has been added to the \backend/src/main/scala/cromwell/backend/dummy\ package, providing a mock backend for testing purposes that simulates job execution without contacting real cloud providers.

backend/src/main/scala/cromwell/backend/standard · high confidence

New services sub-project for modular service infrastructure

The codebase has introduced a new \services\ sub-project to centralize and modularize service-related components. This change includes the creation of a \ServiceRegistryActor\ for dynamic service discovery and configuration, along with dedicated stores for engine and metadata database interactions (\EngineServicesStore\, \MetadataServicesStore\). It also adds foundational infrastructure such as \EnhancedBatchActor\ and \EnhancedThrottlerActor\ for standardized instrumentation and load control, and introduces new capabilities like a \GithubAuthVending\ service for GitHub token resolution and a \GcpCostCatalogService\ for GCP cost lookups.

services · high confidence

Refactor web service routes into modular support traits

The web service API routes have been reorganized from a single monolithic service into separate, composable traits: \CromwellApiService\ (engine and workflow submission/abort/release endpoints), \MetadataRouteSupport\ (metadata, status, labels, and cost queries), and \WomtoolRouteSupport\ (workflow description). This structural change improves maintainability by isolating distinct functional areas of the Cromwell API while preserving all existing endpoint behaviors.

engine/src/main/scala/cromwell/webservice/routes · high confidence

Behavioural changes

AWS integration test workflow for HaplotypeCaller added

Added AWS-specific integration test files (inputs, options, and WDL) for the HaplotypeCaller workflow, enabling users to run the GATK4 HaplotypeCaller in GVCF mode on AWS using S3 data sources and Broad's reference bucket. The new test configuration specifies the GATK 4.0.0.0 Docker image, sets Java memory options to 8 GB for HaplotypeCaller and 30 GB for MergeGVCFs, allocates 100 GB of disk space per task, and configures preemptible instance retries to 3.

centaur/src/main/resources/integrationTestCases/germline/haplotype-caller-workflow · high confidence

CI infrastructure restructured into modular, shared test scripts

The CI test execution logic has been refactored from monolithic scripts into a modular system. A new \src/ci/bin/test.inc.sh\ provides shared build variables and environment setup, while backend-specific tests (AWS, GCP Batch, TES, Slurm, etc.) are now driven by dedicated entry scripts (e.g., \testCentaurAws.sh\) that source common libraries like \test\_gcpbatch.inc.sh\ and \test\_singularity.inc.sh\. This change introduces a standardized \cromwell::build\ namespace for test functions, improves environment isolation, and adds a new \get\_cromwell\_hosts.py\ utility to dynamically discover Cromwell container hosts for load-balanced testing.

src/ci · high confidence

Call caching blacklist support and root workflow file hash caching

This change introduces a local blacklist mechanism for call caching to prevent repeated failures when copying cache hits from problematic sources. The new \BlacklistCache\ and \CallCachingBlacklistManager\ components allow Cromwell to track and skip blacklisted cache entries or storage buckets based on configuration, while \CopyingActorBlacklistCacheSupport\ integrates this logic into the cache hit copying process to automatically blacklist failed sources and whitelist successful ones. Additionally, \RootWorkflowFileHashCacheActor\ adds a caching layer for root workflow file hashing to reduce redundant I/O operations and handle timeouts more gracefully, improving reliability and performance during workflow execution.

backend/src/main/scala/cromwell/backend/standard/callcaching · high confidence

Centaur integration test suite restructured with new restart and upgrade capabilities

The Centaur testing framework has been reorganized into a dedicated subdirectory with a new README and a \centaur.sh\ wrapper for Docker-based execution. The test runner script (\test\_cromwell.sh\) now supports running against a specific Cromwell branch or JAR, allows parallel test execution, and enables tag-based inclusion/exclusion. A new \engineUpgradeTestCases\ directory introduces an \engine\_upgrade\ tag and a \cromwell\_restart.test\ case, which validates Cromwell's ability to recover and continue workflow execution after a server restart.

centaur · high confidence

Centaur integration tests now use a dedicated Cromwell client with robust retry logic

The Centaur test suite now uses a new \CentaurCromwellClient\ that wraps Cromwell API calls in \cats-effect IO\ and implements a configurable retry mechanism for transient network or server errors. This change ensures that test failures caused by temporary Cromwell unavailability (such as connection timeouts or 404s during startup) are automatically retried up to five times with a fixed delay, rather than failing immediately. It also introduces a dedicated daemon thread pool for blocking HTTP operations to prevent thread starvation and adds a lightweight \isAlive\ health check that queries the Cromwell database directly to verify service readiness before submitting workflows.

centaur/src/main/scala/centaur/api · high confidence

Centaur test infrastructure restructured with new upgrade and parallel execution suites

The Centaur integration test framework has been reorganized to support distinct test execution modes and engine upgrade validation. A new base class, AbstractCentaurTestCaseSpec, centralizes test discovery, duplicate name detection, and retry logic using Cats Effect IO. Two new abstract suites, AbstractCromwellEngineOrBackendUpgradeTestCaseSpec and its concrete implementations (EngineUpgradeTestCaseSpec, PapiUpgradeTestCaseSpec), enforce empty database constraints before and after engine upgrades. ParallelTestCaseSpec and SequentialTestCaseSpec now explicitly separate test execution based on the test format, allowing parallel runs by default while ensuring restart-heavy tests run sequentially. An ExternalTestCaseSpec allows running single test files via environment variable, and a new CromwellDatabaseCallCaching utility supports clearing call cache entries during retries.

centaur/src/it/scala/centaur · high confidence

Centaur test runner refactored to support managed Cromwell modes

The Centaur test infrastructure has been restructured to introduce explicit support for managing Cromwell server lifecycles via JAR execution or Docker Compose, in addition to the existing unmanaged mode. New configuration classes (JarCromwellConfiguration, DockerComposeCromwellConfiguration) and a CromwellManager handle starting, stopping, and health-checking the Cromwell process, including a configurable timeout and custom exit code on failure. The test runner now reads its settings from a centralized CentaurConfig, and a new CromwellTracker component allows asserting 'horicromtality' (balanced workload distribution) across backends using statistical tests.

centaur/src/main/scala/centaur · high confidence

Centralized backend configuration and factory initialization

The engine now centralizes backend setup through a new \BackendConfiguration\ module that reads backend definitions from the application configuration. This change introduces a global \CromwellBackends\ singleton that instantiates \BackendLifecycleActorFactory\ instances for all configured backends at startup, ensuring that backend factories are available for use by the workflow engine. Users will benefit from a more robust and centralized backend initialization process that validates backend names and configurations early.

engine/src/main/scala/cromwell/engine/backend · high confidence

Configurable read limits and new value evaluation utilities

The WDL model now supports configurable limits for reading various data types (lines, bool, int, float, string, json, tsv, map, object) via a new \FileSizeLimitationConfig\ that loads settings from the \system.input-read-limits\ configuration block, defaulting to \Int.MaxValue\. Additionally, new utility objects \FileEvaluatorUtil\ and \ValueEvaluation\ have been added to support value serialization to JSON and TSV, and to identify files for delocalization within WDL expressions.

wdl/model/shared · high confidence

Configured logging for integration tests

Added application.conf and logback-test.xml to the integration test resources to control log output. Akka logs are now routed through SLF4J with dead-letter logging disabled, and specific libraries (MockServer, Liquibase, HikariCP, and Google Cloud Storage) are set to WARN or ERROR levels to reduce noise in test logs.

centaur/src/it/resources · high confidence

CromIAM API service routes and authorization logic

The CromIAM webservice now exposes a comprehensive set of API endpoints for workflow management, including submission, querying, status checks, metadata retrieval, logging, aborting, and cost estimation. The service enforces user authorization via SamClient checks before forwarding requests to Cromwell, supports collection-based access control, and includes a restructured Swagger UI with OAuth2 support. Workflow submissions now support both source code and URL-based inputs, with optional hold states and dependency handling.

CromIAM/src/main/scala/cromiam/webservice · high confidence

CromIAM configuration and logging infrastructure

CromIAM now includes its own standalone configuration and logging setup. The new application.conf defines service endpoints for CromIAM (port 8001), SAM (port 443), and Cromwell (port 8000), and sets HTTP request and connection timeouts to 40 seconds. A dedicated logback.xml configures logging with enhanced thread/date formatting, supports Standard, Pretty, and FileRoller modes, and routes WARN-level events to Sentry. A sentry.properties file is also added to initialize the Sentry DSN.

CromIAM/src/main/resources · high confidence

CromIAM introduces Sam client with submit whitelist and instrumentation

The CromIAM service now includes a dedicated Sam client (SamClient) that enforces a configurable submit whitelist check before allowing workflow submissions, rejecting unauthorized users with a specific denial response. This change also adds comprehensive StatsD instrumentation for Sam interactions and introduces new authentication models (User, Collection) to support these authorization checks.

CromIAM/src/main/scala/cromiam/sam · high confidence

CromIAM now strips Host and Timeout-Access headers when forwarding requests to Cromwell

The new CromwellClient implementation in CromIAM explicitly removes the 'Host' and 'Timeout-Access' HTTP headers before forwarding requests to the underlying Cromwell service. Stripping the Host header prevents host-based routing failures in front-end proxies by ensuring Cromwell receives the correct target host rather than CromIAM's, while removing the Timeout-Access header resolves Akka HTTP warnings regarding disallowed headers. This change ensures reliable request forwarding and cleaner logging for CromIAM users.

CromIAM/src/main/scala/cromiam/cromwell · high confidence

CromIAM server configuration model and validation

The CromIAM server configuration is now defined by a new \CromIamServerConfig\ case class that aggregates settings for CromIAM, Cromwell, the SAM client, and Swagger OAuth. This change introduces a robust validation layer using \ErrorOr\ to ensure configuration paths are present and correctly typed, specifically adding a \check-submit-whitelist\ boolean flag to the \SamClientConfig\ to control access restrictions.

CromIAM/src/main/scala/cromiam/server/config · high confidence

CromIAM server initialization and shutdown behavior

The CromIAM server entry point (CromIamServer.scala) has been introduced to handle server startup and shutdown. Upon starting, the server logs its version number using VersionUtil. It initializes an Akka ActorSystem, materializer, and execution context, and configures routes including Swagger UI. A custom shutdown hook is registered to log 'Shutting down the server' and signal completion when the process receives a shutdown signal (e.g., Ctrl-C), replacing the default behavior.

CromIAM/src/main/scala/cromiam/server · high confidence

Database migration for workflow options encryption and key renaming

This change introduces database migration logic to handle workflow options stored in the metadata and workflow store. It adds a migration to rename legacy workflow option keys (such as 'defaultRuntimeOptions' and 'workflowFailureMode') to their new snake\_case equivalents (e.g., 'default\_runtime\_attributes') in the METADATA\_ENTRY table. Additionally, it provides infrastructure to encrypt workflow options in the WORKFLOW\_STORE\_ENTRY table and clear encrypted values from the METADATA\_ENTRY table, ensuring that sensitive data is properly secured and formatted according to the new configuration.

database/migration/src/main/scala/cromwell/database/migration/workflowoptions · high confidence

Database migration refactoring and job store simpleton handling

The database migration logic has been restructured to improve reliability and maintainability. A new abstract base class, AbstractRestartMigration, now standardizes how migrations interact with the database connection, ensuring transactions are properly committed. The JobStoreSimpletonMigration has been updated to correctly handle nullable job store simpletons, preventing failures that occurred when Cromwell recovered after a migration due to incorrect quoting of file values. Additionally, the RenameWorkflowOptionKeysMigration now explicitly handles the renaming of workflow option keys in the WORKFLOW\_STORE table, ensuring consistency in stored workflow configurations.

database/migration/src/main/scala/cromwell/database/migration/restart · high confidence

Database migration to normalize failure metadata structure

A new database migration has been added to the failure metadata module to restructure how workflow failure information is stored in the METADATA\_ENTRY table. The migration introduces two batched tasks: ExpandSingleFailureStrings, which splits flat failure entries (e.g., 'failures\[0\]') into distinct message and causedBy components (e.g., 'failures\[0\]:message' and 'failures\[0\]:causedBy\[\]'), and DeduplicateFailureMessageIds, which resolves duplicate failure message records by assigning them unique identifiers. This ensures that failure metadata is consistently structured for downstream processing.

database/migration/src/main/scala/cromwell/database/migration/failuremetadata · high confidence

Database schema updates for call caching, execution tracking, and workflow metadata

This release applies a series of Liquibase migration changesets to the database schema. It introduces a new GROUP\_METRICS\_ENTRY table to track Cloud Quota exhaustion events and adds a HOG\_GROUP column to the WORKFLOW\_STORE\_ENTRY table. Call caching is enhanced with a new CALL\_CACHING\_AGGREGATION\_ENTRY table for optimized hash lookups, a CREATED\_AT timestamp on CALL\_CACHING\_ENTRY, and a new CALL\_CACHING\_JOB\_DETRITUS table to store job output files. The EXECUTION table is expanded with an ATTEMPT column (making it part of the unique constraint), an IDX column, and START\_DT/END\_DT timestamps, while also adding columns for result reuse tracking (ALLOWS\_RESULT\_REUSE, DOCKER\_IMAGE\_HASH, EXECUTION\_HASH). Additionally, the schema adds a WORKFLOW\_NAME column to WORKFLOW\_EXECUTION, a WORKFLOW\_URL column to WORKFLOW\_STORE\_ENTRY, and a new BACKEND\_KV\_STORE table for backend-specific key-value pairs.

database/migration/src/main/resources/changesets · high confidence

Enhanced metadata flattening to include attempt numbers in shard keys

The \JsonUtils\ module in the Centaur testing framework now supports flattening Cromwell metadata responses in a way that preserves both shard indices and attempt numbers. Previously, flattening shard arrays only included the shard index in the resulting key structure; the new logic detects arrays containing both \shardIndex\ and \attempt\ fields and generates flattened keys that incorporate both (e.g., \shardIndex.attemptNumber\). This ensures that metadata for multiple attempts of the same shard is distinct and not lost during conversion, while maintaining backward compatibility with older structures that only used shard indices.

centaur/src/main/scala/centaur/json · high confidence

Improved GCS error handling and credential retry logic

The GCS filesystem now provides more robust error reporting and automatic recovery for common issues. It automatically retries credential acquisition up to three times to handle transient authentication failures. Additionally, it detects 'Requester Pays' bucket errors where the project ID is missing and transparently retries the request with the project ID set, while also converting specific 'Not Found' storage exceptions into standard Java FileNotFoundExceptions for clearer error messages.

filesystems/gcs/src/main/scala/cromwell/filesystems/gcs · high confidence

Improved database schema synchronization and migration reliability

The database migration system now includes a new DiffResultFilter to ignore harmless schema differences (such as VARCHAR length variations and column reordering) during synchronization checks, ensuring that minor cosmetic differences do not trigger false alarms. Additionally, Liquibase API access is now synchronized via a mutex to prevent concurrent access issues, and logging output from Liquibase is properly routed to the application logger instead of standard output.

database/migration/src/main/scala/cromwell/database/migration/liquibase · high confidence

Improved logging accuracy and reliability with Akka-aware converters and file handle management

This change introduces a new logging infrastructure in the core module to provide more accurate timestamps and thread information in logs, while also fixing file handle leaks. It adds custom Logback converters (EnhancedDateConverter, EnhancedThreadConverter) and an Akka logger (EnhancedSlf4jLogger) that prioritize Akka-specific MDC properties (akkaTimestamp, sourceThread) for more precise log entries. Additionally, it implements a WorkflowLogger that synchronizes file appender creation and teardown to prevent ConcurrentModificationExceptions and ensure file handles are properly closed, and includes a Java Logging Bridge to unify Java Util Logging with SLF4J.

core/src/main/scala/cromwell/core/logging · high confidence

Improved server shutdown and logging behavior

Cromwell now implements a coordinated, graceful shutdown process that stops new workflow submissions, unbinds the HTTP server, and orderly terminates internal actors (such as the I/O and job store actors) to prevent data loss. Additionally, log noise is reduced during shutdown by silencing dead-letter messages and specific Akka stream termination errors that occur while the system is winding down.

engine/src/main/scala/cromwell/server · high confidence

Introduce Call Caching v3 engine implementation

The engine's call caching mechanism has been replaced with a new v3 implementation. This change introduces a new set of actors to manage the hashing, reading, and writing of cache entries, including \EngineJobHashingActor\ to coordinate the process, \CallCacheHashingJobActor\ and \CallCacheReadingJobActor\ to handle individual job hashing and cache lookups, and \CallCacheWriteActor\ and \CallCacheReadActor\ to batch and execute database operations via the new \CallCache\ accessor. Additionally, a \CallCacheDiffActor\ is added to allow users to compare the caching metadata of two different calls to identify why they did not match.

engine/src/main/scala/cromwell/engine/workflow/lifecycle/execution/callcaching · high confidence

Introduce dedicated job and subworkflow preparation actors

The engine now uses new \JobPreparationActor\ and \SubWorkflowPreparationActor\ components to handle the preparation phase for workflow calls. The \JobPreparationActor\ manages the sequence of evaluating call inputs and runtime attributes, performing Docker hash lookups when necessary, and fetching Key/Value store entries before constructing the final \BackendJobDescriptor\. It also includes a safety check to validate GPU requirements against backend capabilities. The \SubWorkflowPreparationActor\ specifically handles the evaluation of inputs for nested subworkflows, ensuring that all required inputs are provided before passing the prepared descriptor to the execution engine.

engine/src/main/scala/cromwell/engine/workflow/lifecycle/execution/job/preparation · high confidence

Introduce standard async backend execution traits and models

Adds a new \AsyncBackendJobExecutionActor\ trait and supporting data structures (\ExecutionHandle\, \ExecutionResult\, \KnownJobFailureException\) to the backend. This provides a standardized, retry-aware foundation for asynchronous job execution, handling execution modes (execute, recover, abort), robust polling with exponential backoff, and specific failure categorization (retryable vs. non-retryable, memory issues, wrong return codes) to improve reliability and error reporting for users running jobs on async backends.

backend/src/main/scala/cromwell/backend/async · high confidence

Introduces configurable call caching modes and file hashing strategies

This change introduces new core abstractions for call caching, including \CallCachingMode\ to explicitly control read/write behavior (off, read-only, write-only, or read-write) and \FileHashStrategy\ to define prioritized lists of hash types (such as Crc32c, Md5, Etag, and Sha256) for file integrity checks. It also adds \MaybeCallCachingEligible\ to handle Docker image hashing eligibility, distinguishing between fixed hashes and floating tags that may prevent caching.

core/src/main/scala/cromwell/core/callcaching · high confidence

Introduces new database table models for call caching, workflow metadata, and job storage

This change adds a comprehensive set of Scala case classes in the \cromwell.database.sql.tables\ package to represent database tables, including new entries for call caching aggregation and detritus, group metrics, sub-workflow tracking, and metadata summarization queues. It also updates existing models like \WorkflowStoreEntry\ to support new fields such as \customLabels\, \hogGroup\, and \importsZip\, while standardizing primary key types to \Long\ and ensuring LOB columns use \SerialClob\ for database compatibility.

database/sql/src/main/scala/cromwell/database/sql/tables · high confidence

Introduction of AWS authentication modes with credential validation

The AWS authentication system now supports multiple modes (no\_auth, custom\_key, default, and assume\_role) via the new AwsAuthMode abstraction. A key behavioral change is that credentials are validated immediately upon construction for static and default modes, and during the assume-role flow, ensuring that invalid credentials fail fast rather than at runtime during job execution.

cloudSupport/src/main/scala/cromwell/cloudsupport/aws/auth · high confidence

Introduction of ExecutionStore and ValueStore for workflow state management

Cromwell introduces two new internal components, ExecutionStore and ValueStore, to manage workflow execution state. ExecutionStore tracks the status of all job keys (including calls, expressions, scatters, conditionals, and subworkflows) and enforces a configurable limit on the total number of backend jobs a root workflow can create, preventing resource exhaustion. ValueStore manages the resolution and collection of output values, specifically handling the gathering of scattered shards and conditional outputs. These stores replace previous ad-hoc state tracking, providing a structured way to monitor job readiness and value availability during workflow execution.

engine/src/main/scala/cromwell/engine/workflow/lifecycle/execution/stores · high confidence

Introduction of asynchronous I/O actor with backpressure and retry support

The core I/O subsystem has been refactored to use an asynchronous Akka-based I/O actor. This change introduces an \AsyncIo\ client that exposes Futurized methods for file operations (read, write, copy, hash, etc.), allowing workflows to perform I/O without blocking. The implementation includes exponential backoff for retrying operations that encounter backpressure, configurable timeouts for different I/O actions, and a command builder pattern that allows backends to optimize specific commands (e.g., GCS batch operations). Additionally, the system now distinguishes between general I/O failures and specific read-forbidden errors, improving error reporting for users.

core/src/main/scala/cromwell/core/io · high confidence

Introduction of structured reference configuration files

Cromwell now provides a new, structured set of default configuration files (reference.conf, reference\_database.inc.conf, and reference\_local\_provider\_config.inc.conf) that define baseline settings for the web service, Akka dispatchers, system limits, database migration batching, and the local backend. This replaces the previous single-file approach, allowing users to include these references in their own cromwell.conf for clearer, modular configuration management.

core/src/main/resources · high confidence

Migrate Cromwell web service from Spray to Akka HTTP

The Cromwell REST API backend has been migrated from the legacy Spray framework to Akka HTTP. This change introduces new service traits (SwaggerService, SwaggerUiHttpService, WebServiceUtils) and actors (EngineStatsActor, LabelsManagerActor) to handle HTTP routing, Swagger UI serving, and workflow statistics/label management. Users will see the API continue to function via the same endpoints, but the underlying HTTP implementation is now Akka HTTP, which also includes a fix for a Swagger UI CVE ([GHSA redacted]) by redirecting legacy /swagger paths and a 304 response handling improvement.

engine/src/main/scala/cromwell/webservice · high confidence

New I/O Actor with backpressure and retry logic

The engine now uses a new IoActor to handle file operations asynchronously via Akka Streams, introducing configurable backpressure mechanisms that throttle I/O when commands become stale or load is high. It also implements a robust retry strategy for transient failures, specifically handling GCS 503/504 errors, AWS timeouts, and specific error messages like 'User project specified in the request is invalid' to improve resilience during workflow execution.

engine/src/main/scala/cromwell/engine/io · high confidence

New NIO-based file I/O implementation with cloud-native hashing

The engine introduces a new NIO-based file I/O layer (NioFlow and NioHashing) that handles standard file operations (read, write, copy, delete, etc.) via Akka Streams and Cats Effect IO. A key behavioral change is in file hashing: for cloud storage paths (GCS, S3, DRS, HTTP), the system now prioritizes retrieving stored hashes (e.g., GCS MD5/CRC32c, S3 ETag) rather than computing them locally, which prevents excessive CPU usage and network costs on large files. Local MD5 hashing is only performed for non-cloud paths. This implementation also includes retry logic and backpressure support for I/O commands.

engine/src/main/scala/cromwell/engine/io/nio · high confidence

New abort response types for workflow abort operations

Added a new \AbortResponse\ sealed trait and associated case classes (\WorkflowAbortFailureResponse\, \WorkflowAbortedResponse\, \WorkflowAbortRequestedResponse\) to model the outcomes of workflow abort requests. This provides a structured way to distinguish between successful aborts (whether immediately aborted or requested for later execution) and failures, improving error handling and status reporting for abort operations.

core/src/main/scala/cromwell/core/abort · high confidence

New execution keys for conditional, scatter, and expression nodes

The workflow engine now uses dedicated execution keys to manage the lifecycle of conditional blocks, scatter loops, and expression nodes. ConditionalKey handles the evaluation of 'if' conditions and manages the execution of scoped children, while ScatterKey and ScatterCollectorKey orchestrate the distribution of work across array shards and the subsequent gathering of results. ExpressionKey is introduced to evaluate WDL expressions using the backend's IoFunctionSet, and SubWorkflowKey manages the execution of nested workflows. These changes refine how the engine tracks and processes complex workflow structures.

engine/src/main/scala/cromwell/engine/workflow/lifecycle/execution/keys · high confidence

New metadata migration implementation for Cromwell database

The database migration module now includes a new set of Scala classes to migrate legacy execution, symbol, and workflow execution tables into the new METADATA\_JOURNAL schema. This change introduces specific migration handlers for execution events, failure events, call inputs/outputs, and workflow-level metadata, ensuring that historical workflow data is correctly transformed and stored in the updated metadata structure.

database/migration/src/main/scala/cromwell/database/migration/metadata · high confidence

New parallel GCS batch I/O flow with request isolation

The GCS I/O layer now uses a new Akka Stream-based batch processing architecture (GcsBatchFlow and ParallelGcsBatchFlow) to handle Google Cloud Storage requests. This change introduces parallelism by distributing batch commands across multiple worker flows, configurable via parallelism, max batch size, and max batch duration settings. To prevent request accumulation and timeouts caused by failed batch executions, the system creates a fresh BatchRequest instance for every batch execution rather than reusing a single object, ensuring that internal queues are not polluted by previous failures. The flow also implements backpressure mechanisms and handles specific error cases, such as read-forbidden access errors and 404 file-not-found scenarios, by providing clearer error messages and appropriate retry or failure responses.

engine/src/main/scala/cromwell/engine/io/gcs · high confidence

New unified CLI entry point with structured logging and hardened server configuration

The server now uses a new \CromwellApp\ entry point that supports \server\, \run\, and \submit\ commands, replacing the previous startup mechanism. This change introduces a dedicated \application.conf\ that hardens the Akka HTTP server by suppressing the \server\ header (preventing version disclosure) and increases the max response reason length to 1024 characters to capture longer error messages. It also extends the logger startup timeout to 30 seconds to prevent initialization failures. Logging is now managed via a new \logback.xml\ that supports \STANDARD\, \PRETTY\, and \FILEROLLER\ modes, integrates with Sentry for error tracking, and silences noisy libraries like HikariCP and Liquibase. Additionally, the entry point enforces a 3-minute DNS TTL to avoid infinite caching issues.

server · high confidence

Refactor backend I/O into modular path and glob function traits

The backend I/O logic in \cromwell.backend.io\ has been restructured into distinct, reusable traits to separate concerns. \WorkflowPaths\ and \JobPaths\ (including Docker-specific variants) now handle the construction and resolution of execution, input, and output directories, supporting configurable roots and sub-workflow nesting. \GlobFunctions\ and \DirectoryFunctions\ manage file discovery, with globs now resolved via list files to decouple pattern matching from execution, while \FileEvaluationIoFunctionSet\ prevents globbing during file evaluation to avoid side effects. This modularization standardizes how paths are generated and accessed across different backend implementations.

backend/src/main/scala/cromwell/backend/io · high confidence

Refactor workflow execution into lifecycle actors with dedicated metadata tracking

The workflow execution engine has been restructured to use a lifecycle-based actor model, introducing \WorkflowExecutionActor\ and \SubWorkflowExecutionActor\ to manage the execution state machine and job scheduling. A new \CallMetadataHelper\ trait centralizes the logic for publishing call-level metadata (such as status transitions, inputs, outputs, and execution events) to the Metadata Service, ensuring consistent and atomic reporting. Execution state and data are now encapsulated in \WorkflowExecutionActorData\ and \WorkflowExecutionDiff\, allowing for precise tracking of job statuses, value stores, and cumulative outputs without leaking internal state to other actors.

engine/src/main/scala/cromwell/engine/workflow/lifecycle/execution · high confidence

Refactor workflow lifecycle actors and output path handling

This change introduces new lifecycle infrastructure and refines how workflow outputs are located. A new \OutputsLocationHelper\ trait centralizes the logic for mapping output file paths, supporting a new \UseRelativeOutputPaths\ workflow option that allows outputs to be stored relative to the execution directory rather than the backend root. The \WorkflowLifecycleActor\ trait is restructured to manage backend actor lifecycles, specifically handling \ActorInitializationException\ by reporting failures and stopping the actor, while \TimedFSM\ provides basic timing instrumentation for state transitions.

engine/src/main/scala/cromwell/engine/workflow/lifecycle · high confidence

Refactored GCP Batch backend to use a new request manager architecture

The GCP Batch backend has been restructured to replace the previous execution model with a new \BatchApiRequestManager\ and associated worker actors. This change introduces a centralized request handling system that manages job submission, status polling, and abortion through dedicated actors (\BatchApiRunCreationClient\, \BatchApiStatusRequestClient\, \BatchApiAbortClient\), improving reliability and separation of concerns in how Cromwell interacts with the GCP Batch API.

supportedBackends/google/batch · high confidence

Refactored GCP authentication to use modern Google Auth Credentials API

The GCP authentication layer in cloudSupport has been updated to replace the deprecated \Credential\ class with the modern \Credentials\ adapter from the Google Auth library. This change introduces a new \GoogleAuthMode\ trait that standardizes credential creation and validation across service account, user, and application default modes. Users will benefit from improved scope handling, where specific scopes (like KMS and Genomics) are now managed explicitly rather than being bundled into the base mode, and credential lookups from workflow options are now strictly typed as \String =\> String\ maps. The refactoring also ensures that credentials are validated upon creation, providing earlier feedback on invalid configurations.

cloudSupport/src/main/scala/cromwell/cloudsupport/gcp/auth · high confidence

Refactored database migration to use batched processing

The database migration logic has been refactored to process large tables in batches rather than loading all rows into memory at once. This change introduces a new \BatchedTaskChange\ trait and supporting pagination utilities (\QueryPaginator\, \ResultSetIterator\) that allow migrations to read and write data in configurable chunks (controlled by \database.migration.read-batch-size\ and \database.migration.write-batch-size\). This improves stability and performance for workflows with large metadata tables by preventing memory exhaustion during migration. An example migration, \PreemptedIsABackendStatus\, now uses this batched approach to update execution statuses.

database/migration/src/main/scala/cromwell/database/migration/custom · high confidence

S3 filesystem provider upgraded to AWS SDK v2

The S3 filesystem implementation in \filesystems/s3/src/main/java\ has been migrated from the AWS SDK v1 to the AWS SDK v2. This change introduces new core classes (such as \AmazonS3ClientFactory\, \S3FileSystem\, and \S3FileChannel\) that utilize the v2 client libraries, enabling support for modern S3 features like cross-region access and updated authentication mechanisms while maintaining the JSR-203 NIO2 interface.

filesystems/s3/src/main/java · high confidence

Sanitized LanguageFactory interface and introduced ValidatedWomNamespace

The core language factory interface has been refactored to provide a cleaner, more explicit contract for workflow processing. The new LanguageFactory trait defines specific methods for generating WomBundles, creating executables, and validating namespaces, alongside configuration-driven controls for enabling factories and strict validation modes. Additionally, a new ValidatedWomNamespace case class now encapsulates the result of workflow validation, explicitly bundling the executable, input mappings, and imported file contents for downstream use.

languageFactories/language-factory-core/src/main/scala/cromwell/languages · high confidence

Standardized local database setup scripts for PostgreSQL and MySQL

The local development environment now uses dedicated shell scripts to launch PostgreSQL and MySQL containers with configurations aligned to Cromwell's CI requirements. The PostgreSQL setup includes a new initialization script that sets \max\_connections\ to 300 and enables the \lo\ (large object) extension, while the MySQL script launches a container compatible with the expected CI parameters. These changes replace previous ad-hoc or undocumented methods for spinning up local databases.

scripts/local-docker-database · high confidence

Support for OCI manifest formats in Docker image hashing

The Docker registry v2 client now accepts OCI manifest media types (application/vnd.oci.image.manifest.v1+json and application/vnd.oci.image.index.v1+json) in addition to standard Docker manifest types. If a request for a Docker manifest fails, the system automatically falls back to requesting an OCI manifest, allowing compatibility with registries and images that use the OCI specification.

dockerHashing/src/main/scala/cromwell/docker/registryv2 · high confidence

Token dispenser now enforces per-group limits and pauses distribution under high load

The job token dispenser in the engine now prevents any single workflow group from monopolizing resources by enforcing configurable 'hog limits' that cap the number of tokens a group can hold. Additionally, the system integrates with the Load Controller to automatically freeze token distribution when high load is detected, resuming only when load returns to normal. These changes ensure fairer resource sharing and improved stability during system stress.

engine/src/main/scala/cromwell/engine/workflow/tokens · high confidence

Updated GATK joint-discovery test inputs to use new reference bucket

The GATK joint-discovery integration test inputs have been updated to point to the new \gcp-public-data--broad-references\ bucket for all reference files (such as the hg38 FASTA, DBSNP, and 1000 Genomes resources). This change ensures the test workflows can successfully locate the required reference data in the new storage location.

centaur/src/main/resources/integrationTestCases/germline/joint-discovery-gatk · high confidence

Updated integration test cases for arrays and single-sample workflows

The integration test cases for the arrays and single-sample workflows have been updated to use the new public reference bucket \gcp-public-data--broad-references\ instead of the previous \broad-references\ path. Additionally, the arrays test configuration now sets preemptible tries to zero and uses a development Docker image, while the single-sample test options have been adjusted to include additional availability zones.

centaur/src/main/resources/integrationTestCases/green · high confidence

Upgrade test script for Cromwell 25 to 26 migration

The \scripts/test\_upgrade\ directory now contains a complete test harness (\test\_upgrade.sh\) and associated WDL workflow (\scatter\_files.wdl\) and input files to validate Cromwell upgrades. The script automates a rolling upgrade test by submitting jobs via the Cromwell API using the previous version (cromwell-25-31ae549-SNAP.jar), shutting down the service, restarting it with the new version (cromwell-26-88630db-SNAP.jar), and resubmitting jobs to ensure continuity and correctness across the version boundary.

_scripts/test\upgrade · high confidence

WOM model refactoring and WDL 1.1 support

The WOM (Workflow Object Model) core has been refactored to introduce a new \Callable\ abstraction for defining workflows and tasks, replacing previous structures. This change adds support for WDL 1.1 features, including new runtime attributes like \container\, \gpu\, \returnCodes\, and \localizationOptional\. The update also introduces a new \MemoryUnit\ enum and \MemorySize\ parser for handling memory specifications, adds YAML parsing limits in \reference.conf\ to mitigate Denial of Service risks, and includes utility functions for string normalization and UTF-8 cleaning.

wom · high confidence

Workflow deletion now handles intermediate files and call cache invalidation

The new DeleteWorkflowFilesActor introduces a structured workflow for cleaning up execution artifacts. It now identifies and deletes intermediate output files (distinguishing them from final outputs), updates metadata status during the process, and automatically invalidates associated call cache entries upon completion. The actor also gracefully handles scenarios where files are already missing (treating them as successful deletions) and ensures the workflow transitions to a terminal state even if no intermediate files require deletion.

engine/src/main/scala/cromwell/engine/workflow/lifecycle/deletion · high confidence

Workflow store refactoring with heartbeat monitoring and coordinated access

The workflow store implementation has been restructured to introduce a new \WorkflowStoreAccess\ abstraction that supports both uncoordinated and coordinated (serialized) database access to prevent deadlocks during concurrent operations like fetching, aborting, and writing heartbeats. A new \WorkflowStoreHeartbeatWriteActor\ now manages periodic heartbeat writes, enabling the system to detect stale heartbeats and automatically shut down if write failures persist beyond a configured duration. Additionally, an \AbortRequestScanningActor\ was added to periodically scan for and process pending abort requests, while in-memory and SQL store implementations were updated to support these new states and coordination mechanisms.

engine/src/main/scala/cromwell/engine/workflow/workflowstore · high confidence

Test coverage

Add SmartSeq2 single-sample integration test case; Added 'hello' external test case for workflow execution validation; Added CNV somatic test cases for GATK-launch; Added Centaur test cases for WDL 'after' keyword and workflow abort scenarios; Added CromIAM webservice test suite and mock clients; Added backend test infrastructure and validation specs; Added integration test for high-scale scattered workflows; Added integration tests for SmartSeq2, HaplotypeCaller, Mutect2, and CNV workflows; Added metadata parsing and comparison utilities for Centaur tests; Added test case for Cromwell restart with GCP Batch backend; Added test coverage for common collection utilities and exception handling; Added test coverage for common utility modules; Added test for graceful handling of unparseable Docker registry responses; Added test infrastructure and specs for core components; Added test infrastructure for backpressure report; Added test infrastructure for engine components; Added test logging configuration for WDL Draft 2; Added test resources for CromIAM; Added test resources for PAPI v1/v2 metadata comparison and forkjoin workflow validation; Added test resources for workflow metadata and configuration; Added test utilities and specifications for core Cromwell components; Added tests for AWS configuration parsing and validation; Added tests for Cromwell API client request construction and error handling; Added tests for Cromwell API model JSON formatters; Added tests for CromwellFileSystems configuration and validation; Added tests for DRS localizer downloaders and checksum handling; Added tests for DefaultPathBuilder and PathBuilderFactory; Added tests for Docker image hashing and mirroring logic; Added tests for GCS path building and requester-pays error handling; Added tests for GCS storage configuration and options; Added tests for GoogleConfiguration parsing and validation; Added tests for GroupMetricsActor and StandardValidatedRuntimeAttributesBuilder; Added tests for JobPaths and WorkflowPaths path resolution; Added tests for S3 storage configuration and client creation; Added tests for WDL draft-3 language factory caching behavior; Added tests for WDL draft2 language factory features; Added tests for WDL import resolution and version detection; Added tests for WomValueBuilder simpleton conversion logic; Added tests for call caching blacklist and file hashing actors; Added tests for call caching hash strategies and core I/O actor behavior; Added tests for common validation utilities and mock helpers; Added tests for core actor batching, robust client, and stream helpers; Added tests for label key and value validation; Added tests for retry and backoff logic; Added tests for standard Centaur test case execution and format definitions; Added unit tests for AWS ECR and ECR Public registry flows; Added unit tests for Centaur test infrastructure components; Added unit tests for CromwellClient workflow ID and collection resolution; Added unit tests for DRS Localizer command-line parsing and main execution logic; Added unit tests for DRS filesystem components; Added unit tests for GCP authentication modes in cloudSupport; Added unit tests for GCS batch IO commands; Added unit tests for S3 file system operations; Added unit tests for S3 path building and validation; Added unit tests for SamClient authorization and submission logic; Added unit tests for core logging components; Added unit tests for metadata comparison scripts; Added unit tests for the DRS Cloud NIO implementation; Added validation tests for runtime attributes and return codes; New test formula library for workflow execution and validation; New test utilities for assertions, timeouts, and repetition; Refactored Centaur test infrastructure with new exception handling and cloud-agnostic file checking; Refactored Centaur test workflow configuration and execution model.

Dependencies

Dependency updates for CromwellRefdiskManifestCreator and Java client codegen

The CromwellRefdiskManifestCreator Maven module now uses Jackson Databind 2.13.4.1, Log4j 2.17.1, and JUnit 4.13.1. The Java client codegen build (codegen\_java) targets Scala 2.13.9 and includes OkHttp 4.9.3, Gson 2.9.0, Commons Lang3 3.12.0, Swagger Annotations 1.6.5, and JUnit 5.8.2. Additionally, the Cromwell Monitor requirements.txt for the Google Pipelines backend now lists google-cloud-monitoring, psutil, and requests.

(dependencies) · high confidence

New sbt build infrastructure for CI, dependencies, and Docker publishing

The build system introduces dedicated Scala files to manage project configuration: ContinuousIntegration.scala adds a \renderCiResources\ task that uses the \broadinstitute/dsde-toolbox\ Docker image to render templates via Hashicorp Vault and validates that all sub-projects are correctly aggregated; Dependencies.scala centralizes library versions (including Akka 2.5.32, AWS SDK 2.29.20, Jackson 2.13.2, and Netty 4.2.16.Final); Merging.scala defines custom assembly merge strategies to handle conflicts in AWS/Google jars and metadata files; and Publishing.scala configures Docker image tagging (including \latest\ for releases) and entrypoints that support \JAVA\_OPTS\ and \--add-opens\ flags for Java 17 compatibility.

project · high confidence

Housekeeping

Repository initialization and configuration standardization

The repository has been initialized with standard configuration files to streamline development and deployment. This includes a \.dockerignore\ symlink to \.gitignore\, \.gitattributes\ for consistent text file handling, and a comprehensive \.gitignore\ covering build artifacts, IDE files, and local execution logs. Build and formatting are standardized via \.sbtopts\ (setting JVM memory limits), \.scalafmt.conf\ (Scala 2.13 formatting rules), and \.sdkmanrc\ (pinning Java 17). Documentation is configured for Read the Docs via \.readthedocs.yaml\, and the project now includes \AUTHORS\, \CITATION.md\, \CONTRIBUTING.md\, and \LICENSE-ASL-2.0\ files.

(repo-wide) · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

Score

  • CAI 51 → 61 (+10.1)
  • Rubric changed (rubric-2026.09.8 → rubric-2026.09.16) — scores are not directly comparable.

Lenses

  • Code Health 87 → 87 (-0.0)
  • Architecture 96 → 85 (-11.0)
  • Maturity 44 → 65 (+20.6)
  • Readiness 55 → 65 (+10.5)
  • Security 43 → 51 (+7.9)

Resolved (3)

  • Documentation: no installation or build instructions (docs/README.md)
  • Documentation: no usage examples (docs/README.md)
  • Off-boarding risk: anonymized user #1

New (6)

  • Documentation: no licence statement (docs/index.md)
  • Duplicated block (11 lines × 2) (scripts/metadata_comparison/metadata_comparison/lib/operations_digesters.py)
  • Edited copy of a member (19 corresponding lines) (scripts/metadata_comparison/metadata_comparison/lib/operations_digesters.py)
  • No ADRs found
  • Off-boarding risk: anonymized user #1
  • Projects may be oversized for their cohesion

Changes since last survey

  • 3 commits — 1 feature/other, 2 fixes

By area

  • (root) — 1 commit
  • docs/LanguageSupport.md — 1 commit
  • supportedBackends/google — 1 commit

Notable commits

  • fix: CTM-655 Fix CI issues with gcloud CLI (#7902)
  • fix: [no JIRA] Fix Markdown navigation in GCP Batch docs (#7903)
  • change: [No JIRA] Update supported WDL versions (#7901)

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

broadinstitute/cromwell was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 28 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit f1b30430fa32539e0b73932ebc160cd5335792c8 — the exact code this score is about.
  • Scored under rubric-2026.09.16 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-2d9048c36d26.