projectglow/glow
69.1
Adequate · 20 September 2026
21.6k
lines of production code
Python
primary language
1
measurement over time
What this system is
This system is a genomic data processing library built on Apache Spark, providing tools for variant file manipulation and statistical genetics analysis. It supports reading and writing various genomic formats like VCF, BGEN, and PLINK, while offering transformers for normalizing and lifting over variants. The system also includes Python APIs for performing genome-wide association studies (GWAS) and weighted genomic regression using linear and logistic models.
How it got here
2019 — Initial scaffolding and genomic data support
18 changes.
The project established its foundational structure, including build infrastructure, licensing, and documentation, while introducing a YAML-driven code generation system for SQL functions. Significant development focused on core genomic capabilities, adding support for BGEN and PLINK file formats, service-based variant transformers, and a new Python API with interval join features.
2020–2021 — GWAS regression and WGR module development
11 changes.
This period focused on introducing the Weighted Genomic Regression (WGR) module for ridge and logistic regression, alongside significant enhancements to the GWAS module with pandas-based linear and logistic regression functions featuring approximate Firth correction. The work was heavily supported by the creation of comprehensive test suites and extensive test fixtures to validate these new statistical modeling capabilities and ensure accuracy against established baselines.
Features
Added HLS event logging capability
A new logging module has been introduced to record High-Level Service (HLS) events. This feature allows users to log specific events by providing a tag and optional arguments, which are then forwarded to the underlying JVM-based event recorder via PySpark.
python/glow/logging · high confidence
Added Spark submit script
A new shell script named spark-submit has been added to the bin directory. This script launches the Apache Spark submitter class using the configured classpath and Java options, providing a simple entry point for submitting Spark applications.
bin · high confidence
Added VCF processing utility scripts
New shell and Python scripts have been added to the test-data/vcf/scripts directory to support variant data processing workflows. These include gwas.sh for simulating a GWAS interface from VCF to CSV, gwas-region.py for filtering association results based on gene regions, and helper utilities (prepend-chr.sh, remove-info.sh, remove-rows.sh) for manipulating VCF format fields, along with a group\_file.txt for defining gene-region mappings.
test-data/vcf/scripts · high confidence
Initial Python package structure and build tooling
This change introduces the foundational structure for the Python package, including \setup.py\ for distribution, \environment.yml\ for development dependencies (such as PySpark 3.5.1, pandas, and scikit-learn), and a \render\_template.py\ script that generates language-specific client code from YAML definitions. It also adds configuration files like \.style.yapf\ for code formatting, \MANIFEST.in\ for packaging assets, and tests for the template rendering logic.
python · high confidence
Initial project scaffolding and configuration
The repository is initialized with foundational configuration and documentation files. This includes the Apache 2.0 license, a Contributor Covenant Code of Conduct, and contributor guidelines (CONTRIBUTING.md, GIT-PROCESS.md). Build and development tooling is established via \.scalafmt.conf\ (v2.7.5), \scalastyle-config.xml\, and a \functions.yml\ manifest that defines the Spark SQL functions for code generation. Testing infrastructure is added through \pytest.ini\, \conftest.py\ (handling Spark session fixtures and version-specific test skipping), and \codecov.yml\. Documentation hosting is configured via \.readthedocs.yml\, and release workflows are documented in \RELEASE.md\.
(repo-wide) · high confidence
Introduce Glow Python package with interval overlap joins and NumPy support
The \python/glow\ package is now available, providing a Python API for genomic data processing on Spark. This release adds \left\_overlap\_join\ and \left\_semi\_overlap\_join\ functions in \glow.sql.functions\ to perform interval-based joins optimized with range-join logic, and introduces \glow.register\ to initialize SQL extensions and Py4J converters. It also enables direct use of 1D and 2D NumPy double arrays in Spark expressions via automatic conversion to Java arrays and \DenseMatrix\ literals.
python/glow · high confidence
Introduce Weighted Genomic Regression (WGR) module for ridge and logistic regression
The \python/glow/wgr\ package is added, providing a new set of tools for genomic analysis including \RidgeReduction\ for feature space reduction, \RidgeRegression\ for quantitative trait modeling, and \LogisticRidgeRegression\ for binary classification. These classes leverage Spark DataFrames and Pandas for efficient computation of ridge models with cross-validation, supporting covariates and intercepts. The module also includes utility functions in \wgr\_functions.py\ for handling sample IDs, blocking variants, and reshaping labels for GWAS workflows.
python/glow/wgr · high confidence
Introduce service-based transformer registration and BGEN file support
The core module now registers genomic data transformers (LiftOverVariants, BlockVariantsAndSamples, NormalizeVariants, SplitMultiallelics, Pipe, and CleanupPipe) via Java ServiceLoader, allowing them to be discovered and applied dynamically through the Glow API. Additionally, this change adds complete support for reading and writing BGEN files, including schema inference, genotype reading for uncompressed and compressed (zlib, zstd) data, and record writing with configurable probability bit depths and ploidy handling.
core/src/main · high confidence
Introduces pandas-based linear and logistic regression with approximate Firth correction
The GWAS module now provides new \linear\_regression\ and \logistic\_regression\ functions that perform analysis using pandas DataFrames and Spark Pandas UDFs, replacing the previous implementation. These functions support LOCO (Leave-One-Chromosome-Out) offsets for distributed processing and include an approximate Firth bias-reduced correction for logistic regression to handle rare variant associations more accurately.
python/glow/gwas · high confidence
Behavioural changes
1 commit (0 fixes) modifying test-data/tabix-test-vcf
A change to existing behaviour in test-data/tabix-test-vcf — 1 commit, 9 files.
test-data/tabix-test-vcf · medium confidence · unverified
Automated generation of Scala SQL functions from YAML definitions
The core module now generates its Scala SQL function definitions (\core/functions.scala.TEMPLATE\) from a \functions.yml\ configuration file instead of maintaining them as static source code. This change introduces a template-based approach where function signatures, documentation, and argument handling (including optional arguments) are derived directly from the YAML definitions, simplifying the addition and maintenance of new SQL functions.
core · high confidence
Test coverage
Add PLINK test data fixtures; Added GFF3 test fixtures for schema inference and embedded FASTA; Added Picard-compatible liftover test fixtures; Added REGENIE test data for binary phenotype analysis; Added VCF test fixtures for edge cases and annotation formats; Added comprehensive test coverage for GWAS regression functions; Added liftover test data and documentation; Added test coverage for Glow Python integration; Added test data for BEDTools intersect operation; Added test data for binary logistic regression and variant sample block generation; Added test fixtures for variant splitter and normalizer; Added test infrastructure and suites for Glow data sources; Added thresholded VCF test data for BGEN validation; Expanded test data for VCF and CSV parsing; Initial test suite for Glow WGR module.
Dependencies
Build infrastructure and dependency updates
The build system has been upgraded to SBT 1.10.4, and several SBT plugins have been updated or added, including sbt-assembly (2.3.0), sbt-pgp (2.3.0), sbt-sonatype (3.12.2), sbt-scalafmt (2.5.2), sbt-scoverage (2.2.2), and sbt-header (5.10.0). The ScalaTest dependency has been upgraded from version 3.0.5 to 3.2.18, and a library dependency scheme for scala-xml has been configured to allow always-compatible updates.
project · high confidence
Upgrade Spark support to 3.5.1 and 4.1.0 with Scala 2.12/2.13
The build configuration has been updated to support Apache Spark 3.5.1 and Spark 4.1.0, replacing the previous Spark 2.4.1 baseline. This change introduces a shim-based architecture to handle API differences across Spark versions and updates the default Scala version to 2.12.19 (with 2.13.15 support for Spark 4+). Users can now build and run Glow against newer Spark releases by setting the SPARK\_VERSION environment variable, ensuring compatibility with modern Spark ecosystems.
(dependencies) · high confidence
Housekeeping
Added documentation for VCF merge test data
A README file has been added to the test-data/vcf-merge directory to document the source of the test VCF files, specifying that they contain the first 1000 sites of chromosome 22 from 1000 Genomes samples HG000096 and HG00097.
test-data/vcf-merge · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Baseline
- First survey — no prior run to compare against. CAI 69.
Lenses
- Code Health 100
- Architecture 100
- Maturity 68
- Readiness 71
- Security 62
Changes since last survey
- 300 commits — 260 feature/other, 40 fixes
By area
- docs/source — 73 commits
- (root) — 71 commits
- core/src — 38 commits
- python/glow — 32 commits
- .github/workflows — 28 commits
- docker/databricks — 13 commits
- project/plugins.sbt — 13 commits
- .circleci/config.yml — 6 commits
- docs/dev — 5 commits
- project/build.properties — 4 commits
- (repo) — 2 commits
- python/setup.py — 2 commits
- test-data/regenie — 2 commits
- .github/PULL_REQUEST_TEMPLATE — 1 commit
- docker/README.md — 1 commit
- docker/open-source-glow — 1 commit
- docs/README.md — 1 commit
- python/LICENSE.txt — 1 commit
- python/environment.yml — 1 commit
- python/version.py — 1 commit
Notable commits
- fix: Accept different datatypes, values column expression for linear regression (#312)
- fix: Add docs for pandas based linear regression; misc doc improvements (#314)
- fix: Add explicit linkcheck timeout; ignore flaky 1kg link; fix setup.py syntax error (#573)
- fix: Add logistic regression (#245)
- fix: Add offset option to logistic regression doc (#299)
- fix: Documentation fixes for DBFS API (#516)
- fix: Extending sample masking functionality in gwas linear regression (#416)
- fix: Fix BgenRowConverterSuite with AQE enabled (#292)
- fix: Fix Databricks SDK compatibility in build script (#785)
- fix: Fix Glow version in doc (#364)
- fix: Fix Infinity/NaN parsing to allow full set of values from VCF specification (#519)
- fix: Fix IntelliJ import (#223)
- fix: Fix VEP parsing failures stemming from indels (#402)
- fix: Fix bgz registration (#359)
- fix: Fix broken validation (#343)
- fix: Fix critical security vulnerability and CI workflow issues (#788)
- fix: Fix default shell in release job (#631)
- fix: Fix for Databricks ES-257648 (#488)
- fix: Fix linkcheck errors (#253)
- fix: Fix name for production release action (#640)
- …and 280 more
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
projectglow/glow was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 20 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 0539c76b95c7cde241f2d468acfad4a0ece77b1f — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-b51f968c9b10.