AbsaOSS/cobrix
59.2
Adequate · 20 September 2026
35.4k
lines of production code
Scala
with Java
1
measurement over time
What this system is
This system is a Spark connector library designed to read and write mainframe COBOL data files within distributed computing environments. It provides robust parsing capabilities for complex copybook structures, including hierarchical records, variable-length formats, and various binary encodings, while exposing a builder-based API for seamless integration with Spark DataFrames. The library supports both ingestion and generation of COBOL datasets, handling features such as GPG-encrypted file decryption, custom record header parsing, and performance-optimized indexing for large-scale data processing.
How it got here
2018 — Initial project scaffolding and core parser development
20 changes.
This period established the foundational architecture for the Cobrix project, including repository setup, CI/CD pipelines, and the initial ANTLR-based COBOL copybook parser. It introduced the core Spark integration components, such as the builder-based API and record readers, alongside comprehensive unit and integration tests to validate schema derivation and data extraction. The work also included building example applications and performance benchmarking infrastructure to demonstrate the connector's capabilities.
2019–2020 — Build modularization and test expansion
9 changes.
The project refactored its build infrastructure into modular sbt configurations to support cross-compilation and centralized dependency management, while isolating example applications into a separate module. Concurrently, extensive regression and unit tests were added to cover COBOL parsing, serialization, and Spark job execution, ensuring robust handling of complex data structures and edge cases.
2021–2026 — Test coverage expansion
4 changes.
This period focused on expanding test coverage for the Spark-COBOL project, particularly for the COBOL writer component. Comprehensive test suites were added to validate fixed and variable-length EBCDIC output, nested data structures, and specific encodings. Additionally, parameter validation logic for writing operations was rigorously tested to ensure correct handling of unsupported features and configuration constraints.
Features
Add COBOL copybook and source examples for data type coverage
The examples/example\_data directory now includes new COBOL copybook (.cpy) and source (.cob) files that demonstrate various data types and structures, including binary, BCD, and string formats, as well as redefines and occurs clauses. These files serve as reference data for parsing and encoding examples.
_examples/example\data · high confidence
Added example COBOL copybook with Apache 2.0 license
The example application now includes a sample COBOL copybook (example\_copybook.cpy) in the resources directory, defining a COMPANY-DETAILS structure with fields for company identification, static details, and contact information. This file is distributed under the Apache License, Version 2.0, ensuring proper licensing compliance for users of the example.
examples/spark-cobol-app/src/main/resources · high confidence
Added performance benchmark data and visualization scripts
The performance directory now includes raw benchmark data (CSV) and Gnuplot scripts (.plot) for three experiments: raw records processing, multi-segment narrow records, and multi-segment wide records. These files enable the generation of SVG charts visualizing processing time, throughput (in records per second and MB/s), and efficiency against the number of executors. Helper scripts (generate.cmd and generate.sh) are provided to automate the chart generation using these data sources.
performance · high confidence
Initial parser and reader architecture introduction
This change introduces the foundational components for the COBOL parser and record reader, including the ANTLR-based \ParserVisitor\, AST node definitions (\Primitive\, \Usage\), and expression evaluation infrastructure. It establishes the core data structures for handling COBOL copybooks, such as \RecordFormat\ (Fixed, Variable, ASCII), \RecordMetadata\, and \RawRecordContext\, along with configurable policies for comments, fillers, and metadata. This initial import sets the stage for parsing and extracting data from mainframe files.
repository · high confidence
Initial project scaffolding and repository setup
The repository is initialized with the core project structure, including the Apache 2.0 license, a Maven-based build configuration (POM), and a Jenkinsfile for CI/CD. Documentation is established via a comprehensive README.md detailing usage with Spark, linking coordinates, and examples. Build tooling is configured with sbt scripts for publishing and dependency management, alongside a .fossa.yml file for open-source license compliance scanning. Development environment hygiene is improved by updating .gitignore to exclude IDE-specific files (VS Code, Eclipse, IntelliJ) and OS artifacts.
(repo-wide) · high confidence
Introduce builder-based API and new reader classes for Spark integration
The spark-cobol module now provides a new builder-based API for configuring and loading COBOL data, including \SparkCobolProcessor\ for raw record processing and \SparkCobolBuilder\ for creating DataFrames from RDDs. Additionally, new reader classes such as \FixedLenNestedReader\, \FixedLenTextReader\, and \VarLenNestedReader\ have been added to handle fixed and variable length records, supporting both binary and text data sources. These changes enhance the flexibility and usability of the spark-cobol connector by offering more granular control over data processing and schema generation.
spark-cobol/src/main · high confidence
New standalone example application for Cobrix on Spark
The spark-cobol-app module now includes a standalone example application that demonstrates how to read mainframe data into Spark DataFrames. This includes a custom record header parser for non-standard RDW headers, a hierarchical record reader example showing segment joining, and a basic type variety example for fixed-length records.
examples/spark-cobol-app/src/main/scala · high confidence
Architecture
Build infrastructure refactored into modular sbt configuration files
The project's build configuration has been reorganized into dedicated Scala files within the project directory to improve maintainability and support cross-compilation. Dependencies are now centrally managed in Dependencies.scala, which defines specific versions for libraries such as Jackson (2.15.4), ANTLR (4.9.3), and BouncyCastle (1.84), along with a version matrix for Spark (2.4.8, 3.4.4, 3.5.7) mapped to Scala versions (2.11, 2.12, 2.13). Compiler options are standardized in ScalacOptions.scala, enforcing UTF-8 encoding and specific JVM targets. Additionally, the build now generates a cobrix\_build.properties resource containing the project version and build timestamp, and integrates custom JaCoCo plugins for code coverage reporting.
project · high confidence
Behavioural changes
COBOL copybook parser replaced with ANTLR-based implementation
The \cobol-parser\ module now uses a new ANTLR 4 grammar (with generated Java lexer and parser classes) to parse COBOL copybooks, replacing the previous parsing logic. This change introduces a dedicated \Logging\ trait for SLF4J integration and a custom JSON parser to handle configuration, providing more robust syntax error reporting and laying the groundwork for advanced parsing features.
cobol-parser · high confidence
Cobrix examples moved to a separate project
The example applications and test data generators have been relocated from the main production codebase into a new, standalone \examples/examples-collection\ project. This change isolates demonstration code—including Spark jobs for reading COBOL files, streaming examples, and various data generators—from the core library, simplifying the main build and clarifying the boundary between production artifacts and usage samples.
examples/examples-collection · high confidence
Test coverage
Added integration tests for non-terminals, custom RDW parsers, copybook merging, file headers, and hierarchical data; Added mock record extractors and AST transformer for testing custom parsing logic; Added regression tests for Cobrix source reading behaviors; Added test configuration for SLF4J logging; Added test data and expected outputs for copybook parsing scenarios; Added test fixtures for temporary file creation and binary/text comparison; Added test infrastructure and schema validation tests for Spark-Cobol; Added test resources for copybook and GPG key handling; Added test runners for Spark COBOL example applications; Added test stubs for CobolSchema and FixedLenReader; Added test suites for COBOL writer capabilities; Added test utilities and suites for custom code pages, LRU caching, and record header parsing; Added test utilities for running Spark jobs locally; Added tests for COBOL record serialization to JSON and XML; Added tests for COBOL writing parameter validation; Added tests for FileStreamer behavior; Added unit tests for ASCII text file parsing; Added unit tests for Spark-Cobol utility classes; Added unit tests for SparkCobolProcessor, schema generation, and row extraction; Added unit tests for index building and location balancing; Added unit tests for spark-cobol source components; Updated test logging configuration to suppress noisy output.
Dependencies
Cobrix build system upgraded to Scala 2.12/2.13 and Spark 3.5 with ANTLR shading
The project's build configuration (build.sbt and POM files) has been updated to target Scala 2.12.21 and 2.13.18, and Spark 3.5.7. The \cobol-parser\ module now includes ANTLR 4.9.3 and applies shading rules to rename the ANTLR runtime package to \za.co.absa.cobrix.cobol.parser.shaded\, preventing binary incompatibilities with Spark's bundled ANTLR version. Additionally, the \spark-cobol\ module now includes Bouncy Castle for GPG support and marks Spark dependencies as provided.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Baseline
- First survey — no prior run to compare against. CAI 59.
Lenses
- Code Health 86
- Architecture 97
- Maturity 63
- Readiness 46
- Security 70
Changes since last survey
- 300 commits — 240 feature/other, 60 fixes
By area
- cobol-parser/src — 126 commits
- (root) — 79 commits
- spark-cobol/src — 62 commits
- (repo) — 12 commits
- .github/workflows — 9 commits
- examples/spark-cobol-app — 3 commits
- examples/spark-cobol-s3-standalone — 2 commits
- cobol-parser/pom.xml — 1 commit
- data/test40_data — 1 commit
- data/test40_data_ascii — 1 commit
- data/test41_expected — 1 commit
- examples/spark-cobol-s3 — 1 commit
- project/Dependencies.scala — 1 commit
- project/build.properties — 1 commit
Notable commits
- fix: #25 Fix a PR suggestion - thanks @coderabbitai!
- fix: #415 Fix PR suggestions (Thanks @coderabbitai), add new unit test cases.
- fix: #415 Fix valuable nitpick PR suggestions from @coderabbitai.
- fix: #415 Implement a basic fixed record length record combiner.
- fix: #723 Fix PR nitpicks.
- fix: #723 Fix numerous PR suggestions about corrupt field generation.
- fix: #723 Make corrupted fields generation slightly more efficient (CPU and memory wise), fix PIC X USAGE COMP cases.
- fix: #727 Fix parsing of '88 VALUE1 VALUE +999.9999.' statement that threw 'SytnaxErrorException'.
- fix: #744 Add a unit test for the fixed record length extractor with record length mapping with default record length.
- fix: #757 Fix BouncyCastle provider resolution when shaded, and close the raw stream if GPG decryption fails.
- fix: #757 Fix Hadoop FSInputStream leak for GPG-encrypted files. Ensure the raw underlying stream is closed.
- fix: #757 Fix another potential leek of Hadoop FSInputStream in BufferedFsDataInputStream.
- fix: #759 Fix order-dependent unit test
- fix: #763 Fix PR suggestions.
- fix: #769 Fix PR nitpick suggestions from CoPilot PR review.
- fix: #769 Fix the processor not processing the last record, add unit tests.
- fix: #776 Fix PR suggestions.
- fix: #777 Fix spark-cobol writer working with file-based committers, and other PR fixes (Thanks @coderabbitai).
- fix: #780 Fix PR suggestions. Fix RDW+BDW encoders from Reader parameters converters.
- fix: #780 Move the fixed record length record extractor to the method where other record extractors are created.
- …and 280 more
Architecture
- 0 containers · 1 bounded contexts · 0 dependency edges (baseline)
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
AbsaOSS/cobrix was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 20 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 339ddbad9d89b3430731ac2f1a27e5dd9c0605c6 — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-b51f968c9b10.