starlake-ai/starflow
48.1
Weak · 20 September 2026
64.1k
lines of production code
Scala
primary language
1
measurement over time
What this system is
Starflow is a code-first data engineering platform that automates the ingestion, transformation, and quality validation of data across diverse sources and destinations. It generates executable SQL and orchestration workflows (Airflow, Dagster) from declarative YAML metadata, supporting native execution on engines like BigQuery, Snowflake, and DuckDB alongside Spark. The system provides comprehensive tooling for schema inference, data lineage, access control, and semantic model export, enabling end-to-end pipeline management from raw files to analytical tables.
How it got here
2018–2021 — data quality and ingestion expansion
52 changes.
This period focused on expanding the platform's data ingestion and quality capabilities by introducing row-level validation, data expectations, and native audit logging. The team also modernized the underlying architecture by upgrading to Spark 4 and Java 17, refactoring the schema and workflow modules, and adding support for new sinks like DuckDB and BigQuery.
2022–2024 — multi-engine integration and native execution
80 changes.
This period focused on expanding the platform's support for diverse data warehouses and databases, introducing native execution paths for BigQuery, Snowflake, PostgreSQL, Redshift, and DuckDB to bypass Spark overhead. It also established a unified framework for adaptive write strategies, data lineage, and project scaffolding, while significantly enhancing test coverage and distribution tooling for local development.
2025–2026 — SQL generation and multi-engine support
48 changes.
The project introduced a new SQL statement generation model and standardized load-strategy templates to support a wider range of database engines, including Snowflake, MySQL, SQL Server, and Trino. Significant enhancements included the addition of data quality expectation macros, REST API extraction capabilities, and semantic model export features for external BI tools. The period also saw the expansion of sample projects and comprehensive test coverage for native loaders and dependency management.
Features
Add BigQuery schema conversion utilities
The project now includes repackaged BigQuery schema conversion utilities (BigQuerySchemaConverters and SupportedCustomDataType) within the ai.starlake.utils.repackaged package. These classes handle the mapping between BigQuery and Spark data types, including support for custom Spark ML data types like vectors and matrices, and address specific conversion issues such as handling pseudo-columns and nested structures.
src/main/java/ai/starlake/utils · high confidence
Add DuckDB data source integration via StarLake provider
Users can now read from and write to DuckDB databases using Spark SQL by specifying the 'starlake-duckdb' data source. This new DuckDBRelationProvider extends the existing JDBC integration, allowing standard DataFrame operations to interact with DuckDB tables, including support for overwrite, append, error-if-exists, and ignore save modes.
src/main/scala/org/apache/spark/sql/execution/datasources/duckdb · high confidence
Add DuckDB sample with metadata, data, and Airflow DAG templates
This change introduces a complete sample configuration for DuckDB within the \samples/duckdb\ directory. It includes metadata definitions for loading HR and sales data (JSON/CSV), transforming KPIs, and validating data quality via expectations. The sample provides specific type mappings for DuckDB (e.g., \VARCHAR\, \INTEGER\, \TIMESTAMPTZ\) and configures Airflow DAG generation templates to execute load tasks using the native DuckDB JDBC driver. Sample data files and test fixtures are also added to demonstrate the end-to-end pipeline.
samples/duckdb · high confidence
Add KPI transform samples for order and product analysis
New SQL transform files and configuration have been added to the starbake sample to demonstrate Key Performance Indicator (KPI) calculations. This includes a configuration file setting the write strategy to OVERWRITE, and SQL scripts that compute metrics such as total order value, product profit, total units sold, and total revenue by joining order lines, products, and orders tables.
samples/starbake/metadata/transform · high confidence
Add KPI transformation sample for Postgres
A new sample configuration and SQL transformation have been added to the Postgres metadata transform area. The \\_config.sl.yml\ file sets the default write behavior to OVERWRITE for the KPI transform, and the \byseller.sql\ script defines a transformation that aggregates sales amounts by seller email, joining the \hr.sellers\ and \sales.orders\ tables.
samples/postgres/metadata/transform · high confidence
Add OpenAPI schema extraction support
Users can now extract data schemas from OpenAPI specifications. This change introduces a new OpenAPI schema extractor that parses OpenAPI files (via Swagger V3 parser) and transforms API definitions into Starflow-compatible domain and table structures. The feature supports configuration via YAML/JSON to define domains, routes, path patterns, HTTP methods (GET/POST), field exclusions, and schema explosion strategies (ALL, ARRAY, OBJECT). It handles JSON schema type mapping and generates normalized table names based on API paths and optional rename rules.
src/main/scala/ai/starlake/extract/impl/openapi · high confidence
Add SQL Server and Trino write strategy templates
New Jinja2 templates have been added for SQL Server and Trino to support data write operations. These templates cover append, create, delete-then-insert, overwrite, overwrite-by-partition, SCD2 (Slowly Changing Dimension Type 2), upsert-by-key, upsert-by-key-and-timestamp, and view creation strategies, enabling users to configure these specific database engines for their data pipelines.
src/main/resources/templates/write-strategies/sqlserver, src/main/resources/templates/write-strategies/trino · high confidence
Add Snowflake sample with 1.0.0 versioning and environment setup
The Snowflake sample directory has been updated to version 1.0.0-SNAPSHOT, introducing a new \.snowflake-env\ file for configuring connection details (account, user, password, warehouse, database) and an \env.sh\ script that validates the \SL\_HOME\ path and required environment variables. The sample now includes a structured set of shell scripts (\0.init.sh\ through \3.transform.sh\) to guide users through importing data, watching for changes, and running a specific transformation (\kpi.byseller\), alongside metadata configuration files for the transform step.
samples/snowflake · high confidence
Add Summarize command for table statistical summaries
Users can now run the 'summarize' command to generate a statistical summary for a specific domain and table, providing key data profiling insights such as column names, types, and descriptions. This new capability is implemented via the SummarizeCmd and SummarizeConfig classes, which accept domain and table arguments and output the results in a specified report format.
src/main/scala/ai/starlake/job/tools · high confidence
Add parquet2csv command to convert Parquet files to CSV
Users can now convert Parquet files to CSV format using the new \starlake parquet2csv\ command. This feature allows specifying input and output directories, filtering by domain or schema, configuring output partitions, setting write modes (e.g., overwrite, append), and passing custom Spark options. It also supports deleting source Parquet files after conversion.
src/main/scala/ai/starlake/job/convert · high confidence
Add sample data for HR and sales domains
Added sample data files to support HR and sales scenarios, including location and seller records in JSON format under the HR directory, and customer, order, product, and category data in CSV/PSV formats under the sales directory.
samples/sample-data · high confidence
Add support for loading data into Elasticsearch
Users can now load datasets in Parquet, JSON, or JSON-array formats into Elasticsearch indices using the new \esload\ (and \index\) command. This feature allows specifying custom mappings, document IDs, timestamp-based index suffixes, and Spark-Elasticsearch configuration options, while automatically registering and updating Elasticsearch index templates via HTTP API before writing data.
src/main/scala/ai/starlake/job/sink/es · high confidence
Added BigQuery and Snowflake sample KPI transformation definitions
New sample files have been added to the BigQuery and Snowflake Spark metadata directories to demonstrate a KPI transformation. This includes a configuration file specifying an OVERWRITE write strategy and a SQL script that calculates the sum of sales amounts grouped by seller email, joining seller and order data.
samples/bigquery/metadata/transform, samples/snowflake-spark/metadata/transform · high confidence
Added DDL generation and data extraction templates for multiple database platforms
New Jinja2 templates have been added to support DDL operations (create, alter, drop) and data extraction scripts for BigQuery, PostgreSQL, Snowflake, Synapse, and MySQL. The BigQuery and Synapse templates now include logic for adding, dropping, and modifying columns with metadata options, while PostgreSQL and MySQL extraction scripts provide full and delta export capabilities with audit status tracking.
(repo-wide) · high confidence
Added Jinja2 macro library for Snowflake data expectations
A new Jinja2 template file (default.j2) has been added to the Snowflake metadata samples, providing a set of reusable SQL macros for defining data quality expectations. These macros, including is\_col\_value\_not\_unique, is\_row\_count\_to\_be\_between, col\_value\_count\_greater\_than, count\_by\_value, and column\_occurs, allow users to parameterize and standardize common validation logic such as uniqueness checks, row count ranges, and value frequency counts within Snowflake tables.
samples/snowflake/metadata/expectations · high confidence
Added MySQL write strategy templates
New Jinja2 templates have been added to the MySQL write-strategies directory, enabling support for various data loading and transformation operations. These include append, create, delete-then-insert, overwrite, overwrite-by-partition, SCD2 (Slowly Changing Dimension Type 2), upsert-by-key, upsert-by-key-and-timestamp, and view creation strategies, allowing users to configure how data is written to MySQL tables.
src/main/resources/templates/write-strategies/mysql · high confidence
Added Postgres sample with bootstrap tutorial
A new sample located in samples/postgres provides a complete bootstrap tutorial for integrating Starlake with a PostgreSQL database. This includes shell scripts to initialize data, import it, load it via the Starlake CLI, and run a specific transformation (kpi.byseller), along with environment configuration files that set up the necessary connection parameters and Starlake variables for a local Postgres instance.
samples/postgres · high confidence
Added Snowflake-Spark sample with bootstrap scripts and configuration
A new sample integration for connecting Spark with Snowflake has been added to the repository. This includes a set of executable shell scripts (0.init.sh through 3.transform.sh) that guide users through initializing data, importing it via Starlake, watching for changes, and running a specific transformation (kpi.byseller). The sample also provides environment configuration files (env.sh and snowflake-env) that set up necessary variables like SL\_HOME, Snowflake connection details, and feature flags for metrics and assertions, allowing users to quickly bootstrap and run the sample workflow.
samples/snowflake-spark · high confidence
Added StarBake sample project for data lineage demonstration
The \samples/lineage\ area now includes a complete, self-contained sample project named 'StarBake' designed to demonstrate data ingestion, transformation, and analytics capabilities using Starlake Starflow. This sample provides fictional e-commerce bakery data (customers, orders, products) in various formats (CSV, JSON, XML) and defines a transformation pipeline that generates analytical tables such as CustomerLifetimeValue, ProductPerformance, and KPIs. The entry includes the sample's documentation (README), raw dataset archives, lineage diagrams defining table relationships and column mappings, and the necessary metadata configuration files (application.sl.yml, DAG definitions) to run the pipeline, offering users a practical reference for implementing similar data workflows.
samples/lineage · high confidence
Added default SQL template macros for data quality expectations
A new Jinja2 template file (default.j2) has been introduced in the samples/starbake/metadata/expectations directory, providing reusable SQL macros for common data validation checks. These macros enable users to define expectations such as verifying column uniqueness, checking row count ranges, counting values greater than a threshold, filtering by specific value patterns, and ensuring a column occurs a specific number of times.
samples/starbake/metadata/expectations · high confidence
Added reference configuration files for Azure, GCP, HDFS, and local filesystem deployments
New reference configuration files have been added to the \src/main/resources-other\ directory to support distinct deployment environments. These include \azure/reference.conf\ and \fs/reference.conf\ for Azure Blob Storage and local filesystems respectively, \gcp/reference.conf\ and \gcp/core-site.xml\ for Google Cloud Storage, and \hdfs/reference.conf\ and \hdfs/core-site.xml\ for Hadoop Distributed File System. Each configuration set defines environment-specific defaults for datasets, metadata, temporary directories, file systems, and Spark settings, allowing users to select the appropriate configuration profile for their infrastructure.
src/main/resources-other · high confidence
Added sample Starbake analytics and KPI transform definitions
The sample project template now includes metadata and SQL definitions for two new analytics domains: 'starbake\_analytics' and 'starbake\_kpis'. The analytics domain provides transforms for 'customer\_purchase\_history' (tracking order counts, spending, and dates) and 'order\_items\_analysis' (aggregating items and values per order). The KPIs domain adds an 'overall\_kpis' transform that calculates high-level metrics such as total customers, revenue, average order value, and daily rates by joining data from the analytics transforms. These additions allow users to immediately run pre-configured analytical queries on the sample data.
src/main/resources/templates/bootstrap/samples/sample-project/metadata/transform · high confidence
Added sample data and validation templates for BigQuery, Postgres, and Snowflake samples
This change adds sample data files and Jinja2 metadata expectation templates to the BigQuery, Postgres, and Snowflake sample directories. The sample data includes HR location and seller records in JSON, customer records in pipe-separated values, and sales orders in CSV, providing realistic test datasets for these integrations. Additionally, default.j2 templates are introduced for each target, defining reusable SQL macros for data quality checks such as uniqueness, row count ranges, and value frequency, enabling users to validate their data loads against these expectations.
(repo-wide) · high confidence
Added sample data files for Snowflake tutorial
New sample data files have been added to the \samples/snowflake/sample-data\ directory to support the Snowflake integration tutorial. This includes JSON files for HR data (\locations\, \sellers\), a pipe-separated values file for customer data (\customers\), a CSV file for order data (\orders\), and a shell script (\gcloud beta run jobs execute starlake-am\) demonstrating how to load sales data into BigQuery using Starlake. These files provide the necessary test data and execution examples for users following the tutorial.
samples/snowflake/sample-data · high confidence
Added sample data files for the Starbake bakery scenario
New sample data files have been added to the \samples/starbake/sample-data\ directory to support the Starbake demo. This includes \products.xml\ defining 12 bakery items (cakes, cookies, breads, and pies) with prices and costs, \orders\_202403011414.json\ containing 32 order records with customer IDs and statuses, and \order-lines\_202403011415.csv\ providing 100 line-item records linking orders to products with quantities and sale prices.
samples/starbake/sample-data · high confidence
Added sample dataset files for the Starbake project template
The bootstrap sample-project template now includes realistic sample data files for the 'Starbake' scenario. Specifically, it adds a customers.csv file containing 10 customer records, a products.json file listing 20 bakery products with pricing and category details, and an orders.json file with 20 sample order records linking customers to products. These files provide immediate, valid data for users to test the sample project's data ingestion and processing capabilities.
src/main/resources/templates/bootstrap/samples/sample-project/datasets · high confidence
Added utility to split Starlake JSON schemas
A new Python utility script has been added to the utils directory that extracts specific JSON schema definitions (such as TypeV1, DagGenerationConfigV1, and DomainV1) from the main starlake.json file and saves them as separate, smaller JSON files. This change is designed to reduce token consumption by breaking down large schema files into more manageable, individual components for downstream processing.
utils · high confidence
Initial DuckDB write-strategy templates
Added a complete set of Jinja2 templates for DuckDB write strategies, including append, create, delete-then-insert, overwrite, overwrite-by-partition, SCD2, upsert-by-key, upsert-by-key-and-timestamp, and view creation. These templates define the SQL logic for inserting, updating, and managing data in DuckDB tables and materialized views.
src/main/resources/templates/write-strategies/duckdb · high confidence
Introduce BigQuery support in Kafka data pipelines
The Kafka job component now supports BigQuery as both a source and a sink format. Users can offload data from Kafka topics to BigQuery tables or streams, and ingest data from BigQuery tables or SQL queries into Kafka topics. This change adds specific configuration requirements for BigQuery, such as mandatory \table\ and \temporaryGcsBucket\ options for writes, and introduces validation to ensure BigQuery is not used as a streaming source.
src/main/scala/ai/starlake/job/sink/kafka · high confidence
Introduce DataFrame transformation pipeline for sink jobs
A new \DataFrameTransform\ trait and utility object have been added to the sink job module, enabling data transformation logic to be applied to DataFrames before they are written to a sink. This introduces a pluggable mechanism where custom transformers (such as the included \IdentityDataFrameTransformer\ and \HeaderDataFrameTransformer\) can modify the data structure or content, with the \HeaderDataFrameTransformer\ specifically demonstrating how to inject structured headers using Kafka Avro serialization.
src/main/scala/ai/starlake/job/sink · high confidence
Introduce KafkaClient utility for topic management and offset handling
Added a new KafkaClient utility class that provides centralized management for Kafka topics and consumer offsets. This component supports creating and deleting topics, retrieving end offsets, and saving/retrieving offsets in either streaming (Kafka) or file-based modes, enabling more robust control over Kafka data flow within the application.
src/main/scala/ai/starlake/utils/kafka · high confidence
Introduce Quack command-line interface for managing DuckDB query servers
Users can now manage Quack DuckDB query servers directly from the command line using the new \starlake quack\ command. This addition introduces four key actions: \serve\ to run the server in the foreground, \start\ to detach it as a background daemon, \stop\ to terminate a specific instance, and \list\/\stop-all\ to manage running servers. The implementation includes a new \QuackServer\ class that handles the in-JVM serve loop with shutdown hooks, a \QuackState\ module for persisting server metadata (PID, bind address, port) to local state files, and a \QuackConfig\ model for parsing CLI arguments. The daemon mode (\start\) automatically handles process spawning, log rotation, and state tracking, while the foreground mode (\serve\) blocks until termination.
src/main/scala/ai/starlake/job/quack · high confidence
Introduce SQL type mappings and JDBC dialects for BigQuery, Snowflake, DuckDB, PostgreSQL, Redshift, and MariaDB
The SQL module now includes explicit type mapping tables for BigQuery, Snowflake, DuckDB, PostgreSQL, and Redshift, ensuring that data types from these sources are correctly translated to Spark types during ingestion. Additionally, custom JDBC dialects have been registered for Snowflake, DuckDB, and MariaDB to handle specific driver behaviors, such as timestamp timezones, huge integer support, and bit-vector conversions, improving compatibility and data accuracy when connecting to these databases.
src/main/scala/ai/starlake/sql · high confidence
Introduce data freshness monitoring command
Added a new \starlake freshness\ command that checks data freshness across tables and datasets by querying an audit table for the latest modification timestamps. The command supports configurable connections and write modes, handles epoch seconds and milliseconds in timestamp data, and outputs a JSON report indicating the status (INFO, WARN, or ERROR) for each monitored table, exiting with an error code if any errors are found.
src/main/scala/ai/starlake/extract/freshness · high confidence
Introduce dedicated schema inference command and configuration
The \infer-schema\ CLI command is now implemented via new \InferSchemaCmd\, \InferSchemaConfig\, and \InferSchemaJob\ classes in the \ai.starlake.job.infer\ package. This change consolidates schema inference logic into a dedicated command that supports options for specifying input paths, output directories, file formats (CSV, JSON, XML, Parquet), encoding, and XML row tags. It also introduces a \--variant\ flag to infer schemas as single variant attributes and a \--from-json-schema\ option to infer schemas from JSON Schema definitions. The configuration class handles automatic table name extraction from file names and determines write modes (append vs overwrite) based on naming patterns.
src/main/scala/ai/starlake/job/infer · high confidence
Introduce file listing strategies for load jobs
Added a new \LoadStrategy\ trait and two concrete implementations (\IngestionNameStrategy\ and \IngestionTimeStrategy\) to the load job module. These components allow the system to list input files for ingestion jobs either sorted by name or by time, providing a structured way to define how files are discovered during the load process.
src/main/scala/ai/starlake/job/load · high confidence
Introduce local single-user HTTP server with Caffeine caching
The \src/main/scala/ai/starlake/serve\ package now provides a new local HTTP server (\MainServerCmd\) that exposes Starflow CLI commands via REST API endpoints (mapped to \/api/v1/cli\). This server uses a new \CaffeineSettingsManager\ to cache settings, schema handlers, storage handlers, and Snowflake tokens, improving performance by avoiding repeated initialization. The server is designed for local, single-user development and testing, allowing users to run Starflow commands remotely via HTTP requests.
src/main/scala/ai/starlake/serve · high confidence
Introduce native BigQuery ingestion and branching capabilities
The BigQuery sink now supports a native execution path alongside the existing Spark-based approach. A new BigQueryNativeJob class allows direct data loading from GCS or local files using the BigQuery Storage Load API, including support for access tokens and local file uploads. Additionally, a BigQueryBranchHandler enables data branching for transforms by cloning source and target tables into a separate dataset and rewriting SQL references, allowing safe experimentation without affecting production data.
src/main/scala/ai/starlake/job/sink/bigquery · high confidence
Introduce native audit logging and autoload command for ingestion
The ingestion module now includes a dedicated \AuditLog\ model and \AuditTaskBuilder\ to natively record load and transform outcomes (including job ID, domain, schema, accepted/rejected counts, and duration) into an audit table, with robust escaping of SQL literals to prevent injection or parsing errors. Additionally, a new \autoload\ command has been added to automatically infer schemas from incoming files and load them into the data warehouse in a single step, streamlining the initial data ingestion workflow.
src/main/scala/ai/starlake/job/ingest · high confidence
Introduce native loaders for BigQuery, DuckDB, and Snowflake
Added new native loader implementations for BigQuery, DuckDB, and Snowflake that bypass Spark for data ingestion, supporting features like two-step loads for complex schemas, POSITION format handling, and native reject capture. The change includes infrastructure for capturing rejected lines (RejectCapture, NativeRejectedSink) and generating replay files for failed loads, enabling direct database operations for improved performance and reduced resource usage.
src/main/scala/ai/starlake/job/ingest/loaders · high confidence
New 'any-engine' sample with multi-environment orchestration and Airflow/Dagster DAG generation
The \samples/any-engine\ directory now provides a complete, runnable sample demonstrating data ingestion, transformation, and lineage across multiple execution environments (LOCAL, HDFS, SNOWFLAKE, BigQuery, PostgreSQL, Redshift, and DuckDB). It includes environment-specific shell scripts (e.g., \0.init.sh\, \1.import.sh\, \2.load.sh\) that configure Spark and StarLake settings, and a rich metadata structure defining tables, connections, and expectations. A key addition is the \metadata/dags/\ directory, which contains templates and configurations for generating Apache Airflow and Dagster DAGs, enabling users to deploy their pipelines as scheduled workflows. The sample also features Jinja-based DDL generation for BigQuery and Synapse, and demonstrates various load strategies like SCD2 and upserts.
samples/any-engine · high confidence
New BigQuery and DuckDB schema conversion utilities
Added new utility objects in the conversion package to handle schema mapping and data type conversions for BigQuery and DuckDB. BigQueryUtils provides functions to convert Spark schemas to BigQuery schemas, normalizes integer types to long types for compatibility, and handles field value extraction from BigQuery responses. DuckDbUtils converts Spark data types to DuckDB column definitions, supporting nested structures, decimals, and JSON types. Conversions.scala adds an implicit conversion to treat Hadoop RemoteIterators as Scala Iterators.
src/main/scala/ai/starlake/utils/conversion · high confidence
New BigQuery sample with Spark loader and deployment guides
Added a new BigQuery sample (\samples/bigquery\ and \samples/bigquery-spark-loader\) that provides a complete workflow for importing, loading, and transforming data into BigQuery using Starlake and Spark. The sample includes shell scripts for each pipeline stage (init, import, load, transform), environment configuration files, and documentation for running the workload on Google Cloud Dataproc, Docker, and Cloud Run. It also includes sample data files for HR and sales scenarios to demonstrate the integration.
samples/bigquery · high confidence
New BigQuery schema and metadata extraction capabilities
This change introduces dedicated commands and extraction logic for Google BigQuery, allowing users to extract table schemas and dataset metadata directly. The \extract-bq-schema\ command retrieves BigQuery table schemas and dataset metadata, while the new \bq-info\ command persists information about BigQuery tables and datasets (such as creation time, size, and row counts) into an audit sink. These additions are implemented in new files including \BigQueryProjectInfo\, \BigQueryTableInfo\, \ExtractBigQuerySchema\, and their corresponding command objects, alongside a new \DucklakeAttachment\ utility for parsing Ducklake attachment strings.
src/main/scala/ai/starlake/extract · high confidence
New BigQuery/Spark loader sample with multi-dialect metadata and expectations
The BigQuery/Spark loader sample now includes a comprehensive metadata configuration that defines table schemas, load/transform pipelines, and data quality expectations. The sample configures BigQuery connections with direct/indirect write methods and schema evolution options, while also providing DDL templates for BigQuery, Snowflake, Postgres, and Synapse to support multi-database deployments. It introduces a Jinja2-based expectations framework with macros for data validation (e.g., uniqueness checks, row count ranges) and demonstrates their use in load definitions. The sample also includes external audit and KPI table definitions, HR and sales load configurations with primary keys and access policies, and a KPI transformation SQL script, providing a complete reference for setting up data ingestion, quality checks, and transformation workflows.
samples/bigquery-spark-loader/metadata · high confidence
New CLI command to generate a standalone documentation site
Users can now run the \starlake site\ command to generate a static, standalone documentation website from their project metadata (schemas, domains, jobs, and lineage). The implementation replaces the previous Docusaurus-based approach with a custom Jinjava-based renderer, producing a responsive HTML site with navigation for domains and jobs. The command supports configuring the output directory, selecting a template (defaulting to 'standalone'), choosing between HTML or JSON output formats, and cleaning the output directory before generation.
src/main/scala/ai/starlake/job/site · high confidence
New CLI command to load files directly into JDBC tables
Users can now use the new \starlake cnxload\ command to load Parquet, CSV, or JSON files directly into a JDBC-connected database table. This feature introduces a dedicated CLI interface (\JdbcConnectionLoadCmd\) and backend logic (\SparkJdbcWriter\) that supports configurable connection options and write strategies (APPEND or OVERWRITE). The implementation automatically handles schema management by creating tables if they do not exist and altering them to add or drop columns as needed. For DuckDB targets specifically, the system uses a temporary Parquet intermediate file to bypass Spark JDBC driver limitations, ensuring reliable data insertion.
src/main/scala/ai/starlake/job/sink/jdbc · high confidence
New CLI commands for ACL export, ACL dependency graphs, and column-level lineage
This change introduces three new command-line capabilities in the lineage module. The \starlake acl --export\ command allows users to export Access Control List (ACL) and Row Level Security (RLS) entries from domain and task definitions into a YAML file. The \starlake acl-dependencies\ command generates GraphViz dependency graphs (as DOT, SVG, PNG, or JSON) to visualize ACL relationships across tables and domains. Additionally, the \starlake col-lineage\ command builds and exports column-level data lineage for specific tasks as JSON, enabling users to trace data flow and provenance of computed fields across transformations.
src/main/scala/ai/starlake/lineage · high confidence
New CLI commands for DAG generation, deployment, and Excel-to-YAML conversion
The \src/main/scala/ai/starlake/schema/generator\ package now includes new commands to support workflow orchestration and configuration management. Users can generate DAG files for tasks and domains using \dag-generate\ (with optional tag filtering and role definitions) and deploy them to orchestration tools like Airflow using \dag-deploy\, which also handles library installation and \.airflowignore\ setup. Additionally, the \xls2yml\ and \xls2ymljob\ commands allow users to convert Excel spreadsheets describing domains, schemas, and transform jobs into Starflow YAML configuration files, including support for IAM policy tags.
src/main/scala/ai/starlake/schema/generator · high confidence
New DuckLake sample for TPC-H data pipeline
Added a new sample project in samples/ducklake that demonstrates an end-to-end data pipeline using Starflow and DuckLake. The sample includes configuration for local (DuckDB) and cloud (PostgreSQL) connections, TPC-H table schemas, SQL transformation logic, orchestration templates for Airflow, Dagster, and Snowflake, and a comprehensive set of data quality expectation macros.
samples/ducklake · high confidence
New Iceberg sample with Starbake dataset and multi-engine DAG templates
The \samples/iceberg\ area now includes a complete sample pipeline for the 'Starbake' dataset, featuring source data (CSV, JSON, XML) and metadata definitions for loading \order\_lines\, \orders\, and \products\ tables into an Iceberg warehouse. It provides configuration for multiple execution engines (Spark, DuckDB, BigQuery, Snowflake, PostgreSQL, Redshift) and includes pre-built DAG templates for Airflow and Dagster, supporting both top-down and bottom-up dependency strategies for load and transform tasks.
samples/iceberg · high confidence
New JSON ingestion utility for schema validation and type coercion
Added JsonIngestionUtil.scala, a new utility object that provides schema validation and type coercion logic for JSON data sources. It includes a compareTypes method to validate dataset types against a defined schema (handling structs, arrays, and numeric promotions) and a findTightestCommonTypeOfTwo function to determine compatible types, mirroring Spark's TypeCoercion behavior to improve JSON ingestion robustness.
src/main/scala/org/apache/spark/sql/execution/datasources/json · high confidence
New Java-based Setup utility for dependency management
A new \Setup.java\ class has been introduced to handle the installation, configuration, and dependency management for the Starflow project. This utility automatically downloads and configures required components such as Spark, DuckDB, Snowflake, Redshift, and Kafka, supporting both Unix and Windows environments. It features interactive CLI menus for selecting dependencies, handles proxy configuration (HTTP, HTTPS, SOCKS) with authentication, and manages SSL settings, including an option to disable certificate validation via the \SL\_INSECURE\ environment variable. The tool also ensures Windows compatibility by installing necessary VC++ runtime DLLs for Hadoop utilities and generates version configuration files for the target platform.
src/main/java · high confidence
New Markdown CLI documentation template
A new Jinja2 template (md-cli.j2) has been added to generate Markdown documentation for CLI tools. This template structures output with a YAML front matter section (including sidebar position, title, description, and keywords), followed by Synopsis, Description, and a formatted Parameters table that iterates over command-line options to display their names, values, required status, and descriptions.
src/main/resources/templates/cli · high confidence
New Project Comparison Report Template
A new Jinja2 template (index.html.j2) has been added to render a visual Project Comparison Report. This report displays summary cards for added, deleted, and updated domains, and provides collapsible sections detailing changes at the table and attribute levels, including specific metadata field updates with before/after values.
src/main/resources/templates/compare · high confidence
New REST API data and schema extraction commands
Users can now extract data and schemas from generic REST API endpoints using the new \starlake extract-rest-data\ and \starlake extract-rest-schema\ CLI commands. The data extraction command supports multiple pagination strategies (offset, cursor, link header, page number), various authentication methods (Bearer, API Key, Basic, OAuth2), and output formats including CSV and JSON Lines. It also includes features for incremental extraction, resume support, and parent-child endpoint relationships. The schema extraction command infers table structures from sample API responses and detects schema evolution.
src/main/scala/ai/starlake/extract/impl/restapi · high confidence
New Redshift write strategy templates
Added new Jinja2 SQL templates for Amazon Redshift write strategies, including append, create, delete\_then\_insert, overwrite, overwrite\_by\_partition, scd2, upsert\_by\_key, upsert\_by\_key\_and\_timestamp, and view. These templates provide the underlying SQL logic for various data loading and transformation operations on Redshift tables.
src/main/resources/templates/write-strategies/redshift · high confidence
New SPI for generic REST and OpenAPI schema extraction
The schema extraction subsystem now supports extracting schemas from generic REST APIs and OpenAPI specifications, in addition to existing JDBC sources. This is enabled by new \SchemaExtractor\ and \SchemaExtractorFactory\ components that instantiate specific extractors (OpenAPI or REST) based on configuration, and a \SchemaExtractorWorkflow\ that manages the extraction process, including merging extracted metadata with existing domain and table configurations and serializing the results to YAML.
src/main/scala/ai/starlake/extract/spi · high confidence
New SQL expectation macros for data quality validation
The empty-project template now includes a comprehensive set of Jinja2 macros in the \metadata/expectations\ directory (Completeness, Numeric, Schema, UnexpectedRowsExpectation, Uniqueness, Validity, and Volume) that generate SQL queries for validating data quality. These macros enable checks for null values, numeric ranges and distributions, schema structure and types, uniqueness constraints, value validity and patterns, and row counts, allowing users to define and run data quality expectations directly against their database tables.
src/main/resources/templates/bootstrap/samples/empty-project/metadata/expectations · high confidence
New SQL generation logic for views and materialized views
The system now supports generating SQL for views and materialized views through the new TransformStrategiesBuilder. This component handles the creation of view-specific SQL templates and automatically corrects trailing semicolons in queries to ensure compatibility with the underlying SQL engine.
src/main/scala/ai/starlake/job/strategies · high confidence
New \`starlake migrate\` command for automatic project configuration updates
Users can now run the \starlake migrate\ command to automatically upgrade Starflow project configuration files (YAML) to the latest version format. This new capability handles the migration of various metadata areas—including extract, load, refs, application, types, DAGs, and environment settings—by detecting eligible files, transforming their structure, and performing post-migration actions such as renaming folders or deleting obsolete files, thereby reducing the manual effort required for version upgrades.
src/main/scala/ai/starlake/migration · high confidence
New bootstrap command to scaffold Starflow projects from templates
Users can now initialize a new Starflow project using the \starlake bootstrap\ command. This new CLI command scaffolds the required directory structure, copies sample configuration files (such as \application.sl.yml\), and sets up IDE support files (VS Code settings, \.gitignore\, and JSON schema) based on a selected template (e.g., \quickstart\ or \userguide\). It also supports specifying a custom project location and allows users to select database connection types during the initialization of the \initializer\ template.
src/main/scala/ai/starlake/job/bootstrap · high confidence
New configuration and connection management infrastructure
This change introduces a new configuration layer for the Starlake platform, centered around a new \ConnectionInfo\ case class that explicitly models database connections (including JDBC, BigQuery, and DuckDB variants) with support for encrypted passwords, access tokens, and specific engine dialects. The \Settings\ object has been refactored to use Pureconfig for loading application configuration and now manages these connections via a \Connections\ map. Additionally, new utilities have been added to handle dataset area paths (\DatasetArea\), Spark session building with support for Spark Connect and various local catalogs (\SparkEnv\, \SparkSessionBuilder\), and Kryo serialization for these new configuration models. A new \SettingsCmd\ CLI command allows users to display resolved settings or test database connections directly.
src/main/scala/ai/starlake/config · high confidence
New configuration-driven schema melding and template utility utilities
This change introduces new core utilities to support schema reconciliation and templating. The \LoadConfigMelder\ adds a configurable strategy for merging extracted metadata with existing domain and table schemas, allowing users to define precedence rules (e.g., keeping current vs. extracted values) for attributes like privacy, tags, and comments. Additionally, \CaseClassToPojoConverter\ enables the conversion of Scala case classes to Java maps for use in Jinjava templating, while \NamingUtils\ and \StringUtils\ provide robust string normalization (e.g., handling OpenAPI parameters, accents, and camelCase) for table and attribute names.
src/main/scala/ai/starlake/core · high confidence
New data quality expectation templates for sample projects
The sample project template now includes a comprehensive set of Jinja2 macros for generating SQL-based data quality expectations. New files in the metadata/expectations directory provide checks for Completeness (null handling), Numeric statistics (KL divergence, min/max/mean/median bounds, most common values), Schema validation (column existence, type checking, column count and ordering), Uniqueness (distinct values, compound uniqueness), Validity (value sets, length checks, pattern matching), and Volume (row counts). These templates allow users to bootstrap projects with robust, out-of-the-box data validation logic.
src/main/resources/templates/bootstrap/samples/sample-project/metadata/expectations · high confidence
New data quality metrics and expectation validation capabilities
This change introduces new components in the metrics package to support data quality profiling and expectation assertions. Users can now compute and publish continuous, discrete, and frequency metrics for loaded tables via the new \MetricsCmd\ CLI command and \MetricsJob\ engine. Additionally, a new expectation handling system is added, featuring \ExpectationJob\ to execute SQL-based assertions, \ExpectationAssertionHandler\ implementations for Spark, BigQuery, and JDBC sinks, and \ExpectationReport\ structures to log results. This allows users to define data quality rules that are validated during ingestion or transformation, with failures optionally causing the job to fail based on configuration.
src/main/scala/ai/starlake/job/metrics · high confidence
New extract sample demonstrating REST API data ingestion
The samples/extract directory now includes a new example for ingesting data from SaaS APIs via a generic REST API extractor. This addition provides a configuration file (rest-api-extract.sl.yml) showcasing support for bearer token authentication, multiple pagination strategies (offset, cursor, and page number), child endpoint relationships, and POST requests with request bodies. The sample is accompanied by setup scripts and template files to help users understand how to configure and run extractions from external REST endpoints.
samples/extract · high confidence
New interactive console command for local data exploration
Users can now run the \starlake console\ command to launch an interactive shell for exploring metadata and executing commands locally. This new feature provides a command-line interface that supports command history, environment switching, SQL execution via the \.\ prefix, and integration with the existing Starflow services, allowing for easier local development and debugging without needing the full web interface.
src/main/scala/ai/starlake/console · high confidence
New project comparison tool and adaptive write strategies
This update introduces two new capabilities in the schema module. First, it adds a 'compare' command (ProjectCompareCmd) that allows users to compare two versions of a Starflow project by file path, git commit, or git tag, generating a diff report via a customizable Jinja template. Second, it introduces AdaptiveWriteStrategy, which dynamically selects write strategies (such as APPEND, OVERWRITE, UPSERT\_BY\_KEY, DELETE\_THEN\_INSERT, etc.) based on file metadata and conditions, grouping files by the resulting strategy for ingestion.
src/main/scala/ai/starlake/schema · high confidence
New release, Docker, and utility scripts for Starflow
The scripts directory now includes a comprehensive set of tooling for the Starflow release process and container management. A new idempotent local-release.sh script manages the end-to-end release workflow, including version synchronization between core and API repositories, building the UI, creating GitHub releases with assets, and announcing releases on Discord. Supporting this are release-lib.sh for shared helpers and announce-release-discord.sh for Discord notifications. Docker operations are streamlined with docker-build.sh, docker-prepare.sh, and docker-multi-arch-build.sh, which handle building and publishing images for various environments (cloud, local, dev) and architectures. Additional utilities include pyspark-smoke-test.sh for verifying PySpark functionality in Docker images, docker-bash.sh for interactive debugging, and versions.sh for managing version variables.
scripts · high confidence
New row-level validation framework for data ingestion
The \src/main/scala/ai/starlake/job/validator\ directory now contains a new row-level validation system that validates and transforms incoming data against schema definitions. This includes a \RowValidator\ class that handles privacy transformations, type casting, and pattern matching, along with specific validator implementations (\AcceptAllValidator\, \FlatRowValidator\, \TreeRowValidator\, \NativeValidator\) that orchestrate the validation workflow. The system produces structured rejection records (\SimpleRejectedRecord\) containing error messages and file paths, allowing users to identify and handle invalid data rows during ingestion.
src/main/scala/ai/starlake/job/validator · high confidence
New semantic-export command to convert models to LookML, Power BI TMDL, and Apache Ossie
Users can now export semantic models defined in Starflow to three external formats using the new \starlake semantic-export\ command. The command supports Apache Ossie (the default), LookML (generating a project with view and model files, requiring a \--connection\ flag for the Looker connection), and Power BI TMDL (generating a folder structure with database, model, and table files, using \--connection\ to derive Power Query sources). The export logic handles Snowflake-style semantic model structures, mapping dimensions, facts, metrics, and relationships while preserving Starflow-specific attributes in Ossie custom extensions and handling TMDL-specific constraints like composite keys and DAX measure translation.
src/main/scala/ai/starlake/semantic · high confidence
New standalone installation and runtime scripts for Windows and Unix
The distrib area now includes standalone setup and runtime scripts (setup.sh, setup.ps1, starlake.sh, starlake.cmd) and a log4j2 configuration, replacing the previous distribution method. The new setup scripts allow users to install Starflow from GitHub Releases with an interactive menu or specific version flag, automatically resolving Java versions (defaulting to an embedded JDK if needed) and configuring the PATH. The runtime scripts (starlake.sh/cmd) handle environment setup, including preferring embedded JDKs, managing proxy settings via environment variables, and enforcing a consistency check to detect and repair interrupted upgrades or reinstalls. A new .gitattributes file ensures correct line endings for Windows batch files.
distrib · high confidence
New unit testing framework for load and transform tasks
Introduces a new \starlake test\ command that allows users to run unit tests for load and transform tasks on specific domains and tables. The framework supports filtering by domain, table, or specific test, and can generate HTML reports or JUnit XML output. It utilizes DuckDB for in-memory testing, supports pre-test SQL scripts, and provides detailed result summaries including success rates and durations.
src/main/scala/ai/starlake/tests · high confidence
New utility modules for encryption, CLI configuration, and template loading
The utils package now includes several new components: AESEncryption for secure handling of sensitive data, CliConfig and CliEnvConfig for standardized command-line argument parsing and documentation generation, AnyTemplateLoader for loading DAG templates from resources or external paths, and CometJacksonModule for custom Jackson serialization of Spark types and protection of Scala singletons. These additions support improved security, configuration management, and template-driven orchestration.
src/main/scala/ai/starlake/utils · high confidence
Register DuckDB as a Spark SQL data source
The application now registers the DuckDB relation provider with Apache Spark SQL, enabling users to query DuckDB data directly using standard Spark SQL syntax via the registered data source interface.
src/main/resources/META-INF · high confidence
Schema model refactoring and new security/merging capabilities
The schema model layer has been significantly refactored to support new data governance and schema evolution features. A new \AccessControlEntry\ class has been added to define and apply granular access control policies (grants) across various database engines like Spark, BigQuery, and Snowflake. Schema merging logic has been extracted into a dedicated \AttributeMerger\ component, introducing configurable strategies (\AttributeMergeStrategy\) to control how incoming schema changes are applied to existing definitions (e.g., keeping source diffs or dropping them). Additionally, new model classes for \DagInfo\ and \AutoJobInfo\/\AutoTaskInfo\ have been introduced to support automated DAG generation and more detailed task configuration, while \AnyRefDiff\ provides generic utilities for comparing and diffing complex model objects.
src/main/scala/ai/starlake/schema/model · high confidence
Support for expectations
The system now supports defining and validating data expectations, allowing users to specify quality rules for their datasets. This feature enables automated checks to ensure data integrity and consistency during ingestion and processing workflows.
(repo-wide) · high confidence
Architecture
Workflow logic restructured into focused trait mixins
The monolithic workflow implementation has been split into distinct, composable traits (InferWorkflow, IngestionWorkflow, MetricsSecurityWorkflow, SinkWorkflow, TestWorkflow, and TransformWorkflow) that are mixed into the main IngestionWorkflow class. This refactoring isolates specific capabilities—such as schema inference, data ingestion staging, security policy application, sink operations, testing, and transformations—into their own modules, improving code organization and maintainability without changing the external user-facing behavior.
src/main/scala/ai/starlake/workflow · high confidence
Behavioural changes
BigQuery write strategies now use native MERGE and partition-aware operations
The BigQuery write templates have been replaced with new implementations that leverage BigQuery's native MERGE statement for upserts, SCD2, and delete-then-insert strategies, replacing previous approaches. These templates now support partition pruning via DECLARE variables to optimize performance on large partitioned tables, handle materialized views explicitly, and ensure partition filter requirements are managed correctly during overwrites.
src/main/resources/templates/write-strategies/bigquery · high confidence
CLI command execution now returns proper exit codes
The Starflow CLI now correctly propagates command results to the operating system via exit codes. Previously, the main entry point did not distinguish between successful and failed executions; now, the \Main\ object inspects the result of each command and calls \System.exit\ with 0 for success and 1 for any failure (including exceptions or soft failures like empty preload results). This allows scripts and orchestration tools to reliably detect when a Starflow command has failed.
src/main/scala/ai/starlake/job · high confidence
Data Catalog site replaces Docusaurus with a standalone Jinja/React implementation
The documentation and catalog interface has been migrated from Docusaurus to a lightweight, standalone site rendered with Jinja templates and styled with custom CSS. The new layout features a responsive sidebar for navigating domains and jobs, and table/task detail pages now use tabs to organize General, Attributes, SQL/Python, and Expectations content. Interactive data lineage and access control diagrams are rendered client-side using React and the @xyflow/react library, replacing the previous static or framework-dependent visualizations.
src/main/resources/templates/site · high confidence
Fallback for Dataset.showString in Spark 4 compatibility layer
A new \DatasetLogging\ trait provides an implicit helper to safely call \showString\ on Spark \Dataset\ objects. Since Spark 4 moved the \Dataset\ implementation to \org.apache.spark.sql.classic\, this helper detects classic datasets and delegates to their \showString\ method; for other dataset types (such as Connect datasets), it falls back to returning the schema tree string instead of failing.
src/main/scala/org/apache/spark/sql · high confidence
Installer now downloads only changed dependencies
The setup process has been updated to perform incremental installs and upgrades. Instead of re-downloading all dependencies every time, the installer now compares the desired artifacts against files already present on disk, downloading only those that are missing, outdated, or have changed size. It also automatically removes obsolete or disabled dependency files to keep the installation directory clean.
src/main/java/ai/starlake/setup · high confidence
Introduce modular reference configuration files and sample data
The monolithic reference configuration has been split into dedicated files (reference-audit.conf, reference-connections.conf, reference-dags.conf, reference-engines.conf, reference-expectations.conf, reference-extra.conf, reference-general.conf, reference-internal.conf, reference-kafka.conf, reference-metrics.conf, reference-refs.conf, and reference-service.conf) to organize settings by domain. Additionally, sample data files (SCHEMA-VALID.dsv, SCHEMA-VALID-NOHEADER.dsv) and an \_\init\\_.py file have been added to the resources directory.
src/main/resources · high confidence
Introduce native BigQuery and JDBC transformation tasks with expectation support
The transform module now includes dedicated task implementations for native execution on BigQuery and JDBC databases, alongside the existing Spark-based processing. BigQueryAutoTask enables direct SQL execution on BigQuery without Spark overhead, supporting native branching via SL\_DATA\_BRANCH and OAuth access tokens. JdbcAutoTask provides native transactional execution for JDBC-compatible databases (PostgreSQL, MySQL, Snowflake, Redshift, DuckDB) with support for MERGE and SCD2 write strategies. Both new task types integrate with the expectation framework, allowing data quality assertions to be run interactively or persisted to audit tables. The refactored AutoTask base class and its subclasses (SparkAutoTask, BigQueryAutoTask, JdbcAutoTask) unify the transformation execution model while delegating to the appropriate engine based on connection type.
src/main/scala/ai/starlake/job/transform · high confidence
Migration tooling now supports versioned metadata and updated type mappings
The migration module has been updated to support a new versioned metadata format (version 1) for configuration files, introducing a structured schema for application connections, DAG templates, environment settings, and extract/load/transform definitions. This change includes a comprehensive update to the default type definitions, which now provide explicit DDL mappings for BigQuery, Snowflake, Postgres, and Synapse across various data types (e.g., string, integer, timestamp). Additionally, the migration logic has been adjusted to handle legacy 'unversioned' metadata structures, ensuring compatibility during the transition to the new versioned format.
migration · high confidence
New SQL statement generation model for task workflows
A new \TaskSQLStatements\ and \WorkflowStatements\ data model has been introduced to structure SQL generation for tasks. This model explicitly supports a \syncStrategy\ (defaulting to ADD or NONE based on target schema presence) and organizes SQL actions into distinct phases (pre, main, post) along with schema creation and SCD2 column additions. The model provides a standardized map output for downstream processing, including support for expectations, audit, and ACL statements within the workflow context.
src/main/scala/ai/starlake/job/common · high confidence
New Spark SQL templates for write strategies
The Spark write-strategy templates have been replaced with new Jinja2 files that generate specific SQL operations. The system now supports append, delete-then-insert, overwrite (with Delta Lake and materialized view handling), overwrite by partition, SCD2 (Slowly Changing Dimension Type 2), upsert by key, upsert by key and timestamp, and view creation. These templates define the exact SQL logic used for data loading and transformation in Spark environments.
src/main/resources/templates/write-strategies/spark · high confidence
New default write-strategy template with column quoting fixes
A new default Jinja2 template for write strategies has been introduced, providing macros to generate SQL for strategy keys, joins, and column lists. This template includes a specific fix to unquote column names before re-quoting them, ensuring that identifiers are handled correctly without double-quoting or syntax errors when generating SQL statements.
src/main/resources/templates/write-strategies · high confidence
New exception types for ingestion and validation errors
The system now introduces specific exception classes to handle distinct failure scenarios during data processing. Users will see \DisallowRejectRecordException\ when a rejected record count limit is exceeded, \DataExtractionException\ when data extraction fails for a specific domain and table, and \SchemaValidationException\ when JSON schema validation fails. These changes improve error granularity and clarity for ingestion and validation issues.
src/main/scala/ai/starlake/exceptions · high confidence
Package namespace renamed from com.ebiznext to ai.starlake
The project's Java/Scala package namespace has been renamed from com.ebiznext to ai.starlake. This change affects all source files and imports within the project, requiring users to update any external code or configurations that reference the previous package names.
project · high confidence
PostgreSQL write strategies now use MERGE and standard SQL constructs
The PostgreSQL write-strategy templates have been rewritten to use modern SQL syntax, replacing previous implementations with MERGE statements for upserts (upsert\_by\_key, upsert\_by\_key\_and\_timestamp) and SCD2, and standard INSERT/DELETE patterns for append, overwrite, and delete\_then\_insert. This change improves compatibility with PostgreSQL 15+ and ensures consistent behavior across strategies like SCD2, which now correctly handles temporal updates using MERGE and temporary views.
src/main/resources/templates/write-strategies/postgresql · high confidence
Restored Oracle database export script template and expected outputs
The \samples/database\ area now includes the \EXTRACT\_TABLE.sql.j2\ Jinja2 template for generating Oracle SQL\*Plus scripts, along with \expected\_script\_payload.txt\ and \expected\_script\_payload2.txt\ as reference outputs. This restores the capability to generate incremental and full data exports from Oracle tables, handling configuration such as custom delimiters, date formats, and export status tracking via the \COMET\_EXPORT\_STATUS\ table.
samples/database · high confidence
Schema handler refactored into focused helper classes
The monolithic SchemaHandler has been refactored into a set of focused helper classes to improve code organization and maintainability. New components include AccessControlQueries for managing ACLs, Row-Level Security, and IAM Policy Tags; DagHandler for loading and validating DAG generation configurations; SchedulingQueries for handling table and task schedules; and InferSchemaHandler for improved schema inference logic. This change also introduces the InvalidFieldNameException for stricter field name validation and updates storage handlers (HdfsStorageHandler, LocalStorageHandler) to support the new configuration and file management patterns.
src/main/scala/ai/starlake/schema/handlers · high confidence
Snowflake write strategies now support hybrid tables, materialized views, and clustered columns
The Snowflake write-strategy templates have been updated to support new Snowflake capabilities. Users can now create and overwrite HYBRID TABLES and MATERIALIZED VIEWs in addition to standard tables. The templates also support defining clustered columns via the sinkTableClusteringClause and utilize the QUALIFY clause for deduplication logic in source-and-target strategies. Existing strategies like SCD2, upsert, and delete-then-insert have been refined to handle these new object types and clustering configurations correctly.
src/main/resources/templates/write-strategies/snowflake · high confidence
Standardized load-strategy templates and added Snowflake context support
The load-strategy templates for BigQuery, DuckDB, JDBC, PostgreSQL, Redshift, and Spark have been replaced with empty placeholders, while the Snowflake template now explicitly exposes the Python JSON context, statements, expectation items, audit data, and expectations to the rendering engine. A new shared default template provides reusable macros for handling pseudo-columns, join conditions, and column quoting, ensuring consistent SQL generation logic across supported databases.
src/main/resources/templates/load-strategies · high confidence
Unify JDBC write strategy templates
The SQL templates for JDBC write strategies (append, create, delete\_then\_insert, overwrite, overwrite\_by\_partition, scd2, upsert\_by\_key, upsert\_by\_key\_and\_timestamp, and view) have been consolidated into a single unified location. This change standardizes how data is written to target tables, ensuring consistent behavior across different write modes such as upserts, SCD2, and overwrites.
src/main/resources/templates/write-strategies/jdbc · high confidence
Updated bootstrap samples with new documentation and configuration files
The bootstrap samples now include a README explaining how to run extract, load, and transform workflows for various data warehouses (BigQuery, Snowflake) and local environments. Additionally, an extensions.json file recommends specific VS Code extensions (IBM.output-colorizer, samuelcolvin.jinjahtml, Starlake.starlake), and a gitignore file is provided to exclude most files except metadata and datasets directories.
src/main/resources/templates/bootstrap/samples · high confidence
Test coverage
Added BigQuery integration test suite; Added DDL template tests for BigQuery and Synapse; Added Jinja2 macro templates for test expectations; Added Kafka integration test suites; Added SQL test cases for view rewriting logic; Added comprehensive test coverage for data extraction and migration features; Added comprehensive test coverage for schema model components; Added comprehensive tests for RowValidator; Added configuration and connection-info unit tests; Added disabled Snowflake freshness integration test; Added expected DOT graph outputs for ACL and domain model tests; Added integration test infrastructure for BigQuery and JDBC; Added integration tests for BigQuery, Postgres, Redshift, and Snowflake transforms; Added integration tests for Starbake workflows; Added integration tests for Starlake CLI commands and utilities; Added integration tests for data and schema extraction; Added integration tests for data loading across multiple storage backends; Added integration tests for lineage and dependency commands; Added native test resources for sales data loading; Added sample Python Pi job and its SL configuration; Added sample data files for schema inference tests; Added sample metadata and Jinja templates for testing transform configurations; Added sample test data for sales quickstart; Added sample test resources for schema validation and ingestion; Added test coverage for JSON, position, and load strategy ingestion; Added test coverage for Starlake utility functions; Added test coverage for metrics jobs and expectation report escaping; Added test coverage for schema generator components; Added test fixtures for CSV schema inference; Added test fixtures for DSV max-errors validation; Added test fixtures for DSV rename and position-based header replay; Added test fixtures for HR quickstart validation; Added test fixtures for JSON DuckDB native loader; Added test fixtures for Jinja delimiter escaping in DSV loading; Added test fixtures for POSITION format validation and DSV options; Added test fixtures for SQL export script generation; Added test fixtures for adaptive write, merge, and schema extraction scenarios; Added test fixtures for merge operations; Added test fixtures for merge scenarios; Added test fixtures for position-based parsing and custom types; Added test for ESLoad command usage output; Added test for JDBC-to-JDBC connection jobs; Added test resource configuration and data fixtures; Added test resource marker file; Added test resources for DSV DuckDB replay failure scenarios; Added test resources for DSV pre-existing table loading; Added test resources for Dream sample data and schema; Added test resources for DuckDB native DSV reject handling; Added test resources for DuckDB native loader POSITION format; Added test resources for XSD-based XML validation and type definitions; Added test resources for adaptive write strategies; Added test resources for position-based encoding with ISO-8859-1 data; Added test resources for position-based ingestion with regex filtering and ignore logic; Added test resources for the test framework; Added test template for DAG generation with raw domain support; Added tests for BigQuery and DuckDB schema conversion utilities; Added tests for BootstrapConfig usage output; Added tests for CaseClassToPojoConverter template rendering; Added tests for DuckDB native load second-step failure and retry scenarios; Added tests for JDBC connection load job command usage; Added tests for JSON Schema inference and CLI usage verification; Added tests for JSON ingestion schema validation; Added tests for LoadConfigMelder domain merging logic; Added tests for Parquet to CSV conversion job; Added tests for Quack CLI parsing, state management, and server integration; Added tests for SettingsManager configuration updates; Added tests for installer dependency sync and version pinning; Added tests for lineage dependency graph cycle handling and table metadata emission; Added tests for native loader reject capture, replay file naming, and Snowflake POSITION support; Added tests for schema inference and domain directory handling; Added tests for semantic model export to LookML, TMDL, and Ossie formats; Added tests for the generic REST API extractor; Added tests for the privacy engine's data transformation and hashing capabilities; Added tests for transform job handling and DuckDB write strategies; Added tests for transform strategy SQL generation across multiple engines; Added unit tests for SQL reference substitution and table extraction; Added unit tests for the Starlake test framework; Expanded test coverage for DuckDB native ingestion and audit behaviors; New HTML test report templates for Starflow; New test infrastructure for JDBC validation and adaptive write strategies.
Dependencies
Upgrade to Spark 4.1.3 and Java 17 with updated Google Cloud dependencies
The build configuration has been updated to target Apache Spark 4.1.3 and Java 17 (via \--release 17\), replacing previous Spark 3 and Java 8/11 support. Scala 2.13.18 is now the default and only supported version, with Scala 2.12 and 2.11 dropped. Google Cloud dependencies have been aligned to BigQuery 2.68.0, DataCatalog 1.101.0, Logging 3.36.0, and Protobuf 4.35.1 (with a runtime override to 4.36.1 to satisfy gencode requirements). The assembly process now excludes Conscrypt to prevent JVM crashes on macOS arm64 and Linux, and uses specific merge strategies for Spark service files.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Baseline
- First survey — no prior run to compare against. CAI 48.
Lenses
- Code Health 86
- Architecture 92
- Maturity 74
- Readiness 37
- Security 49
- Accessibility 49
Changes since last survey
- 300 commits — 228 feature/other, 72 fixes
By area
- (root) — 72 commits
- src/main — 70 commits
- (repo) — 41 commits
- src/test — 34 commits
- distrib/setup.jar — 23 commits
- .github/workflows — 12 commits
- docs/superpowers — 9 commits
- distrib/starlake.cmd — 8 commits
- project/Versions.scala — 8 commits
- project/Dependencies.scala — 5 commits
- distrib/setup.ps1 — 4 commits
- distrib/python-libs — 3 commits
- scripts/local-release.sh — 3 commits
- docs/flightsql.md — 2 commits
- scripts/announce-release-discord.sh — 2 commits
- distrib/setup.sh — 1 commit
- docs/GITHUB_ACTIONS.md — 1 commit
- project/PythonLibs.scala — 1 commit
- samples/templates — 1 commit
Notable commits
- fix: Merge branch-1.8 into master: OVERWRITE by-name fix, sync hint, syncSqlWithYaml (#1722, #1723)
- fix: Merge pull request #1694 from starlake-ai/fix/setup-ps1-java-probe-eap
- fix: Merge pull request #1695 from starlake-ai/fix/winutils-vcruntime
- fix: Merge pull request #1696 from starlake-ai/fix/api-install-move-and-verify
- fix: Merge pull request #1720 from starlake-ai/fix/semantic-export-mdx-placeholder
- fix: Merge pull request #1721 from starlake-ai/fix/cli-doc-links-starflow-prefix
- fix: Merge pull request #1758 from starlake-ai/fix/transform-bq-fixture
- fix: Merge pull request #1762 from starlake-ai/fix/conscrypt-exclusion
- fix: Merge pull request #1766 from starlake-ai/fix/installer-version-drift
- fix: Merge pull request #1793 from starlake-ai/fix/1792-presql-separator
- fix: Merge pull request #1796 from starlake-ai/fix/ci-snowflake-jdbc-noconscrypt
- fix: chore: fix stale references surfaced by the naming audit
- fix: config(spark4): keep ANSI off for Spark 3 parity; fix stale datetime rebase keys
- fix: fix(autoload): skip inference for already-defined tables instead of failing
- fix: fix(bigquery): terminate every presql statement in the native script
- fix: fix(bootstrap): valid schedule and autoload-ready data layout in sample-project
- fix: fix(bq): build the DataCatalog policy client only when CLS is needed
- fix: fix(ci): guard release trigger from snapshot tags, shasum -c compatible snapshot checksums, drop orphaned publishSigned
- fix: fix(docker): install python3 so PySpark transform tasks can execute
- fix: fix(duckdb): drop temp tables left behind by a failed two-step load
- …and 280 more
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
starlake-ai/starflow was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 20 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 994def4a73cf1b25c68e7c30c70f19d21371660b — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-b51f968c9b10.