Skip to content
CAI
Software that uses CAICheck a score

byzer-org/byzer-lang

44.2

Weak · 27 September 2026

71.9k

lines of production code

Scala

with Java, Python

4

measurements over time

CAI band scale
CAI trend line
CAI lens gauges

What this system is

This release introduces a comprehensive SQL Profiler plugin with query optimization, indexing, and dialect support, alongside a new auto-suggestion engine for code completion. It significantly expands data processing capabilities with new ETS commands for Delta Lake, HDFS, and Kafka, while adding distributed TensorFlow and Python-based ML algorithm support. The update also includes a MySQL meta-store implementation, Kubernetes health check endpoints, and extensive integration tests for streaming and batch workflows.

Features

Add AutoML algorithm for automated model selection

A new SQLAutoML algorithm has been introduced to automate the training and selection of the best performing model among a set of classifiers (GBT, Logistic Regression, Naive Bayes, and Random Forest). This feature allows users to train multiple models simultaneously and automatically pick the best one based on a specified performance metric, streamlining the machine learning workflow.

streamingpro-mlsql/src/main/java/tech/mlsql/ets/algs · high confidence

Add DeltaUtils helper for Delta table statistics

A new utility class, DeltaUtils, has been added to the ETS delta package. It provides a method to extract key statistics about a Delta table, including size in bytes, number of files, metadata, protocol, removes, and set transactions, by leveraging the DeltaLog snapshot.

streamingpro-mlsql/src/main/java/tech/mlsql/ets/delta · high confidence

Add Docker support for MLSQL Cluster

Added shell scripts to build and run the MLSQL Cluster as a Docker container. The new \build-docker.sh\ script compiles the cluster JAR and builds a Docker image tagged with the current snapshot version. \run-mlsql-cluster.sh\ launches the cluster container on port 8080, while \run-db.sh\ starts a MySQL 5.7 container to initialize the \mlsql\_cluster\ database. Helper scripts \docker-command.sh\ and the Dockerfile (implied by the build script) support this new containerized workflow.

streamingpro-cluster/dev · high confidence

Add Dockerfile and configuration for MLSQL Cluster containerization

Users can now run the MLSQL Cluster service in a Docker container. This change introduces a new Dockerfile that sets up a Debian-based environment with Java 8, installs necessary dependencies, and configures the MLSQL Cluster application using a dedicated Docker configuration file (application.docker.yml) and a startup script (start.sh). The container is configured to accept the MLSQL Cluster JAR and configuration file as build arguments, allowing for flexible deployment of the cluster service within a containerized environment.

streamingpro-cluster/dev/mlsql-cluster-docker · high confidence

Add HDFS shell command implementations for file operations

New Java classes are added to the \org.apache.hadoop.fs.shell\ package to implement HDFS shell commands, including \WowCommandFactory\, \WowFsCommand\, \WowCopyCommands\ (for \cp\ and \getmerge\), \WowCount\, \WowDelete\ (for \rm\, \rmr\, \rmdir\), \WowLs\ (for \ls\ and \lsr\), \WowMkdir\, and \WowMoveCommands\ (for \mv\). These classes provide the underlying logic for listing, counting, deleting, copying, and moving files and directories on HDFS, with support for options like recursive deletion, preserving file attributes, and JSON-formatted output for \ls\.

streamingpro-mlsql/src/main/java/org/apache/hadoop · high confidence

Add InitializationCompositor for early-stage SQL execution

A new \InitializationCompositor\ component has been introduced to handle initial SQL script execution during the startup phase. This component allows for the execution of SQL statements before the main processing pipeline, supporting features like setting up session contexts and managing job execution for initialization tasks.

streamingpro-mlsql/src/main/java/tech/mlsql/compositor · high confidence

Add MySQL-based meta store implementation

The mlsql-mysql-store module now provides a new MySQL-based meta store implementation. This includes a \MetaStoreService\ that initializes the MySQL connection and registers a \MySQLDBStore\ when the \streaming.metastore.db.type\ parameter is set to 'mysql'. The \MySQLDBStore\ class implements the \DBStore\ interface, enabling the system to read, save, and manage configuration data in a MySQL database, supporting complex data types like lists and maps.

external/mlsql-mysql-store · high confidence

Add ResultRenderManager for extensible result rendering

A new ResultRenderManager object is introduced in the result\_render plugin directory. It provides a call method that iterates through registered ResultRender implementations from the AppRuntimeStore and applies each one to the result response, allowing users to extend or modify the final output format through custom renderers.

_streamingpro-core/src/main/java/tech/mlsql/runtime/plugins/result\render · medium confidence

Add SQL Profiler plugin with command-line and controller interfaces

The SQL Profiler plugin is introduced, providing a new 'profiler' command for users to inspect Spark SQL execution plans and view configuration details. The plugin registers a GenSQLController for programmatic access and includes a ProfilerCommand that supports 'conf' (to view Spark configurations as a temporary view), 'sql' (to execute SQL), and 'explain' (to display query execution plans). Additionally, request cleaners for session and REST datasource temp files are registered.

external/mlsql-sql-profiler-30/src/main/java/tech/mlsql/plugins/sql · high confidence

Add SQL profiler utilities and dialect support for Spark 3.0

Introduced new components for SQL profiling and logical plan-to-SQL generation in the mlsql-sql-profiler module. Added CleanerUtils to manage Spark listener buses and filter execution listeners by session. Implemented the SQLDialect trait and BasicSQLDialect to handle SQL generation for MySQL JDBC connections, including quoting, relation formatting, and limit clauses. Added LogicalPlanSQL to convert Spark logical plans into SQL strings, supporting joins, aggregations, and subqueries. Included ViewCatalyst to manage view-to-table mappings for logical plan resolution.

external/mlsql-sql-profiler-30 · high confidence

Add Spark 3.3.0 adaptor components for Python worker and configuration management

The \streamingpro-spark-3.3.0-adaptor\ module introduces a suite of new classes to support Spark 3.3.0, including \APIDeployPythonRunnerEnv\ and \WowPythonWorkerFactory\ for managing Python worker processes, \MLSQLConf\ for handling configuration entries, and \SparkInstanceService\ for resource tracking. These changes implement the necessary adapters and utilities to run MLSQL on the upgraded Spark runtime.

streamingpro-spark-3.3.0-adaptor · high confidence

Add Spark SQL context holder to streamingpro-spark-common

A new file, SQLContextHolder.scala, has been added to the streamingpro-spark-common module. This file defines a basic class for managing Spark SQL context, providing a foundational component for SQL-related operations within the streaming framework.

streamingpro-spark-common/src/main/java/streaming/core/common · high confidence

Add base evaluation traits for classification and clustering algorithms

New base traits have been introduced to standardize model evaluation: BaseClassification now provides a method to compute multiclass classification metrics (F1, weighted precision, weighted recall, and accuracy) using Spark's MulticlassClassificationEvaluator, while BaseCluster introduces a method to evaluate clustering performance via the Silhouette score using Spark's ClusteringEvaluator. These abstractions allow downstream algorithm implementations to easily integrate standardized evaluation logic.

streamingpro-mlsql/src/main/java/streaming/dsl/mmlib/algs/classfication, streamingpro-mlsql/src/main/java/streaming/dsl/mmlib/algs/cluster · high confidence

Add deployment scripts for MLSQL cluster on Aliyun Cloud

New shell scripts are introduced to automate the deployment and management of an MLSQL cluster on Aliyun ECS. The \run-mlsql-cluster.sh\ script provisions the master and slave instances, configures Spark, and handles SSH key management. The \copy-main-jar-to-slaves.sh\ script copies the main JAR and third-party JARs to slave nodes. The \start-slaves.sh\ script creates slave instances and configures their hosts. The \stop-msql-cluster.sh\ script stops all cluster instances. These scripts support configuration for Spark version, memory, and HDFS-to-OSS integration.

_dev/mlsql\_cluster\cloud · high confidence

Add distributed TensorFlow training support

A new \DistributedTensorflow\ class has been added to the MLSQL engine, enabling distributed TensorFlow training across multiple nodes. This implementation manages the coordination of parameter servers and workers, handling resource allocation, Python process lifecycle management, and logging for distributed training tasks.

streamingpro-mlsql/src/main/java/tech/mlsql/ets/tensorflow · high confidence

Add form parameter handling for UI components

A new FormParams utility class has been added to the common library, introducing support for a wide range of form input types including CheckBox, Input, Text, Select, Radio, Switch, Slider, Rate, InputNumber, TreeSelect, Transfer, TimePicker, Upload, Dynamic, and Editor. This change enables the serialization of these form fields into JSON, facilitating the generation of form configurations for the UI layer.

streamingpro-commons/src/main/java/tech/mlsql/common · high confidence

Add health check endpoints for Kubernetes probes

A new HealthController is introduced to expose /health/liveness and /health/readiness endpoints. The liveness probe returns UP if the Spark runtime is active, and the readiness probe returns IN\_SERVICE when the platform is ready, enabling Kubernetes to monitor the application's health status.

external/mlsql-healthy · high confidence

Add new MLlib algorithm modules and utility functions

Added several new algorithm modules in the streamingpro-mlsql package, including SQLALSInPlace (Alternating Least Squares for recommendation), SQLCommunityBasedSimilarityInPlace (community detection via GraphX), SQLConfusionMatrix (classification evaluation metrics), SQLCorpusExplainInPlace (label distribution analysis), SQLDataSourceExt (external data source management), SQLDicOrTableToArray (dictionary/table to array conversion), and SQLDiscretizer (feature binning). Additionally, new utility functions were introduced: Functions.scala provides model configuration and parameter mapping helpers, while MllibFunctions.scala handles output formatting and model selection logic for AutoML workflows.

streamingpro-mlsql/src/main/java/streaming/dsl/mmlib/algs · high confidence

Add new data source implementations for CSV, CarbonData, Console, Crawler, Delta, DirectJDBC, ElasticSearch, HBase, Hive, Image, JDBC, JSON, and Kafka

The system now supports reading from and writing to a wider variety of data sources. Users can now load and save data to/from CSV files (with options for skipping lines and encoding), CarbonData tables, web console streams, crawler SQL, Delta Lake tables (with versioning and path/table mode support), direct JDBC queries, Elasticsearch, HBase, Hive (with optional Delta Lake integration), image files, standard JDBC connections (with MySQL caching and fetch size optimizations), JSON strings, and Kafka topics. This expands the range of external systems that can be queried or written to directly within MLSQL scripts.

streamingpro-mlsql/src/main/java/streaming/core/datasource/impl · high confidence

Add sample data for machine learning algorithms

Added sample data files for various machine learning algorithms, including ALS, GMM, K-Means, PageRank, and Ridge regression, to support testing and demonstration of these models.

streamingpro-mlsql/src/main/resources-online/data · high confidence

Add streaming and batch data processing components

Added new compositor classes to handle data sources, transformations, and outputs for both batch and streaming workflows. This includes \MultiSQLSourceCompositor\ and \MultiSQLOutputCompositor\ for reading and writing data, \PersistCompositor\ and \UnpersistCompositor\ for managing data caching, and various transformation components like \SQLCompositor\, \AlgorithmCompositor\, and \DFScriptCompositor\ to execute SQL queries, machine learning algorithms, and custom scripts. Additionally, a helper \MultiSQLOutputHelper\ and a UDF \FVectors\ for vector operations were introduced.

streamingpro-mlsql/src/main/java/streaming/core/compositor · high confidence

Added Byzer configuration and utility classes

New classes have been added to the streamingpro-commons module to support Byzer configuration management and security utilities. This includes ByzerConfig for loading and managing byzer.properties and byzer.properties.override files, ByzerConfigCLI for command-line access to configuration values, EncryptUtil for AES-based encryption and decryption of sensitive values, and supporting classes like OrderedProperties and Unsafe. These changes enable the system to read configuration from the BYZER\_HOME environment variable and handle encrypted property values.

streamingpro-commons/src/main/java/tech/mlsql/tool · high confidence

Added DownloadRunner and HDFSTarEntry for file downloads

Users can now download files from HDFS via the new DownloadRunner utility, which supports bundling multiple files or directories into a TAR archive or serving raw file content with optional byte-range seeking.

streamingpro-spark-common/src/main/java/streaming/core · high confidence

Added HSQLStringIndex helper for Spark ML string indexing

A new Scala object, HSQLStringIndex, was added to the Spark ML help utilities. This component provides methods to predict string indices and map indices back to labels, registering corresponding UDFs (such as \_array, \_r, and \_rarray) to handle string-to-index conversions and reverse lookups within the Spark session.

streamingpro-mlsql/src/main/java/org/apache/spark/ml/help · high confidence

Added MLSQL Watcher plugin for monitoring and managing executor state

Introduced the MLSQL Watcher plugin, which provides a new command interface for configuring database connections and cleaning up stale executor data. The plugin registers a 'watcher' command that allows users to add database configurations and set cleanup thresholds for executor and job records. A background scheduler periodically saves and prunes executor state data based on a configurable retention period, ensuring that old records are removed from the database to maintain system performance.

_external/mlsql-watcher/src/main/java/tech/mlsql/plugins/mlsql\watcher · high confidence

Added MLSQL console sink for streaming data output

A new MLSQL console sink has been introduced, allowing streaming data to be written to the console output. This implementation, located in the streaming sink package, enables users to direct their streaming results to the console for debugging or immediate visibility, supporting configuration for row limits and truncation.

streamingpro-mlsql/src/main/java/tech/mlsql/stream · high confidence

Added MLSQLJobCollect utility for job and progress tracking

A new utility class, MLSQLJobCollect, was added to the streamingpro-mlsql module to manage and retrieve job information and progress. This class provides methods to list jobs by owner, retrieve specific job details, and fetch the current progress of a streaming job via the Spark streams API, exposing job metadata and real-time execution status to users.

streamingpro-mlsql/src/main/java/streaming/core/datasource/util · high confidence

Added MLlib feature and Spark ML algorithm base classes

The update introduces new Scala source files in the \streamingpro-mlsql\ module, adding foundational components for machine learning workflows. This includes \BaseAlgorithmEstimator\ and \BaseAlgorithmTransformer\ traits that standardize how algorithms are trained and applied. Specific algorithm implementations for ALS, Linear Regression, and Logistic Regression are added as estimators and transformers, enabling users to train and apply these models within the streaming SQL environment. Additionally, a new \IntTF\ class is added to the \mllib.feature\ package to support integer-based term frequency vectorization.

streamingpro-mlsql/src/main/java/org/apache/spark/ml, streamingpro-mlsql/src/main/java/org/apache/spark/ml/algs, streamingpro-mlsql/src/main/java/org/apache/spark/mllib · high confidence

Added Python resources for ML model training and management

Added Python scripts for training and managing machine learning models, including support for TensorFlow (FC, CNN, AttentionLSTM) and Scikit-learn (GradientBoosting, MultinomialNB, RandomForest, SVC) models. The update also includes daemon scripts for managing worker processes and utility modules for data handling and model saving.

streamingpro-mlsql/src/main/resources-online/python · high confidence

Added Python scripts for machine learning model training and daemon management

Added a suite of Python scripts to the streamingpro-mlsql module to support machine learning workflows. This includes daemon scripts (daemon22.py, daemon23.py, daemon232.py, daemon24.py) for managing worker processes, and various model training scripts for scikit-learn (GradientBoostingClassifier, MultinomialNB, RandomForestClassifier, SVC) and TensorFlow (FC, CNN, AttentionLSTM classifiers). The addition also includes utility modules (mlsql.py, mlsql\_model.py, mlsql\_tf.py, msg\_queue.py) that handle data ingestion, model serialization, and inter-process communication, enabling the execution of distributed training jobs and model evaluation within the streaming environment.

streamingpro-mlsql/src/main/resources-local/python · high confidence

Added Python utility to automate Maven POM configuration for Spark 2.4 and 3.0

A new Python script, convert\_pom.py, has been added to the dev/python directory. This utility automatically modifies Maven pom.xml files to configure Spark profiles for either version 2.4 or 3.0. It handles toggling activation flags and commenting/uncommenting specific XML sections, allowing developers to easily switch between Spark versions across the project.

dev/python · high confidence

Added PythonServer for executing Python code via socket

A new PythonServer class has been added to handle remote Python code execution. This server listens on a socket, receives Python code and environment variables, and executes them using the ArrowPythonRunner, returning the results or errors back to the client.

streamingpro-mlsql/src/main/java/tech/mlsql/ets/python · high confidence

Added Ray integration components for data collection and server configuration

Introduced two new Scala classes to support Ray-based distributed computing: \CollectServerInDriver\ implements a socket server on the driver node to collect and aggregate host and port information from executors, while \DataServer\ provides a case class to represent server configuration including host, port, and timezone.

streamingpro-mlsql/src/main/java/tech/mlsql/ets/ray · medium confidence

Added SQL DSL grammar files for streamingpro-dsl

New ANTLR grammar files (DSLSQL.g4 and DSLSQL\_v2) were added to the streamingpro-dsl module. These files define the syntax for SQL-like commands including LOAD, SAVE, SELECT, INSERT, CREATE, DROP, REFRESH, SET, CONNECT, TRAIN, RUN, PREDICT, REGISTER, UNREGISTER, and INCLUDE, enabling the system to parse and execute these specific SQL operations.

streamingpro-dsl/src/main/resources · high confidence

Added SQL generation and aggregation functions for MLSQL SQL Profiler

New files were added to the mlsql-sql-profiler module to support SQL generation and optimized distinct counting. The change introduces a SQL dialect system (SQLDialect, BasicSQLDialect, LogicalPlanSQL) that converts Spark logical plans into executable SQL strings, including support for MySQL JDBC connections. Additionally, new aggregate functions (BitSetMapping, PreCountDistinct, ReCountDistinct) were added to the catalyst expressions package, enabling efficient re-aggregation of distinct counts using RoaringBitmaps.

external/mlsql-sql-profiler/src/main/java/org/apache/spark/sql/catalyst · high confidence

Added SQL profiler analysis and Z-ordering utilities

The SQL profiler now includes new utility classes to support query analysis and optimization. LPUtils provides logical plan normalization and predicate splitting, while MLSQLAnalyzer parses and extracts load statements. LoadRewriter and LoadUtils handle SQL rewriting for indexed loads. Additionally, ZOrderingBytesUtil and UnsafeAccess provide low-level byte manipulation and memory access for Z-ordering index calculations.

external/mlsql-sql-profiler/src/main/java/tech/mlsql/tool · high confidence

Added SQL profiler indexer implementations for query optimization and dialect handling

The SQL profiler module introduces a suite of new indexer implementations in the \tech.mlsql.indexer.impl\ package to support query optimization and database-specific SQL generation. This includes dialect handlers for H2, Kylin, and MySQL (including a generic MLSQL dialect), which manage SQL quoting, explain plans, and metadata retrieval. The update also adds specific indexers: \ZOrderingIndexer\ for optimizing queries using Z-ordering metadata, \NestedDataIndexer\ for handling nested JSON data structures, \PushdownIndexer\ for applying pushdown optimizations, and a \LinearTryIndexerSelector\ to orchestrate the selection and application of these indexers during query rewriting.

external/mlsql-sql-profiler/src/main/java/tech/mlsql/indexer/impl · high confidence

Added SQL query rewriting and indexing infrastructure

Added new classes and traits in the \tech.mlsql.indexer\ package to support SQL query rewriting and indexing. This includes the \IndexBuilder\ for training and predicting with indexers, the \IndexerQueryReWriterContext\ for rewriting logical plans with indexers, and the \MLSQLIndexer\ trait defining the interface for read/write operations. The \pojo.scala\ file introduces data models like \MlsqlIndexerItem\ and \MlsqlOriTable\ to represent indexer configurations and table metadata. Additionally, \ConsoleRequest\ and \RunScript\ objects were added to handle HTTP requests and script execution, enabling the indexer to interact with external services and execute SQL scripts.

external/mlsql-sql-profiler/src/main/java/tech/mlsql/indexer · high confidence

Added Spark entry point in streamingpro-mlsql

A new Scala file, Main.scala, was added to the org.apache.spark package within the streamingpro-mlsql module. This file defines the package structure for the Spark integration, serving as a placeholder or entry point for the streaming SQL functionality.

streamingpro-mlsql/src/main/java/org/apache/spark · medium confidence

Added SparkAgent utility for internal Spark API access

A new Scala file, SparkAgent.scala, was added to the mlsql-sql-profiler module. This object provides a set of helper methods to access package-private or internal Spark APIs, such as retrieving SQL configuration, logical plans, and DataFrames, as well as utility functions for date formatting and decimal type handling.

external/mlsql-sql-profiler/src/main/java/org/apache/spark · medium confidence

Added SparkOperationUtil for streaming job execution and context management

A new utility class, SparkOperationUtil, was added to the streaming module. This trait provides helper methods for managing Spark runtime contexts, executing SQL or script code (both synchronously and asynchronously), and waiting for job states. It includes utilities for generating execution contexts, managing job information, and handling cleanup of temporary resources like the embedded Derby database and temporary directories.

streamingpro-mlsql/src/main/java/org/apache/spark/streaming · high confidence

Added ViewCatalyst for SQL view metadata management

A new Scala trait and implementation, ViewCatalyst, has been added to the SQL profiler module. This component manages a mapping of view names to table metadata (MlsqlOriTable) using a thread-local context, allowing the system to register, retrieve, and list SQL view definitions and their associated database information.

external/mlsql-sql-profiler/src/main/java/tech/mlsql/sqlbooster · high confidence

Added ZooKeeper client and path utilities

New classes have been added to the streamingpro-commons module to support ZooKeeper integration. This includes ZKClient for managing data listeners and configuration retrieval, ZKConfUtil for handling serialization and path operations, Path for parsing and managing ZooKeeper node paths, and ZkRegister (Scala) to register service addresses as ephemeral nodes in ZooKeeper.

streamingpro-commons/src/main/java/streaming/common/zk · high confidence

Added custom output writers for TensorFlow processing

Users can now write data to JSON and Parquet files during TensorFlow processing tasks. The new JsonOutputWriter and ParquetWriter classes implement Spark's OutputWriter interface, enabling structured data export in these specific formats for downstream model consumption.

streamingpro-mlsql/src/main/java/tech/mlsql/ets/tensorflow/files · high confidence

Added if/else/elif/then/finally control flow commands

The streamingpro-mlsql module now supports conditional execution of SQL statements via new command classes: IfCommand, ElifCommand, ElseCommand, ThenCommand, and FiCommand. These components implement branching logic, allowing users to execute specific SQL blocks based on evaluated conditions or to handle default/final execution paths within the ETL pipeline.

streamingpro-mlsql/src/main/java/tech/mlsql/ets/ifstmt · high confidence

Added local Spark application examples for MLSQL

Two new example applications, LocalSparkServiceApp and LocalSparkApp, have been added to the MLSQL streamingpro module. LocalSparkServiceApp demonstrates a local Spark service configuration with REST enabled, while LocalSparkApp provides a local Spark application setup with specific runtime hooks and script execution parameters. These examples illustrate how to configure and run MLSQL in local mode for development and testing purposes.

streamingpro-mlsql/src/main/java/tech/mlsql/example · high confidence

Added metadata classes for MLlib algorithms

Introduced new case classes in the \streaming/dsl/mmlib/algs/meta\ package to define metadata structures for various MLlib algorithms, including TFIDF, Word2Vec, Word2Index, Scale, OutlierValue, MinMaxValue, StandardScaler, Discretizer, Word2Array, and MapValues. These classes encapsulate training parameters, function references, and other algorithm-specific data required for processing.

streamingpro-mlsql/src/main/java/streaming/dsl/mmlib/algs/meta · high confidence

Added new ETS commands for Delta Lake, HDFS, Kafka, and resource management

This update introduces several new commands within the ETS (External Table Service) module. Users can now manage Delta Lake tables with \!delta\ commands for compaction, history, and table listing. HDFS operations are supported via \!hdfs\ and \!fs\ commands, including a new \getmerge\ utility. Kafka integration is enhanced with \!kafkaTool\ for sampling data, inferring schemas, and managing offsets. Additionally, the \!engine\ command allows dynamic adjustment of Spark executor resources (add/remove/set), and \!model\ provides a way to view model training history. These changes expand the toolkit available for data engineering and model management tasks.

streamingpro-mlsql/src/main/java/tech/mlsql/ets · high confidence

Added sample data for MLlib algorithms

New sample data files have been added to the \streamingpro-mlsql/src/main/resources-local/data/mllib\ directory to support machine learning algorithms. These include \sample\_movielens\_ratings.txt\ for the Alternating Least Squares (ALS) model, \test.data\ for ALS testing, \gmm\_data.txt\ for Gaussian Mixture Models, \kmeans\_data.txt\ for K-Means clustering, \pagerank\_data.txt\ for PageRank calculations, \pic\_data.txt\ for graph processing, \lpsa.data\ for Ridge regression, \sample\_article.txt\ for text processing, and \sample\_binary\_classification\_data.txt\ for binary classification tasks. These files provide ready-to-use datasets for testing and demonstrating these specific MLlib capabilities within the streamingpro-mlsql module.

streamingpro-mlsql/src/main/resources-local/data · high confidence

Added scripts to upload and download MLsql releases to/from Aliyun OSS

New Python scripts have been added to the \dev/mlsqltestssupport/aliyun\ directory to manage release artifacts in Aliyun Object Storage Service (OSS). The \upload\_release.py\ script uploads release tarballs to the \mlsql-release-repo\ bucket, while \download\_release.py\ retrieves them. Both scripts rely on environment variables (\MLSQL\_RELEASE\_TAR\, \AK\, \AKS\) for configuration and authentication, with the download script specifically targeting the internal endpoint (\oss-cn-hangzhou-internal.aliyuncs.com\) for optimized access.

dev/mlsqltestssupport/aliyun · high confidence

Added shell utility functions for test support

The \dev/mlsqltestssupport\ package was introduced, providing a new \shellutils\ module that offers helper functions for executing shell commands and managing the file system. This includes utilities to run commands with error handling, remove directories or files, and locate executable programs in the system PATH, which will be used to support test environments.

dev/mlsqltestssupport · high confidence

Added utility classes for string processing and debugging

Added new utility classes in the streamingpro-commons module to support string processing and debugging: DefaultShortNameMapping for mapping short names to implementation classes, PunctuationUtils for identifying and filtering punctuation characters, SpecificPortGenerator for managing specific ports, TimeRecord for tracking execution time in debug mode, and UnicodeUtils for converting strings to Unicode and filtering Chinese characters.

streamingpro-commons/src/main/java/streaming/common · medium confidence

Adds SQL query pushdown optimization for H2, Kylin, and MySQL data sources

The optimizer now supports pushing down SQL queries to external data sources, specifically H2, Kylin, and MySQL. This new capability allows the system to execute more processing on the database side rather than in memory, which can significantly improve performance for large datasets. The implementation includes new source info classes for each database type, a central \Pushdown\ optimizer that identifies and replaces logical plan nodes with optimized database scans, and a \Pushdownable\ trait defining the interface for query pushdown support.

external/mlsql-sql-profiler/src/main/java/org/apache/spark/sql/optimizer · high confidence

Adds default configuration files for local and online environments

New configuration files are introduced for the streamingpro-mlsql module, providing default settings for both local and online deployment modes. This includes application.yml, log4j2.properties, and mlsql-version-info.properties in both resources-local and resources-online directories, alongside Hive configuration (hive-site.xml) and strategy definitions (strategy.v2.json and examples). These files establish baseline logging, datasource, and runtime versioning configurations for the application.

streamingpro-mlsql/src/main/resources-local · high confidence

Byzer-lang project initialization and documentation

The repository is initialized with the Byzer-lang project, including a comprehensive README, Apache 2.0 license, and CI configuration for Scala 2.11.8 and JDK 8. The project provides a SQL-like language for data pipelines and AI, featuring a 'Everything is a table' design protocol. The README includes architecture diagrams, online trial links, installation instructions, code examples, and contribution guidelines.

(repo-wide) · high confidence

Dynamic executor management and status reporting

A new Scala class, SparkInnerExecutors, is introduced to expose internal Spark scheduler state (executor cores, memory, and count) via reflection, enabling a new SparkDynamicControlExecutors utility that allows users to dynamically request or kill executors with a 60-second timeout. This adds the ability to programmatically scale the cluster's executor count and query resource status, addressing previous issues where null references could occur when accessing the executor data map.

streamingpro-spark-common/src/main/java/org · medium confidence

Expanded HDFS shell commands and file operations

The HDFS integration now supports a broader set of file system operations, including merging files (getmerge), deleting files and directories (rm, rmdir, rmr), and listing directory contents (ls, lsr). A new \FSGetmerge\ class and \WowFsShell\ shell implementation enable these capabilities, allowing users to perform common HDFS tasks directly through the streamingpro-mlsql interface.

streamingpro-mlsql/src/main/java/tech/mlsql/ets/hdfs · high confidence

Introduce Dispatcher and StreamingApp entry points for job execution and lifecycle hooks

Added new Scala files, Dispatcher.scala and StreamingApp.scala, which establish the core entry points for the streaming platform. The Dispatcher object handles job configuration loading from classpath, HTTP/HTTPS URLs, or HDFS, and initializes the StrategyDispatcher. The StreamingApp object serves as the main entry point, registering built-in and user-defined platform lifecycle hooks (via the streaming.platform\_hooks parameter) before running the platform. This change enables users to define custom lifecycle hooks that are persisted and executed during application startup.

streamingpro-commons/src/main/java/streaming/core · high confidence

Introduce MLSQL Watcher for Spark runtime monitoring

Added the MLSQL Watcher plugin, which monitors Spark runtime metrics and stores them in a database. The change introduces a new directory at external/mlsql-watcher containing a SQL schema (db.sql) for tables w\_executor, w\_executor\_job, and w\_kv, along with Scala source files (DataCompute.scala, DBAction.scala, PluginDB.scala, tables.scala) that compute and persist executor and job statistics such as shuffle read/write, disk spills, and GC time.

external/mlsql-watcher · high confidence

Introduce MLSQLPlatformLifecycle interface for runtime hooks

A new Scala trait, MLSQLPlatformLifecycle, has been added to the runtime module. This interface defines four lifecycle hooks—beforeRuntime, afterRuntime, beforeDispatcher, and afterDispatcher—that allow external code to execute logic before and after specific runtime and dispatcher phases.

streamingpro-commons/src/main/java/tech/mlsql/runtime · high confidence

Introduce cluster management components for distributed TensorFlow processing

Added four new Scala files in the ML cluster package to support distributed TensorFlow processing: DataManager handles writing local data files (Parquet or JSON) for each algorithm index; LocalDirectoryManager manages task directories and downloads Python projects; PortManager allocates and releases network ports for cluster nodes; and TFContext manages the lifecycle of the TensorFlow cluster, including worker proxy communication, status reporting, and process management. These components enable the system to distribute TensorFlow workloads across a cluster.

streamingpro-mlsql/src/main/java/tech/mlsql/ets/ml · high confidence

Introduce cluster proxy and elastic resource allocation APIs

The streamingpro-cluster module now includes a new proxy application and REST controllers for managing backend instances and elastic resources. Users can now add, list, update, and remove backends via the BackendController, and manage ECS resource pools and elastic monitors via the EcsResourceController. The system supports dynamic resource allocation with strategies like JobNumAware and ResourceAware, allowing the cluster to automatically scale MLSQL instances based on tags and load.

streamingpro-cluster/src · high confidence

Introduce new auto-suggestion engine for MLSQL

Added a new auto-suggestion system for MLSQL, implemented in the \external/mlsql-autosuggest/src/main\ directory. This feature provides code completion and syntax suggestions based on the SQL grammar, supporting features such as table and column name completion, function suggestions, and metadata integration. The implementation includes a new \AutoSuggester\ class and supporting classes for tokenization, parsing, and suggestion generation, enabling users to receive intelligent code completions while writing MLSQL scripts.

external/mlsql-autosuggest/src/main · high confidence

Introduce new session management and PS cluster infrastructure

Added new classes for session management including MLSQLSession, SessionManager, and SparkSessionCacheManager to handle user sessions and SparkSession caching. Additionally, introduced PS (Parameter Server) cluster infrastructure with PSDriverBackend, PSExecutorBackend, and PSDriverEndpoint to support distributed machine learning workloads.

streamingpro-core · high confidence

New ETS plugins for data manipulation and system introspection

The ETS (External Table Source) plugin module introduces several new capabilities for data processing and system management. Users can now repartition tables with configurable hash or range strategies, expand JSON strings into multiple columns, and perform pivot operations. Additionally, the module adds commands to retrieve the last executed command or table name, save binary data as files, and query schema or streaming job information. The suite also includes a syntax analyzer for SQL and a connector for sending messages to Feishu webhooks.

external/mlsql-ets/src/main/java/tech/mlsql/plugins/ets · high confidence

New Python controller components for MLSQL

Added new Java and Scala classes to the Python controller module, including JavaDoc.java, PythonApp.scala, PythonInclude.scala, and quill\_model.scala. These files introduce the core logic for executing Python scripts within MLSQL, supporting both executor and driver modes, and enabling the inclusion of external Python code via the !pyInclude command.

external/python-controller/src/main/java/tech/mlsql/plugins/app · high confidence

New PythonAlg module for training, batch prediction, and API prediction with MLflow support

A new PythonAlg module has been added to the streamingpro-mlsql component, enabling users to train, batch predict, and perform API-based predictions using external Python scripts. The implementation includes dedicated classes for handling Python project loading, command generation, and resource management, supporting both standard Python scripts and MLflow-based projects. Users can now integrate Python-based machine learning workflows directly into their MLSQL pipelines, with support for Conda environments and local data handling.

streamingpro-mlsql/src/main/java/streaming/dsl/mmlib/algs/python · high confidence

New REST API endpoints for model prediction and SQL execution

The streamingpro-mlsql module introduces new REST controllers to expose MLSQL capabilities via HTTP. The \RestController\ adds a \/run/script\ endpoint that allows users to execute MLSQL scripts, supporting both synchronous and asynchronous execution modes with callback headers and retry logic. Additionally, the \RestPredictController\ introduces \/model/predict\ and \/compute\ endpoints, enabling users to perform model predictions on vector, string, or row data types, and execute arbitrary SQL queries via the \/compute\ interface.

streamingpro-mlsql/src/main/java/streaming/rest · high confidence

New REST datasource page strategies for pagination

The REST datasource now supports three distinct pagination strategies for paginated API responses: AutoIncrement, Offset, and Default. Users can configure the \config.page.values\ parameter to select the strategy—\auto-increment\ for simple incrementing page numbers, \offset\ for offset-based pagination with configurable increments, or the default strategy which extracts page values via JSONPath expressions. This allows more flexible handling of different REST API pagination patterns.

streamingpro-mlsql/src/main/java/tech/mlsql/datasource/helper · high confidence

New SQL Profiler plugin with query optimization and indexing capabilities

The SQL Profiler plugin is introduced, adding new capabilities for SQL analysis and optimization. Users can now generate SQL from load/select statements, rewrite queries using indexers (including Z-ordering and pushdown indexers), and access a profiler command for explaining query plans. The plugin also includes a result render that automatically applies query optimizations when enabled, and provides fine-grained access control for the profiler commands.

external/mlsql-sql-profiler/src/main/java/tech/mlsql/plugins/sql · medium confidence

New configuration files for Byzer server and tools logging and runtime modes

Added new configuration files to support Byzer's server and all-in-one runtime modes, along with dedicated logging configurations. The \byzer.properties\ file introduces the \byzer.server.mode\ parameter to switch between 'all-in-one' and 'server' modes, with example files (\byzer.properties.all-in-one.example\, \byzer.properties.server.example\, \byzer.properties.server.yarn-client.example\) providing templates for different deployment scenarios. Additionally, \byzer-server-log4j2.properties\ and \byzer-tools-log4j2.properties\ are introduced to manage logging for the server and tools respectively, allowing users to control log levels and output destinations.

conf · high confidence

New data models for MLSQL job and resource rendering

A new Scala file, Protocal.scala, was added to the mlsql.render.protocal package. It introduces case classes for MLSQLScriptJobGroup, MLSQLScriptJob, MLSQLResourceRender, and MLSQLShufflePerfRender. These models expose detailed metrics about job execution (active, completed, failed, and killed tasks) and resource usage (executor counts, memory, and shuffle performance), enabling the rendering layer to display comprehensive job and resource information to users.

streamingpro-spark-common/src/main/java/tech · medium confidence

New data source implementations for file systems, Kafka, and REST APIs

Added new data source implementations for handling various data formats and sources. This includes a generic file system source (CustomFS) for object stores, an ad-hoc Kafka source (MLSQLAdHocKafka), a binary file source (MLSQLBinaryFile), a binlog source (MLSQLBinlog), a multi-table Delta Lake sink (MLSQLMultiDelta), a REST API source (MLSQLRest) with pagination and retry support, a stream batch processor (MLSQLStreamBatch), an unstructured file source (MLSQLUnStructured), and a generic everything source (MLSQLEverything). These additions expand the range of data sources and sinks available for data ingestion and processing.

streamingpro-mlsql/src/main/java/tech/mlsql/datasource/impl · high confidence

New dev scripts for environment checks, building, and testing

The dev directory now includes a suite of new shell scripts to streamline the build and test workflow. Environment validation is handled by check-env.sh, which orchestrates checks for the operating system, Java version, and port availability. The build process is managed by package.sh and package.cmd, which invoke Maven to compile the project and can optionally upload release artifacts to Aliyun OSS. Testing is supported by run-test.sh, which allows users to execute unit or integration tests against specific Spark versions (3.0 or 3.3). Additionally, helper scripts like change-scala-version.sh and switch.sh are provided to manage Scala and Spark version configurations.

dev · high confidence

New feature engineering and batch prediction components

Added three new components to the Spark ML feature library: DiscretizerFeature, which provides utilities for discretizing continuous features into buckets; IntTF, a transformer that maps sequences of terms to their term frequencies using the hashing trick; and PythonBatchPredictDataSchema, which defines the data schema for Python-based batch prediction workflows.

streamingpro-mlsql/src/main/java/org/apache/spark/ml/feature · high confidence

New feature engineering components for text and numeric data

Added new feature engineering modules to the streamingpro-mlsql library, including BaseFeatureFunctions, DiscretizerIntFeature, DoubleFeature, and StringFeature. These components provide capabilities for handling text analysis (tokenization, stopword filtering, n-grams, TF-IDF), numeric scaling (min-max, log, outlier removal), and vectorization (PCA, assembly).

streamingpro-mlsql/src/main/java/streaming/dsl/mmlib/algs/feature · high confidence

New health check and cleanup plugins for MLSQL

Added a new 'healthy' plugin that provides a health check endpoint and automated cleanup of temporary tables. The plugin registers a controller at /plugins/healthy to report Spark context status and trigger a delayed shutdown, and registers request cleaners to optionally shut down the Spark context and drop temporary views when the context stops or requests complete.

external/mlsql-healthy/src/main/java/tech/mlsql/plugins/healthy · medium confidence

New parameter configuration for Python-based ML algorithms

The \streamingpro-mlsql\ module introduces new parameter definitions for Python-based machine learning algorithms. The \BaseParams\ trait adds \evaluateTable\ for model evaluation and \keepVersion\ to control model versioning. The \SQLPythonAlgParams\ trait adds configuration for Python execution, including \enableDataLocal\ and \dataLocalFormat\ for data handling, alongside parameters for Python script paths, Kafka connections, and entry points for fit, batch prediction, and API prediction. This enables the system to manage Python algorithm projects and their specific execution parameters.

streamingpro-mlsql/src/main/java/streaming/dsl/mmlib/algs/param · high confidence

New shared object pooling infrastructure

The \streaming/core/shared\ package now includes a new shared object management system. This introduces \SharedObjManager\ for managing global state, along with a new \pool\ sub-package containing \BigObjPool\, \DicPool\, and \ForestPool\ classes. These components provide a thread-safe mechanism for storing, retrieving, and clearing large objects and sets of strings, likely to optimize resource usage across streaming operations.

streamingpro-mlsql/src/main/java/streaming/core/shared · high confidence

New utility classes for data processing and environment variable substitution

Added several new utility classes in the tech.mlsql.tool package. ScriptEnvDecode enables replacing variables in the format :variableName with environment variables, addressing issues with Python configuration. MasterSlaveInSpark introduces a new mechanism to start servers in the driver and data servers in tasks, reporting their host:port to Python jobs. TarfileUtil and SparkTarfileUtil provide utilities for walking HDFS directories and creating/extracting TAR files. ScalaObjectReflect offers a method to find object methods via reflection, and YieldByteArrayOutputStream provides a byte array output stream that triggers a callback when full.

streamingpro-mlsql/src/main/java/tech/mlsql/tool · high confidence

Behavioural changes

Add default loader plugin and multi-bucket filesystem support

A new DefaultLoaderPlugin was added to allow users to define custom plugins that execute before load, save, and HDFS/FS commands, enabling pre-processing of dataframes by filtering rows and selecting or dropping specific columns. Additionally, support for multi-bucket filesystems was introduced, allowing configuration of path prefixes and Hadoop/Spark configurations for various data formats (CSV, JSON, Parquet, etc.) when operating in multi-bucket mode.

streamingpro-mlsql/src/main/java/tech/mlsql/plugin · medium confidence

Added .gitkeep file to examples directory

A .gitkeep file was added to the examples directory. This ensures the directory is tracked by version control, preventing the directory from being ignored or removed when empty.

examples · medium confidence

Added automatic cleanup for REST datasource temp files and Spark session listeners

Two new cleaner components have been introduced to prevent resource leaks. The RestDataSourceTempFileCleaner now periodically deletes temporary data directories associated with REST datasource requests, controlled by the enableRestDataSourceRequestCleaner flag. Additionally, the SessionCleaner has been added to clean up Spark session listeners, addressing a memory leak issue in Spark 3.1.1. Both cleaners respect asynchronous execution contexts to avoid interfering with running tasks.

external/mlsql-sql-profiler-30/src/main/java/tech/mlsql/plugins/cleaner · medium confidence

Added new template files for REST and SQL UI interfaces

The template directory now includes three new Velocity templates: \index.vm\ provides a basic React.js-based entry point, while \sqlui.vm\ and \sqlui-result.vm\ supply the HTML structure for the Spark SQL query tool interface and its results page, respectively.

template · high confidence

Introduce PlatformManager and RuntimeOperator abstractions

The platform strategy layer now introduces new abstractions for managing runtime lifecycle and execution. PlatformManager serves as the central coordinator, handling the initialization of REST and Thrift servers, ZK registration, and the dispatching of jobs via a StrategyDispatcher. It also manages lifecycle callbacks (before/after runtime and dispatcher phases) and exposes a listener mechanism for event processing. Concurrently, the new RuntimeOperator trait and the updated StreamingRuntime interface define the contract for runtime instances, including methods for starting, destroying, and awaiting termination, as well as configuring runtime info. This refactors how the streaming engine initializes and manages its execution environment.

streamingpro-commons/src/main/java/streaming/core/strategy/platform · high confidence

Introduce pluggable exception rendering and request cleaning mechanisms

The system now supports pluggable exception rendering and request cleaning through new manager classes. Exception rendering is handled by a new \ExceptionRenderManager\ that iterates through registered \ExceptionRender\ implementations (such as \ArrowExceptionRender\ and \DefaultExceptionRender\) to format error messages, falling back to a default handler. Similarly, \RequestCleanerManager\ is introduced to iterate through registered \RequestCleaner\ implementations to perform cleanup tasks. These changes restructure how exceptions are formatted and how requests are cleaned, providing a more modular approach to error handling and request processing.

_streamingpro-core/src/main/java/tech/mlsql/runtime/plugins/exception\_render, streamingpro-core/src/main/java/tech/mlsql/runtime/plugins/request\cleaner · medium confidence

Refactor and add strategy components in streamingpro-commons

The \streamingpro-commons\ module now includes newly added \JobStrategy\ and \ParamsValidator\ traits, alongside a refactored \SparkStreamingRefStrategy\ (renamed from \SparkStreamingStrategy\) that enforces parameter validation before processing. Additionally, \DebugTrait\ was moved to this module, and all files were updated with the Apache 2.0 license header.

streamingpro-commons/src/main/java/streaming/core/strategy · medium confidence

Removed TestOutputStream helper class

The TestOutputStream class, previously located in the org.apache.spark.streaming package, has been removed from the codebase. This eliminates a test utility used to capture streaming output for verification purposes.

src/main/java/org · high confidence

Removed legacy streaming core components and utility classes

Deleted the legacy streaming application entry points (LocalStreamingApp, StreamingApp) and associated runtime management (PlatformManager, SparkStreamingRuntime). Also removed utility classes (NumberUtil, ParamsUtil) and various processing components (RDDPrintOutputCompositor, KafkaStreamingCompositor, JSONTableCompositor, SQLCompositor, NginxParser). These files were part of the original streaming engine architecture and are no longer present in the codebase.

src/main/java/streaming · high confidence

Renamed web console to Byzer Web Console

The web console title has been updated to 'Byzer Web Console' in the HTML header, and the associated CSS and JavaScript assets have been updated to reflect the new branding and component structure.

streamingpro-mlsql/src/main/resources-online/streamingpro · high confidence

Updated Byzer Web Console UI assets

The static assets for the Byzer Web Console have been updated, including the main HTML entry point, CSS styles, and JavaScript bundles. This reflects the renaming of the web console from 'MLSQL' to 'Byzer' and includes updated styles for the query editor, job list, and table components.

streamingpro-mlsql/src/main/resources-local/streamingpro · high confidence

Updated SQL parser to support new SQL and DSL syntax

The SQL parser has been regenerated using ANTLR 4.7.1, incorporating grammar updates that add support for new SQL commands and DSL keywords. The updated parser now recognizes and parses additional keywords such as 'set', 'connect', 'train', 'run', 'predict', 'register', 'unregister', 'include', 'options', 'overwrite', 'append', 'errorIfExists', and 'ignore', enabling the system to process a broader range of SQL and domain-specific language statements.

streamingpro-dsl/src/main/java · high confidence

Updated Spark MLlib imports to Spark ML

Replaced deprecated Spark MLlib imports with the newer Spark ML API. In the vectorize source files, the code now imports \org.apache.spark.ml.linalg.Vectors\ instead of \org.apache.spark.mllib.linalg.Vectors\. In the REST JSON source files, the implementation was updated to use \org.apache.spark.sql.execution.datasources.json.{JSONOptions, JacksonParser}\ and \org.apache.spark.sql.execution.datasources.json.InferSchema.infer\ instead of the previous \InferSchema\ and \JacksonParser\ usage, aligning the data source implementation with current Spark SQL standards. These changes affect both local and online resource configurations.

streamingpro-mlsql/src/main/resources-local/models, streamingpro-mlsql/src/main/resources-local/source, streamingpro-mlsql/src/main/resources-online/source · high confidence

Web console title updated to 'Byzer Web Console'

The HTML templates for the local and online REST endpoints have been updated to reflect the new product name. Users will now see 'Byzer Web Console' as the title in their browser tabs and headers, aligning the interface with the recent renaming of the web console.

streamingpro-mlsql/src/main/resources-local/rest, streamingpro-mlsql/src/main/resources-online/rest · high confidence

Test coverage

Add UDF test suite for Scala, Python, and Java script UDFs; Add test coverage for job management and killing; Added BasicSparkOperation test base class; Added MySQL integration tests; Added Python code templates for ML model training and prediction; Added REST API test suite for script execution and authentication; Added SQL test cases for ML and NLP features; Added SQL test cases for ML and data processing modules; Added TFSpec test for TensorFlow integration; Added integration test infrastructure for Byzer cluster containers; Added integration test suite for Byzer scripts; Added integration test utilities for Docker and exception handling; Added integration tests for SQL execution, data transformation, and SSB queries; Added integration tests for core, Delta, and UDF features; Added integration tests for streaming data sources; Added performance testing utilities for MLSQL script execution; Added test configuration files for Byzer integration tests; Added test coverage for ANTLRv4 parsing and client mode operations; Added test coverage for CacheExt, Text processing, TreeBuildExt, and TextModule; Added test for stream sub-batch query execution; Added test infrastructure and helper utilities for the streaming core; Added test server implementations for integration testing; Added test utilities for streaming processing; Added tests for Python ML algorithm integration; Added tests for SQLRateSampler validation; Added unit tests for ByzerConfig, Scala reflection, and test utilities; Added unit tests for MLSQL auto-suggest components; Added unit tests for MLSQL configuration, DSL parsing, and template evaluation; Added unit tests for MLSQL data source integrations; Added unit tests for MLSQLRest data source; Added unit tests for MLlib algorithms; Added unit tests for Oracle and MySQL upsert builders; Added unit tests for Python-based ML algorithm execution; Added unit tests for SQL pushdown and byte utility functions; Added unit tests for expression evaluation and stream batch query logic; Added unit tests for template merge functionality.

Dependencies

Added Maven build configurations for new MLSQL modules

Added new Maven build configurations for the mlsql-autosuggest, mlsql-ets, mlsql-healthy, mlsql-mysql-store, mlsql-sql-profiler, mlsql-watcher, and python-controller modules, as well as the streamingpro-assembly, streamingpro-cluster, and streamingpro-commons modules. These configurations define dependencies on Spark 3.3.0, Scala 2.12.15, and various internal MLSQL components, enabling the build system to compile and package these specific parts of the MLSQL project.

(dependencies) · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

Score

  • CAI 35 → 44 (+9.1)
  • Rubric changed (rubric-2026.08.15 → rubric-2026.09.15) — scores are not directly comparable.

Lenses

  • Code Health 44 → 85 (+40.7)
  • Architecture 96 → 97 (+1.2)
  • Maturity 48 → 41 (-7.4)
  • Readiness 11 → 50 (+39.0)
  • Security 94 → 51 (-43.1)
  • Accessibility 37 (new)

Resolved (39)

  • Coverage not measured — test suite did not build
  • Dimension evaluation failed
  • Duplicated block (11 lines × 2) (streamingpro-core/src/main/java/tech/mlsql/runtime/kvstore/LevelDBTypeInfo.java)
  • Duplicated block (11 lines × 4) (external/mlsql-autosuggest/src/main/java/tech/mlsql/autosuggest/app/MysqlType.java)
  • Duplicated block (11 lines × 5) (streamingpro-dsl/src/main/java/streaming/dsl/parser/DSLSQLParser.java)
  • Duplicated block (117 lines × 3) (streamingpro-dsl/src/main/java/streaming/dsl/parser/DSLSQLParser.java)
  • Duplicated block (13 lines × 3) (streamingpro-dsl/src/main/java/streaming/dsl/parser/DSLSQLParser.java)
  • Duplicated block (15 lines × 3) (streamingpro-dsl/src/main/java/streaming/dsl/parser/DSLSQLParser.java)
  • Duplicated block (15 lines × 4) (streamingpro-dsl/src/main/java/streaming/dsl/parser/DSLSQLParser.java)
  • Duplicated block (15 lines × 9) (streamingpro-dsl/src/main/java/streaming/dsl/parser/DSLSQLParser.java)
  • Duplicated block (154 lines × 2) (external/mlsql-sql-profiler/src/main/java/tech/mlsql/tool/ZOrderingBytesUtil.java)
  • Duplicated block (17 lines × 2) (streamingpro-commons/src/main/java/tech/mlsql/tool/OrderedProperties.java)
  • Duplicated block (224 lines × 10) (streamingpro-dsl/src/main/java/streaming/dsl/parser/DSLSQLBaseListener.java)
  • Duplicated block (253 lines × 3) (streamingpro-dsl/src/main/java/streaming/dsl/parser/DSLSQLParser.java)
  • Duplicated block (254 lines × 3) (streamingpro-dsl/src/main/java/streaming/dsl/parser/DSLSQLParser.java)
  • Duplicated block (254 lines × 3) (streamingpro-dsl/src/main/java/streaming/dsl/parser/DSLSQLParser.java)
  • Duplicated block (255 lines × 3) (streamingpro-dsl/src/main/java/streaming/dsl/parser/DSLSQLParser.java)
  • Duplicated block (255 lines × 3) (streamingpro-dsl/src/main/java/streaming/dsl/parser/DSLSQLParser.java)
  • Duplicated block (256 lines × 3) (streamingpro-dsl/src/main/java/streaming/dsl/parser/DSLSQLParser.java)
  • Duplicated block (257 lines × 3) (streamingpro-dsl/src/main/java/streaming/dsl/parser/DSLSQLParser.java)
  • …and 19 more

New (474)

  • AddProject.findPoint (cognitive 21) (external/mlsql-sql-profiler-30/src/main/java/org/apache/spark/sql/catalyst/sqlgenerator/LogicalPlanSQL.scala)
  • AddProject.findPoint (cognitive 21) (external/mlsql-sql-profiler/src/main/java/org/apache/spark/sql/catalyst/sqlgenerator/LogicalPlanSQL.scala)
  • AddProject.findPoint (cyclomatic 18) (external/mlsql-sql-profiler-30/src/main/java/org/apache/spark/sql/catalyst/sqlgenerator/LogicalPlanSQL.scala)
  • AddProject.findPoint (cyclomatic 18) (external/mlsql-sql-profiler/src/main/java/org/apache/spark/sql/catalyst/sqlgenerator/LogicalPlanSQL.scala)
  • AddSubqueryAlias.findPoint (cognitive 25) (external/mlsql-sql-profiler-30/src/main/java/org/apache/spark/sql/catalyst/sqlgenerator/LogicalPlanSQL.scala)
  • AddSubqueryAlias.findPoint (cognitive 25) (external/mlsql-sql-profiler/src/main/java/org/apache/spark/sql/catalyst/sqlgenerator/LogicalPlanSQL.scala)
  • AddSubqueryAlias.findPoint (cyclomatic 18) (external/mlsql-sql-profiler-30/src/main/java/org/apache/spark/sql/catalyst/sqlgenerator/LogicalPlanSQL.scala)
  • AddSubqueryAlias.findPoint (cyclomatic 18) (external/mlsql-sql-profiler/src/main/java/org/apache/spark/sql/catalyst/sqlgenerator/LogicalPlanSQL.scala)
  • AttributeExtractor.extractor (cognitive 29) (external/mlsql-autosuggest/src/main/java/tech/mlsql/autosuggest/AttributeExtractor.scala)
  • AutoSuggestContext._suggest (cyclomatic 17) (external/mlsql-autosuggest/src/main/java/tech/mlsql/autosuggest/AutoSuggestContext.scala)
  • AutoSuggester.fillParserTransitionLabels (cognitive 22) (external/mlsql-autosuggest/src/main/java/com/intigua/antlr4/autosuggest/AutoSuggester.java)
  • AutoSuggester.isParseableWithAddedToken (cognitive 28) (external/mlsql-autosuggest/src/main/java/com/intigua/antlr4/autosuggest/AutoSuggester.java)
  • ByzerConfigCLI.execute (cognitive 57) (streamingpro-commons/src/main/java/tech/mlsql/tool/ByzerConfigCLI.java)
  • ByzerConfigCLI.execute (cyclomatic 23) (streamingpro-commons/src/main/java/tech/mlsql/tool/ByzerConfigCLI.java)
  • Change coupling: LibIncludeSource.scala ↔ ScriptIncludeSource.scala (streamingpro-core/src/main/java/tech/mlsql/dsl/includes/LibIncludeSource.scala)
  • ClassTooLong: DSLSQLParser (streamingpro-dsl/src/main/java/streaming/dsl/parser/DSLSQLParser.java)
  • DSLSQLParser.sql (cognitive 165) (streamingpro-dsl/src/main/java/streaming/dsl/parser/DSLSQLParser.java)
  • DSLSQLParser.sql (cyclomatic 106) (streamingpro-dsl/src/main/java/streaming/dsl/parser/DSLSQLParser.java)
  • Dependency hygiene PARTLY measured — Maven/Gradle declarations read, no dependency graph resolved
  • DistributedTensorflow.train (cognitive 22) (streamingpro-mlsql/src/main/java/tech/mlsql/ets/tensorflow/DistributedTensorflow.scala)
  • …and 454 more

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

byzer-org/byzer-lang was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 27 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit c772189a2089424e873a94dbf2ced7cc23ad7c21 — the exact code this score is about.
  • Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-d00c643c3f66.