databricks/LearningSparkV2
54.5
Adequate · 27 September 2026
701
lines of production code
Scala
with Python, Java
4
measurements over time
What this system is
Features
Add MLflow project example for training a Random Forest model
A new mlflow-project-example directory has been added, containing a complete MLflow project that trains a Spark Random Forest regressor on Airbnb data. The example includes a MLproject definition, a conda.yaml environment file specifying Python 3.7, pandas 0.24, and pyspark 3.0.0, along with a train.py script that logs parameters, metrics (RMSE, R2), and artifacts (feature importance CSV) to MLflow.
mlflow-project-example · high confidence
Add sample datasets and documentation for Spark learning examples
Added new sample datasets and associated documentation to the learning-spark-v2 directory. This includes a Spark README with build and usage instructions, and several new data files: a blogs dataset (JSON), a MNM dataset, a US population dataset, a web access log dataset, and a CCTV video dataset with labels. Additionally, flight data is provided in both CSV (departuredelays.csv, airport codes) and Avro formats, along with summary CSV files for years 2010–2013. These files support standalone Spark applications and tutorials.
databricks-datasets · high confidence
Added Apache web server sample log file
A new sample Apache web server access log file (web1\_access\_log\_20190715-064001.log) has been added to the chapter7/data directory. This file contains 100,000 lines of simulated HTTP request data, providing a realistic dataset for testing log parsing or analytics workflows.
chapter7/data · high confidence
Added Chapter 3 Scala example and data for blog analysis
A new Scala example (Example3\_7) and its associated JSON data file (blogs.json) have been added to the chapter3/scala directory. The example demonstrates reading a JSON file with a predefined schema, displaying the DataFrame schema, and performing column and expression operations such as filtering for high-traffic blog posts. A README.md file was also added to provide build and run instructions for this chapter's Scala code.
chapter3/scala · high confidence
Added Java examples for Chapter 6
Added Java source files (Example6\_3, Person, PersonExample, Usage, UsageCost) and a README for building and running the examples. These files demonstrate creating and manipulating Spark Datasets using Java beans and explicit encoders, providing concrete code samples for Chapter 6.
chapter6/java · high confidence
Added Python examples and test data for Chapter 2
Added Python source files and test data for the M&M example in Chapter 2. This includes the main processing script (mnmcount.py) that aggregates M&M color counts by state, a generator script (gen\_mnm\_dataset.py) to create sample datasets, a small test dataset (test\_agg.csv), and a README with execution instructions.
chapter2/py · high confidence
Added Python examples for working with Spark DataFrames and Rows
New Python scripts (Example-3\_6.py and rows.py) were added to demonstrate creating and manipulating Spark DataFrames using explicit schemas and Row objects. The changes include a README with execution instructions, providing users with concrete examples of defining StructTypes, creating DataFrames from lists of tuples or Row objects, and performing column operations.
chapter3/py · high confidence
Added Scala example for counting M&Ms by state and color
A new Scala example, MnMcount, was added to the chapter2/scala directory. This example demonstrates how to read a CSV dataset of M&Ms, perform aggregations to count occurrences by state and color, and display the results. The accompanying README provides instructions on how to build and run the example using sbt and spark-submit.
chapter2/scala · high confidence
Added Scala examples for Chapter 7: Spark configuration, partitions, map operations, caching, and joins
The chapter7/scala directory now includes new Scala source files demonstrating Spark programming patterns. These include SparkConfig\_7\_1 for configuring Spark sessions, Partitions\_7\_2 for managing data partitions, MapAndMapPartitions\_7\_3 for processing data with map and mapPartitions operations, CachingData\_7\_5 for caching and persisting DataFrames, and SortMergeJoin\_7\_6 and SortMergeJoinBucketed\_7\_6 for performing sort-merge joins on both standard and bucketed DataFrames. A README.md file was also added to provide build and run instructions for these examples.
chapter7/scala · high confidence
Added build script and updated documentation for building Spark JARs
Users can now build all relevant Spark JAR files for chapters 2, 3, 6, and 7 by running the new \build\_jars.py\ script, which automates the process of compiling Scala and Java code using SBT. The top-level README has been updated to document this new workflow and provide instructions for setting up the environment, while the .gitignore file has been updated to exclude IDE and OS-specific temporary files.
(repo-wide) · high confidence
Added sample dataset for state and color counts
A new CSV file, mnm\_dataset.csv, was added to the chapter2/scala/data directory. This file contains 100,000 rows of sample data with columns for State, Color, and Count, providing the necessary data for the associated Scala data processing code.
chapter2/scala/data · high confidence
Added sample datasets for Chapter 3
New sample datasets have been added to the chapter3/data directory to support the examples in Chapter 3. This includes blogs.json, a JSON file containing structured blog post metadata, and sf-fire-calls.csv, a CSV file containing historical San Francisco fire call records. These files provide the necessary data sources for the chapter's exercises and demonstrations.
chapter3/data · high confidence
Behavioural changes
Updated notebook import instructions for Databricks Runtime 9.1 ML LTS
The notebooks/README.md file was added, providing step-by-step instructions for importing the LearningSparkv2.dbc notebook file into Databricks. The guide specifies using a Databricks cluster with Runtime 9.1 ML LTS, which pre-installs common ML libraries like MLflow and XGBoost.
notebooks · medium confidence
Dependencies
Updated Spark dependencies to 3.0.0-preview2
The build configuration for Scala and Java examples in chapters 2, 3, 6, and 7 has been updated to use Spark 3.0.0-preview2 for both spark-core and spark-sql libraries.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Score
- CAI 41 → 55 (+13.1)
- Rubric changed (rubric-2026.08.15 → rubric-2026.09.15) — scores are not directly comparable.
Lenses
- Code Health 100 → 99 (-0.5)
- Architecture 69 → 100 (+31.0)
- Maturity 45 → 45 (-0.3)
- Readiness 17 → 36 (+19.2)
- Security 95 → 95 (+0.0)
Resolved (6)
- Dimension evaluation failed
- LLM evaluation failed
- No automated tests
- No exposed public API
- No tests found
- Test reliability not included
New (5)
- Duplicated block (14 lines × 2) (chapter7/scala/src/main/scala/chapter7/SortMergeJoinBucketed_7_6.scala)
- Duplicated block (6 lines × 3) (chapter7/scala/src/main/scala/chapter7/MapAndMapPartitions_7_3.scala)
- Edited copy of a member (27 corresponding lines) (chapter7/scala/src/main/scala/chapter7/SortMergeJoinBucketed_7_6.scala)
- Orphaned files with no living knowledge
- Secret: generic-api-key (mlflow_on_colab.ipynb)
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
databricks/LearningSparkV2 was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 27 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 8d5bdf6fd7b58da930855eb3d5ead38e6fb039e4 — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-d00c643c3f66.