aws/sagemaker-spark
40.4
Weak · 20 September 2026
11.5k
lines of production code
Scala
with Python
1
measurement over time
What this system is
This system is the SageMaker PySpark SDK, a library that enables users to build, train, and deploy machine learning models using Apache Spark and Amazon SageMaker. It provides estimator wrappers for algorithms such as Linear Learner, XGBoost, Factorization Machines, and LDA, while handling data serialization via Protobuf and managing SageMaker resources like training jobs and endpoints. The SDK supports custom naming policies, safe S3 artifact cleanup, and matrix feature conversion, operating within a CI/CD pipeline built on AWS CodeBuild.
Features
Custom naming policies and Python-to-Scala list conversion support
Users can now provide explicit names for SageMaker entities (training jobs, models, endpoint configs, and endpoints) using the new CustomNamePolicy and CustomNamePolicyWithTimeStampSuffix classes, which are exposed via the sagemaker\_pyspark package. Additionally, the SDK now supports passing Python lists to the underlying Scala/Java layer via a new ScalaList wrapper, enabling seamless integration of Python list data structures in SageMaker operations.
_sagemaker-pyspark-sdk/src/sagemaker\pyspark · high confidence
Initial release of SageMaker Spark SDK documentation and build infrastructure
This change introduces the project's initial public release artifacts, including the CHANGELOG (v1.0.2 through v1.4.5), CONTRIBUTING guidelines, CODEOWNERS, and Code of Conduct. It establishes the CI/CD pipeline using AWS CodeBuild (buildspec.yml) with sbt 1.7.1 and OpenJDK 8, and updates the README to reflect compatibility with EMR 5.11.0+ and provide updated installation instructions.
(repo-wide) · high confidence
New algorithm support and expanded Linear Learner capabilities
The SDK now supports Factorization Machines and Latent Dirichlet Allocation (LDA) algorithms, with corresponding estimator classes and region account maps added for both. Linear Learner has been extended to support multiclass classification, additional loss functions (such as hinge, huber, and quantile), new optimizers (rmsprop, auto), and early stopping parameters. Additionally, the XGBoost estimator has been fixed to use integer types for the maxDepth and seed parameters, correcting previous type mismatches.
sagemaker-spark-sdk/src/main/scala/com/amazonaws/services/sagemaker/sparksdk/algorithms · high confidence
New algorithm wrappers and expanded Linear Learner capabilities
The SDK now includes new estimator wrappers for Latent Dirichlet Allocation (LDA) and Factorization Machines, enabling users to perform topic modeling and recommendation-style regression/classification tasks directly within Spark pipelines. Additionally, the Linear Learner estimator has been significantly enhanced with support for multiclass classification, new loss functions (such as hinge, Huber, and quantile losses), new optimizers (rmsprop, auto), and early stopping parameters, allowing for more flexible model tuning and broader problem coverage.
_sagemaker-pyspark-sdk/src/sagemaker\pyspark/algorithms · high confidence
Support for Matrix features in RecordIO-Protobuf conversion
The ProtobufConverter now handles Matrix types (DenseMatrix and SparseMatrix) in addition to Vector types when converting Spark DataFrame rows to RecordIO-Protobuf format. Users can now include Matrix columns as features, which are encoded using column-major order for dense matrices and Compressed Sparse Column (CSC) format for sparse matrices, with shape and key information preserved in the protobuf output. The SageMakerProtobufWriter also exposes its output path via a new path() method.
sagemaker-spark-sdk/src/main/scala/com/amazonaws/services/sagemaker/sparksdk/protobuf · high confidence
Support for new SageMaker algorithms and increased runtime timeout
The SDK now supports deserializing responses from Linear Learner multiclass classifiers, Factorization Machines (binary classifier and regressor), and Latent Dirichlet Allocation (LDA) models, allowing users to process inference results from these algorithms. Additionally, the SageMaker runtime client's socket timeout has been increased to 80 seconds to handle longer-running transformation requests more reliably.
sagemaker-spark-sdk/src/main/scala/com/amazonaws/services/sagemaker/sparksdk/transformation · high confidence
Behavioural changes
Custom naming policies and safer S3 cleanup for SageMaker jobs
Users can now configure custom names for SageMaker training jobs, models, endpoint configs, and endpoints via new \CustomNamePolicy\ and \CustomNamePolicyWithTimeStampSuffix\ factories, allowing explicit control over entity naming or automatic timestamp suffixing. Additionally, the SDK now restricts S3 training data deletion to objects strictly belonging to the invoking job, preventing accidental deletion of artifacts from other jobs that share a similar staging path prefix.
sagemaker-spark-sdk/src/main/scala/com/amazonaws/services/sagemaker/sparksdk · high confidence
PySpark SDK updates to version 3.3.2 and Python 3.7+ with improved documentation
The sagemaker-pyspark-sdk now requires Python 3.7 or higher and pins the PySpark dependency to version 3.3.2. The README has been updated to reflect modern PySpark practices, recommending the use of SparkSession instead of SparkContext, adding instructions for installing from source, and including new examples for XGBoost and S3 file system schemes (s3:// and s3a://). Additionally, the build process now reads the version from a dedicated VERSION file, and testing infrastructure has been updated to use pytest with parallel execution support.
sagemaker-pyspark-sdk · high confidence
Fixes
Fix MNIST example notebooks and data
Corrected the MNIST examples to ensure they run successfully: updated the Zeppelin notebook to use dynamic IAM role resolution and updated instance types, fixed the data source path in the Zeppelin notebook, added the missing test data file required by the Boto3 invocation example, and removed the matplotlib dependency from the Zeppelin notebook to prevent import errors.
examples · high confidence
Test coverage
Added and updated tests for name policies, S3 resources, wrappers, and estimator attributes; Added tests for Factorization Machines and LDA deserializers; Added tests for Matrix support in ProtobufConverter; Added tests for custom name policies and S3 cleanup safety; Added tests for new algorithm estimators and expanded region coverage; Added tests for new algorithm wrappers and updated Linear Learner and XGBoost test coverage.
Dependencies
Upgrade Spark 2.x to 3.3.0 and Scala 2.11 to 2.12
The SDK now targets Apache Spark 3.3.0 (up from 2.2.0) and Scala 2.12.16 (up from 2.11.7), requiring users to run on compatible Spark and Scala versions. This upgrade is accompanied by updates to AWS SDK dependencies (e.g., aws-java-sdk-s3, hadoop-aws) and test libraries (scalatest, mockito), along with the removal of the Python requirements file for the Spark SDK module.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Baseline
- First survey — no prior run to compare against. CAI 40.
Lenses
- Code Health 96
- Architecture 100
- Maturity 49
- Readiness 43
- Security 95
- Accessibility 19
Changes since last survey
- 131 commits — 116 feature/other, 15 fixes
By area
- (root) — 68 commits
- sagemaker-spark-sdk/src — 19 commits
- sagemaker-pyspark-sdk/src — 9 commits
- sagemaker-pyspark-sdk/README.rst — 7 commits
- sagemaker-pyspark-sdk/setup.py — 6 commits
- examples/notebooks — 4 commits
- sagemaker-pyspark-sdk/docs — 3 commits
- sagemaker-pyspark-sdk/tox.ini — 3 commits
- sagemaker-spark-sdk/build.sbt — 3 commits
- sagemaker-spark-sdk/project — 3 commits
- .github/PULL_REQUEST_TEMPLATE.md — 2 commits
- .github/issue_template.md — 1 commit
- docs/com — 1 commit
- docs/lib — 1 commit
- sagemaker-pyspark-sdk/requirements.txt — 1 commit
Notable commits
- fix: Fix XGBoostEstimator maxDepth parameter from Double to Int (#30)
- fix: Fix creating model from training job, endpoint and s3 file in pyspark (#39)
- fix: Fix data location in message (#70)
- fix: Fix max depth description in XGBoost Estimator. (#57)
- fix: Fix seed parameter for XGBoost (#41)
- fix: fix has-matching-changes to reflect latest script changes (#105)
- fix: fix output in MNIST-kmeans.json (#48)
- fix: fix: Upgrade pyspark version to 2.3.4 (#112)
- fix: fix: confine deleteTrainingData to the invoking job's own S3 objects (#162)
- fix: fix: freeze pyspark version to 2.3.2 (#83)
- fix: fix: remove compiler-bridge dependency (#148)
- fix: fix: set version of pyspark setup.py (#77)
- fix: fix: upgrade jquery to version 1.9.0 (#129)
- fix: fix: upgrade pyspark (#135)
- fix: notebook fixes: import fix, remove matplotlib cell (#44)
- change: Add CODEOWNERS file (#159)
- change: Add Linear Learner multiclass classifier (#60)
- change: Add PySpark Example For XGBoost Multi-Class Classification on MNIST (#32)
- change: Add build.properties and plugins.sbt (#2)
- change: Add contributing.md (#31)
- …and 111 more
Architecture
- 0 containers · 1 bounded contexts · 0 dependency edges (baseline)
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
aws/sagemaker-spark was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 20 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit f83a1856a9c1f1effc247d3827a5c6f76b3ba43f — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-b51f968c9b10.