Skip to content
CAI
Software that uses CAICheck a score

apache/incubator-toree

71.0

Strong · 28 September 2026

19.5k

lines of production code

Scala

primary language

2

measurements over time

CAI band scale
CAI trend line
CAI lens gauges

What this system is

Apache Toree is a Jupyter kernel that enables interactive code execution and data analysis within Apache Spark notebooks. It provides interpreters for Scala and SQL, allowing users to run Spark jobs, manage dependencies, and visualize results directly in the browser. The system supports rich output formatting, dynamic variable injection in SQL queries, and real-time monitoring of Spark job metrics through a plugin architecture.

How it got here

2014–2015 — Apache release and Pekko migration

17 changes.

This period centered on the initial Apache release of Toree 0.6.0, establishing the project's foundational build infrastructure and documentation. The work involved a significant architectural shift from Akka to Apache Pekko and alignment with Spark 3.4, accompanied by the removal of legacy entry-point code. Extensive test coverage was added to validate the new kernel-api module, client communication protocols, and interpreter functionality.

2016–2025 — Apache migration and Pekko refactoring

47 changes.

The project transitioned from IBM Spark to Apache Toree, updating licensing and package namespaces while migrating the actor infrastructure from Akka to Apache Pekko. This period introduced a comprehensive plugin system, expanded magic command capabilities, and added SQL interpreter support, alongside significant protocol updates to align with Jupyter v5.2 specifications.

Features

Add SQL interpreter support for Spark SQL

This change introduces a new SQL interpreter within the Toree kernel, allowing users to execute SQL queries directly against a Spark SQL context. The implementation includes the SqlInterpreter class to manage the lifecycle and execution of SQL code, a SqlService to handle the submission of queries to the Spark session, and supporting components like SqlTransformer and SqlException to integrate with the existing interpreter framework. Users can now write and run SQL statements in notebooks, with results displayed as output tables.

sql-interpreter/src/main/scala/org/apache/toree/kernel · high confidence

Added Scala 2.13-specific interpreter implementation

A new Scala 2.13-specific implementation of the interpreter logic has been added to the \scala-interpreter\ module. This file provides the concrete details for how the Scala interpreter interacts with the Scala 2.13 compiler and runtime, ensuring compatibility with this specific Scala version.

scala-interpreter/src/main/scala-2.13 · high confidence

Added experimental and internal API annotations

New Scala annotations, Experimental and Internal, have been added to the org.apache.toree.annotations package. The Experimental annotation marks APIs as subject to change between minor versions, while the Internal annotation indicates that an API is not stable or public. These annotations are available for use in Scala code to signal the stability and intended usage scope of specific components.

macros/src · high confidence

Added sample data files and example notebooks for Toree

New example resources have been added to the \etc/examples/notebooks\ directory to help users get started with Toree. This includes \cars.json\ and \people.json\ sample datasets, a \magic-tutorial.ipynb\ notebook that demonstrates line and cell magics (such as \%LsMagic\, \%Truncation\, and \%%sql\), and a \meetup-streaming-toree.ipynb\ notebook that provides a complete demo of streaming data from the Meetup API into a Spark Streaming job and visualizing it with declarative widgets.

etc/examples · high confidence

Added source release ignore list and development kernel configuration

The project now includes an .src-release-ignore file to exclude build artifacts, IDE settings, and example data from source distributions, ensuring cleaner releases. Additionally, a new kernel.json file has been added to define the configuration for the Apache Toree development kernel, specifying environment variables such as PYTHONPATH and SPARK\_HOME, along with the command-line arguments required to launch the kernel via Jupyter.

etc · high confidence

Added support for code completeness checking

The interpreter actor now handles \is\_complete\_request\ messages, allowing clients to check if a block of code is syntactically complete before execution. This is implemented by introducing a new \IsCompleteTask\ child actor and corresponding \IsCompleteTask\ enumeration value within the interpreter package, alongside the existing execute and code-completion tasks.

kernel/src/main/scala/org/apache/toree/kernel/protocol/v5/interpreter · high confidence

Enable pip installation of Apache Toree

Users can now install Apache Toree as a Jupyter kernel via pip (pip install toree). This change introduces the setup.py and MANIFEST.in files required to package the project, including the kernel logo, version info, and legal notices (LICENSE, NOTICE, DISCLAIMER). The package declares dependencies on jupyter\_core, jupyter\_client, and traitlets (all version 4.0 or higher) and registers a console script entry point to facilitate kernel installation.

_etc/pip\install · high confidence

Initial Apache release of Toree (Incubating) 0.6.0

This change introduces the initial release of Apache Toree (Incubating) version 0.6.0. It provides the complete source code, build configuration (Makefile, Dockerfiles), and documentation (README, Release Notes) required to build and install the Jupyter kernel for Apache Spark. The release includes a Python-based integration test suite (test\_toree.py) to verify kernel functionality, along with standard Apache licensing and compliance files (LICENSE, NOTICE, DISCLAIMER, .asf.yaml).

(repo-wide) · high confidence

Initial plugin dependency injection system

Added the core components for a plugin dependency management system within the plugins module. This includes the Dependency case class for holding typed values, a DependencyManager for storing and retrieving dependencies by name or type, and specific exception classes (DependencyException, DepNameNotFoundException, etc.) to handle lookup failures.

plugins/src/main/scala/org/apache/toree/plugins/dependencies · high confidence

Initial plugin system implementation

Introduces a new plugin architecture for Apache Toree, allowing users to load and execute external code modules. This change adds the core plugin infrastructure, including a \PluginManager\ for loading and initializing plugins, a \PluginClassLoader\ for isolated class loading, and a \PluginSearcher\ to discover plugin classes. It defines the \Plugin\ trait for creating custom plugins, \PluginMethod\ for handling method invocation with dependency injection, and \PluginEvents\ (such as \SparkReady\ and \PreRunCell\) to enable event-driven interactions within the kernel.

plugins/src/main · high confidence

Introduce pip-installable Apache Toree kernel with CLI installation options

Users can now install the Apache Toree Jupyter kernel via \pip install toree\ and register it using the \jupyter toree install\ command. This new Python package provides a CLI that allows specifying the Spark home directory, Python executable, and kernel arguments (such as \--toree\_opts\ and \--spark\_opts\) during installation. The installation process supports configuring which interpreters (currently Scala and SQL) are included in the kernel spec.

_etc/pip\install/toree · high confidence

New Comm communication infrastructure for kernel targets

The protocol layer now includes a dedicated communication subsystem for managing kernel-to-kernel or kernel-to-extension interactions. This change introduces a new package, org.apache.toree.comm, containing five core components: CommCallbacks for registering and executing event handlers (open, message, close); CommManager for orchestrating connections and automatically linking/unlinking targets; CommRegistrar for managing target registrations and callback chains; CommStorage for persisting target-to-Comm-ID mappings and callback data; and CommWriter for sending structured Comm messages (open, msg, close) to registered targets. This infrastructure enables the kernel to dynamically register communication targets, handle incoming Comm events via pluggable callbacks, and maintain bidirectional connections with external extensions or other kernels.

protocol/src/main/scala/org/apache/toree/comm · high confidence

New MagicParser class for handling magic command syntax

A new MagicParser class has been added to the kernel protocol v5 magic package to handle the parsing and substitution of line and cell magic commands. This component identifies magic invocations (e.g., %magic or %%magic), validates them against the registered MagicManager, and substitutes them with equivalent kernel object calls, returning errors for invalid or non-existent magics.

kernel/src/main/scala/org/apache/toree/kernel/protocol/v5/magic · high confidence

New SparkMonitor plugin for real-time Jupyter kernel integration

A new SparkMonitor plugin has been added to the Toree plugin system, enabling real-time monitoring of Apache Spark jobs, stages, and tasks directly within Jupyter notebooks. The plugin registers a 'SparkMonitor' comm target to facilitate bi-directional communication and automatically detects when a SparkContext becomes available to attach a listener. This listener forwards comprehensive event data—including application start/end, job and stage lifecycle, task execution metrics, and executor status changes—as JSON messages to the Jupyter client, allowing for live visualization and debugging of Spark workloads.

spark-monitor-plugin/src/main/scala/org/apache/toree/plugins/sparkmonitor · high confidence

New kernel API classes for display, streaming, and factory methods

The kernel now exposes new API classes in the \org.apache.toree.kernel.api\ package to handle output and data streaming. \DisplayMethods\ allows interpreters to send rich display content to the client by MIME type and supports clearing output. \StreamMethods\ provides functionality to stream text data (such as stdout) to the client. \FactoryMethods\ enables the creation of kernel input and output streams, while the \Kernel\ class integrates these capabilities, including magic parsing and interpreter interaction, into the main kernel interface.

kernel/src/main/scala/org/apache/toree/kernel/api · high confidence

New plugin lifecycle and dependency annotations

Added a set of Java annotations in the plugins module to support plugin initialization, destruction, event handling, and dependency injection. Specifically, the diff introduces @Init and @Destroy for marking lifecycle methods, @Event and @Events for defining plugin events, @DepName for specifying dependency names on parameters, and @Priority for setting execution order on types and methods. These annotations provide the metadata foundation for the plugin system's runtime behavior.

plugins/src/main/java/org/apache/toree/plugins/annotations · high confidence

New run.sh script for launching the Spark Kernel

A new \run.sh\ script has been added to \etc/bin\ to handle the execution of the Apache Toree kernel. This script validates that \SPARK\_HOME\ is set, locates the kernel assembly JAR in the \lib\ directory, and invokes \spark-submit\ with the appropriate class and options. It supports runtime configuration via \SPARK\_OPTS\ and \TOREE\_OPTS\ environment variables, which override any options stored during installation.

etc/bin · high confidence

Removals

Removed SparkKernel and SparkKernelOptions entry-point files

The \SparkKernel.scala\ application object and the \SparkKernelOptions.scala\ command-line argument parser have been deleted from the source tree. This removes the previous standalone kernel startup logic and its associated CLI options (such as \--help\, \--verbose\, \--create-context\, and \--profile\) from the \src/main\ location, indicating a shift away from this specific entry-point implementation.

src/main · high confidence

Architecture

New kernel-api module with interpreter and dependency management APIs

The core kernel functionality has been reorganized into a new \kernel-api\ module, introducing the \Interpreter\ trait for code execution and the \DependencyDownloader\ abstraction with \CoursierDependencyDownloader\ and \IvyDependencyDownloader\ implementations for resolving external JARs. This change also adds the \BrokerService\ and \BrokerProcess\ classes to manage external language processes, \StreamState\ for controlling standard I/O redirection, and \LanguageInfo\ for providing syntax highlighting metadata.

kernel-api · high confidence

Behavioural changes

Added bundled third-party license and notice files for binary distributions

The binary distribution now includes a comprehensive set of legal files in \etc/legal/\ to comply with Apache Software Foundation policy for bundled third-party code. \LICENSE\_extras\ and \NOTICE\_extras\ explicitly list and attribute external dependencies such as Apache Pekko, Jackson, Spring Framework, and Coursier, while the \licenses/\ directory contains the full text for non-Apache-2.0 licenses (MPL-2.0, BSD, MIT) required to be distributed with the software. These files are concatenated into the final distribution and embedded in the assembly jar to ensure all bundled artifacts are properly accounted for.

etc/legal · high confidence

Builtin magics migrated to plugin architecture with new capabilities

The built-in magics in the kernel have been refactored into a plugin-based system under the \org.apache.toree.magic.builtin\ package, introducing several new capabilities and behavioral changes. The \%AddDeps\ magic now supports credential files, repository configuration, and module exclusions for dependency resolution, while \%AddJar\ has been expanded to handle cloud storage schemes (HDFS, S3, S3A, S3N, and Google Cloud Storage) and includes validation for JAR integrity. A new \%%dataframe\ magic allows users to convert Spark DataFrames into HTML, CSV, or JSON outputs. Additionally, new magics \%ShowOutput\, \%ShowTypes\, and \%Truncation\ allow users to toggle kernel display options, and \%LSMagic\ provides a dynamic list of available line and cell magics.

kernel/src/main/scala/org/apache/toree/magic · high confidence

Client bootstrap refactored to use Pekko and modular initialization layers

The client boot process has been restructured to use Apache Pekko instead of Akka for the underlying actor system, and the initialization logic is now split into modular traits (SystemInitialization and HandlerInitialization). This change updates the Spark Kernel client to leverage Pekko's actor system for managing heartbeat, stdin, shell, and IOPub connections, while also standardizing how message handlers like ExecuteHandler are registered. Users benefit from a more maintainable and modernized client infrastructure that aligns with the broader project migration to Pekko.

client/src/main/scala/org/apache/toree/kernel/protocol/v5/client/boot · high confidence

Client socket communication migrated to Apache Pekko actors

The client-side socket communication layer (Heartbeat, IOPub, Shell, and Stdin clients) has been rewritten to use Apache Pekko actors instead of the previous ZeroMQ extension. This change replaces direct ZeroMQ socket handling with Pekko-based socket actors (ReqSocketActor, DealerSocketActor, SubSocketActor) and updates all relevant imports to org.apache.pekko, aligning the client protocol implementation with the broader project migration from Akka to Pekko.

client/src/main/scala/org/apache/toree/kernel/protocol/v5/client/socket · high confidence

Client-side kernel communication migrated to Apache Pekko

The client-side kernel protocol implementation has been refactored to use Apache Pekko instead of Akka for actor-based communication. This change introduces new client components, including the ActorLoader interface and SparkKernelClient, which manage actor selection and handle code execution requests via the Jupyter kernel protocol. The underlying messaging infrastructure now relies on Pekko actors for handling execute requests, heartbeats, and stdin responses, ensuring compatibility with the updated runtime environment.

client/src/main/scala/org/apache/toree/kernel/protocol/v5/client · high confidence

Communication layer refactored to use Apache Pekko and JeroMQ

The communication module has been rewritten to replace the previous Akka-based actors with Apache Pekko actors (e.g., DealerSocketActor, PubSocketActor) and to use JeroMQ for ZeroMQ socket management via the new SocketManager. This change introduces a new actor-based architecture for handling request, reply, publish, subscribe, router, and dealer sockets, along with updated security components (Hmac, SignatureCheckerActor) that now rely on Pekko's asynchronous patterns. Users will experience the same kernel messaging capabilities but with the underlying infrastructure migrated to the Pekko ecosystem and JeroMQ bindings.

communication/src/main · high confidence

Enhanced DataFrame display and improved kernel message logging

Users can now view DataFrames in HTML, JSON, and CSV formats, with null values and array structures explicitly rendered for better readability. Additionally, the kernel's internal logging for incoming messages has been expanded to include trace-level details for message IDs, signatures, and parent headers, aiding in debugging and observability.

kernel/src/main/scala/org/apache/toree/utils · high confidence

Fix exception propagation in Scala 2.12 interpreter

The Scala 2.12-specific interpreter implementation now correctly handles and propagates exceptions during code execution. This change ensures that errors occurring within the Scala REPL (IMain) are properly surfaced to the user rather than being silently swallowed or causing undefined behavior, improving reliability when running Scala 2.12 code.

scala-interpreter/src/main/scala-2.12 · high confidence

Interpreter task actors migrated to Apache Pekko

The kernel's interpreter task actors (CodeComplete, ExecuteRequest, IsComplete) and their factory have been moved to the new package org.apache.toree.kernel.protocol.v5.interpreter.tasks and updated to use Apache Pekko instead of Akka for actor management. This change updates the underlying actor framework while preserving the existing task execution behavior for code completion, execution, and completeness checks.

kernel/src/main/scala/org/apache/toree/kernel/protocol/v5/interpreter/tasks · high confidence

Kernel boot process restructured into modular initialization layers

The kernel startup sequence in the boot layer has been refactored into distinct, composable initialization traits (BareInitialization, ComponentInitialization, HandlerInitialization, HookInitialization, and InterpreterManager) to improve modularity and separation of concerns. This change introduces support for alternate interrupt signals via the --alternate-sigint configuration option, allows dependency installation directories to be specified or automatically generated as unique temporary paths, and enables lazy Spark session initialization controlled by the spark\_context\_initialization\_mode configuration. Additionally, the kernel now fires AllInterpretersReady events after interpreter initialization and supports Comm Info messages, while migrating the underlying actor system from Akka to Pekko.

kernel/src/main/scala/org/apache/toree/boot/layer · high confidence

Kernel message relay now ignores unknown message types

The kernel's message relay mechanism has been updated to gracefully handle unknown incoming and outgoing message types. Instead of failing or throwing errors when encountering unrecognized message types, the system now logs a warning and ignores them. This change improves robustness by preventing disruptions from unexpected or future message formats, ensuring the kernel continues to operate smoothly even when it encounters messages it does not explicitly understand.

kernel/src/main/scala/org/apache/toree/kernel/protocol/v5/relay · high confidence

Kernel socket endpoints now use Pekko actors instead of direct ZeroMQ

The kernel's communication layer for control, shell, IOPub, stdin, and heartbeat messages has been refactored to route through Apache Pekko actors (RepSocketActor, RouterSocketActor, PubSocketActor) rather than using direct ZeroMQ socket extensions. This change, driven by the migration from Akka to Pekko, means that socket connections are now managed via the actor system, which may affect how message routing and lifecycle are handled internally, though the external IPython kernel protocol interface remains the same.

kernel/src/main/scala/org/apache/toree/kernel/protocol/v5/kernel/socket · high confidence

New Main entry point for the kernel

The kernel now uses a new Main.scala entry point to initialize and run. This class handles command-line arguments, displays help or version information (including Scala and Spark versions via BuildInfo), and starts the kernel bootstrap process with standard initialization components.

kernel/src/main/scala/org/apache/toree · high confidence

New Scala interpreter implementation with enhanced displayers

The Scala interpreter in the Toree kernel has been replaced with a new implementation that improves how results are displayed in Jupyter notebooks. A new \ScalaDisplayers\ module registers custom handlers for Spark objects, \MagicOutput\, and \Option\ types, ensuring that Spark contexts and SQL rows are rendered as rich HTML tables and links rather than plain text. The interpreter now also supports calling a \toHtml\ method on Scala objects for custom rendering, and ensures that kernel variables and Spark bindings are defined before user code executes to prevent initialization errors.

scala-interpreter/src/main/scala/org/apache/toree/kernel · high confidence

New command-line options for Spark context initialization and repository configuration

The kernel now supports explicit control over Spark context creation via the new \--spark-context-initialization-mode\ option (accepting \none\, \lazy\, or \eager\), which replaces the deprecated \--nosparkcontext\ flag. Users can also configure the timeout for context creation using \--spark-context-initialization-timeout\. Additionally, new options \--default-repositories\ and \--default-repository-credentials\ allow users to specify custom Maven/Ivy repositories and their credentials at launch, while \--alternate-sigint\ enables configuring a custom signal for interrupting long-running cells.

kernel/src/main/scala/org/apache/toree/boot · high confidence

New configuration and logging resources for the kernel

This change introduces new configuration files (application.conf, reference.conf) and logging properties (log4j.properties) for both compile and test scopes. These files define default settings for kernel ports, Spark context initialization modes and timeouts, interpreter plugins (Scala and SQL), and logging levels, establishing the baseline configuration for the kernel's runtime behavior.

resources · high confidence

New release-audit and signing tooling in etc/tools

The etc/tools directory now includes a suite of scripts to enforce Apache Software Foundation release policies. The check-licenses script runs Apache RAT (version 0.12) against the source tree to verify that all files carry the required license headers, using a new .rat-excludes file to ignore generated or non-commentable formats. The release-build.sh script manages the release lifecycle, supporting separate prepare and publish phases, GPG signing via the MAVEN\_GPG\_PASSPHRASE environment variable, and dry-run testing. Additionally, sign-file generates detached GPG signatures and SHA512 checksums for artifacts, while verify-release audits a directory to ensure all release artifacts have valid signatures and matching checksums before distribution.

etc/tools · high confidence

Protocol v5 content models updated to align with Jupyter spec changes

The kernel protocol content models have been updated to match the Jupyter messaging specification. This includes adding support for \comm\_info\_request\ and \comm\_info\_reply\ messages, implementing the \is\_complete\_request\ and \is\_complete\_reply\ pair, and correcting the \kernel\_info\_reply\ structure to include the \language\_info\ field. Additionally, the deprecated \source\ field has been removed from \display\_data\, and the \connect\_request\/\connect\_reply\ messages are now marked as deprecated in the protocol.

protocol/src/main/scala/org/apache/toree/kernel/protocol/v5/content · high confidence

Protocol v5 handlers migrate to Apache Pekko and implement Comm and status messaging

The kernel's protocol v5 handler layer has been rewritten to use Apache Pekko (replacing Akka) and introduces a new \StatusDispatch\ actor to manage busy/idle status messages sent to clients. This change adds support for the Jupyter Comm protocol with new handlers for \comm\_open\, \comm\_msg\, \comm\_close\, and \comm\_info\_request\ messages, allowing external tools to interact with the kernel via custom channels. Additionally, the \ExecuteRequestHandler\ now fires \PreRunCell\ and \PostRunCell\ plugin events around code execution, and the \ShutdownHandler\ distinguishes between a graceful shutdown request and a forced exit, ensuring proper cleanup before termination.

kernel/src/main/scala/org/apache/toree/kernel/protocol/v5/handler · high confidence

Rebrand to Apache Toree and update licensing

The project has been rebranded from IBM Spark to Apache Toree, with all Scala source files moved to the new \org.apache.toree\ package namespace and license headers updated to the Apache Software Foundation (ASF) License, Version 2.0. This change affects the client-side execution handling (including \DeferredExecution\ and \DeferredExecutionManager\), kernel communication components (\KernelCommManager\, \KernelCommWriter\), and global state management (\ExecuteRequestState\, \ExecutionCounter\, \ScheduledTaskManager\), ensuring all code now reflects the new organizational identity and licensing terms.

(repo-wide) · high confidence

Rebranded client communication components to Apache Toree

The client-side communication infrastructure has been rebranded from IBM Spark to Apache Toree, moving the \ClientCommManager\ and \ClientCommWriter\ classes into the \org.apache.toree.comm\ package. These components now handle the creation and sending of kernel messages (open, msg, close) from the client to the kernel via the shell actor, with updated Apache License headers reflecting the project's new ownership.

client/src/main/scala/org/apache/toree/comm · high confidence

Reimplemented Jupyter Protocol v5 message structures and builders

The kernel's protocol layer has been rewritten to align with Jupyter messaging protocol v5.2. This introduces new core data structures for message headers, kernel metadata, and language information, alongside a builder pattern for constructing kernel messages. The change ensures the kernel correctly implements the v5.2 specification, including support for new message types like \is\_complete\_request\ and \comm\_info\, and fixes the structure of kernel info replies to match the updated spec.

protocol/src/main/scala/org/apache/toree/kernel/protocol/v5 · high confidence

SQL magic now resolves Scala variables within statements

The SQL cell magic (\%%sql\) now automatically injects and resolves Scala variables into SQL statements before execution. When a Scala interpreter is available, the magic evaluates the provided code as a Scala string interpolation expression (e.g., \s"SELECT \* FROM $table"\), allowing dynamic table names or values to be substituted directly into the query sent to the SQL interpreter. This enables seamless integration between Scala variables and SQL queries without manual string concatenation.

sql-interpreter/src/main/scala/org/apache/toree/magic · high confidence

Scala magic command now supports MIME type output

The Scala cell magic has been refactored to produce output by MIME type, aligning with the new plugin-based magic architecture. This change allows the Scala interpreter to return structured output data rather than just raw text, improving compatibility with frontends that support rich media rendering.

scala-interpreter/src/main/scala/org/apache/toree/magic · high confidence

Test coverage

Added Scala client usage examples; Added comprehensive test suite for the plugin system; Added integration test specs for client socket communication; Added integration tests for JeroMQ socket communication; Added integration tests for Scala interpreter jar loading and display representations; Added integration tests for signature security actors; Added system test for kernel stdin input handling; Added system tests for client-side Comm API interactions; Added test utilities for Spark client system testing; Added test utilities for deploying Spark Kernel and Client instances; Added tests for Hmac and HmacAlgorithm components; Added tests for OrderedSupport actor behavior; Added unit tests for JeroMQSocket and ZeroMQSocketRunnable; Added unit tests for client communication components; Added unit tests for kernel protocol v5 client components; Added unit tests for protocol v5 content message serialization; Added unit tests for protocol v5 header and message builder components; Added unit tests for the comm module components; Removed SparkKernelOptionsSpec test suite.

Dependencies

Build infrastructure consolidated and upgraded to SBT 1.9.3 with Pekko migration

The build system has been restructured to centralize configuration, introducing a new CommonPlugin that defines dedicated unit, integration, and system test configurations, and a Dependencies object that manages all library versions. This change upgrades the build tool to SBT 1.9.3 and migrates the core actor framework from Akka to Apache Pekko 1.1.5. Additionally, key dependencies have been aligned with Spark 3.4.0, including Jackson Databind 2.14.2 and SLF4J 2.0.16, while the assembly plugin was updated to version 2.3.1.

project · high confidence

Migrate from Akka to Pekko and upgrade to Spark 3.4.4

The build system has been updated to replace the Akka actor framework with Apache Pekko, as seen in the new \Dependencies.pekkoActor\, \Dependencies.pekkoSlf4j\, and \Dependencies.pekkoTestkit\ references in the client, communication, and kernel module build files. Additionally, the project now targets Apache Spark 3.4.4 (configurable via the \APACHE\_SPARK\_VERSION\ environment variable) and supports Scala 2.12 and 2.13, replacing the previous single-version setup. The build also consolidates dependency management into a central \Dependencies.scala\ object and updates assembly strategies to correctly handle legal files under \META-INF\.

(dependencies) · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

Score

  • CAI 70 → 71 (+0.5)
  • Rubric changed (rubric-2026.09.8 → rubric-2026.09.16) — scores are not directly comparable.

Lenses

  • Code Health 93 → 93 (+0.0)
  • Architecture 99 → 97 (-1.5)
  • Maturity 70 → 75 (+5.7)
  • Readiness 65 → 64 (-0.6)
  • Security 70 → 70 (+0.0)

Resolved (3)

  • Concentrated knowledge decay
  • Most significant orphaned file (kernel/src/main/scala/org/apache/toree/utils/ClassPath.java)
  • Most significant orphaned file (spark-monitor-plugin/src/main/scala/org/apache/toree/plugins/sparkmonitor/JupyterSparkMonitorListener.scala)

New (3)

  • Further sole-owners (lower concentration)
  • No ADRs found
  • Off-boarding risk: anonymized user #1

Changes since last survey

  • 25 commits — 19 feature/other, 6 fixes

By area

  • (root) — 16 commits
  • etc/tools — 3 commits
  • .github/workflows — 2 commits
  • kernel-api/src — 2 commits
  • .github/dependabot.yml — 1 commit
  • protocol/src — 1 commit

Notable commits

  • fix: Fix SBT linting and missing credentials warnings
  • fix: Fix SHA512 checksum format regression in sign-file (#285)
  • fix: Fix flaky BrokerTransformerSpec eventually timeout
  • fix: Fix flaky size test in TaskManagerSpec (#283)
  • fix: Fix pip artifact names to include incubating suffix
  • fix: Fix stale toree-assembly artifact names in release-build.sh (#286)
  • change: Add license and description to assembly POM (#257)
  • change: Add missing ASF license header to dependabot.yml
  • change: Align release tooling with ASF distribution policy
  • change: Allow a release to be prepared from another repository
  • change: Attribute SparkMonitor-derived code in spark-monitor-plugin (#282)
  • change: Avoid duplicate CI runs for branches pushed to the main repo
  • change: Compile against the Scala versions Spark ships (#287)
  • change: Fork tests so Mockito's inline mock maker attaches per JVM (#284)
  • change: Prefix binary and source distributions with apache
  • change: Prepare for next development iteration 0.6.0.dev0
  • change: Prepare release 0.6.0-incubating
  • change: Read the GPG passphrase from MAVEN_GPG_PASSPHRASE in Makefile
  • change: Refactor release script to separate prepare from publish
  • change: Refer to Apache Toree (Incubating) (#288)
  • …and 5 more

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

apache/incubator-toree was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 28 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit c9437e38ae3286947accc0a61f431de9e694851f — the exact code this score is about.
  • Scored under rubric-2026.09.16 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-2d9048c36d26.