Skip to content
CAI
Software that uses CAICheck a score

apache/texera

48.8

Weak · 20 September 2026

156.1k

lines of production code

Scala

with TypeScript, Python

1

measurement over time

CAI band scale
CAI lens gauges

What this system is

This system is a distributed, cloud-native data workflow platform that enables users to build, execute, and manage complex data pipelines through a visual interface or Python SDK. It features a high-performance execution engine with support for pipelined and materialized processing, fault tolerance via checkpointing, and integration with Apache Iceberg for scalable data storage. The platform also includes an LLM-powered agent service for conversational workflow editing and management, alongside comprehensive microservices for access control, file storage, and computing unit administration.

How it got here

2016–2025 — Apache Texera incubation and microservices migration

128 changes.

The project was rebranded from Apache Amber to Apache Texera and formally entered the Apache Incubator, establishing standard governance and licensing frameworks. This period focused on migrating the monolithic architecture into a set of standalone microservices, including access control, computing unit management, and file services, while integrating LakeFS and Iceberg for scalable data storage. Significant engineering efforts were also directed at modernizing the execution engine by migrating from Akka to Pekko, implementing Arrow Flight for Python workers, and adding comprehensive test coverage for the new distributed components.

2026 — Agent service launch and Amber engine stabilization

117 changes.

This period introduced the agent-service microservice for LLM-driven workflow management and established a new coordinator architecture for the Amber engine. It was heavily focused on comprehensive test coverage across the engine, operators, and storage layers, alongside significant infrastructure improvements like centralized configuration and modular Kubernetes deployments.

Features

Add Arrow Flight end-to-end micro-benchmark

A new end-to-end micro-benchmark has been added to measure the performance of the Arrow Flight data path through a live PythonWorkflowWorker actor. This benchmark times the round-trip of DataFrames sent to and echoed from a real Python subprocess via Arrow Flight, outputting results to CSV and JSON files for CI tracking. It supports a configurable sweep grid (controlled by BENCH\_MODE) to balance between quick PR checks and comprehensive daily performance monitoring.

amber/src/bench · high confidence

Add ErrorUtils for centralized error handling and reporting

A new ErrorUtils object has been added to the amber error module to centralize error processing logic. This utility provides methods for safely catching exceptions, generating console error messages with source location details, creating structured control errors for different languages (Scala and Python), and reconstructing throwables from serialized error data. It also includes helpers for extracting operator and worker identity information from actor IDs, improving consistency in how errors are reported and handled within the system.

amber/src/main/scala/org/apache/texera/amber/error · high confidence

Add JSON Schema generation support for workflow definitions

The workflow-operator now includes a bundled JSON Schema generator (derived from mbknor-jackson-jsonschema) under the \com.kjetland.jackson.jsonSchema\ package. This addition introduces a \JsonSchemaGenerator\ and a suite of annotations (such as \JsonSchemaBool\, \JsonSchemaDefault\, \JsonSchemaDescription\, and \JsonSchemaInject\) that allow developers to define and customize JSON schemas for workflow data models. Users can now generate structured schema definitions for their workflow configurations, enabling better validation and tooling support.

common/workflow-operator/src/main/scala/com · high confidence

Add Python support for large binary data types

Users can now handle large binary objects (such as files stored in S3) directly within Python workflows. A new \largebinary\ model class has been introduced in the core type system, allowing developers to create, read, and write large binary data via S3 URIs. This type is automatically tracked and cleaned up upon workflow execution completion, enabling seamless integration of large file handling into existing pipeline logic.

amber/src/main/python/core/models/type · high confidence

Add Sign in with Apple and ORCID authentication providers

Users can now log in to Texera using their Apple or ORCID accounts. The new Apple integration verifies the identity token against Apple's public keys and provisions accounts, requiring a verified email address unless the user is on a Work & School account (in which case they are prompted to add an email after sign-in). The ORCID integration supports standard OAuth authorization-code flow, provisioning accounts with the user's ORCID iD and name, also requiring post-sign-in email collection since ORCID does not provide an email under the requested scope. These changes are implemented in the auth resource layer (AppleAuthResource, OrcidAuthResource) and the shared ExternalAuthProvisioner, which handles the creation and linking of user accounts for these new external providers.

amber/src/main/scala/org/apache/texera/web/resource/auth · high confidence

Add user quota and storage usage API endpoints

Introduces the UserQuotaResource to expose user storage quotas via REST endpoints. Users can now retrieve details on their created datasets and workflows, as well as view granular storage usage statistics (result, runtime stats, and log sizes) for their workflow executions through the new /quota API paths.

amber/src/main/scala/org/apache/texera/web/resource/dashboard/user/quota · high confidence

Added Flarum forum installation and configuration scripts

New shell scripts have been added to the \bin/forum\ directory to automate the setup of a Flarum forum on macOS and Ubuntu. The \macos-install.sh\ and \ubuntu-install.sh\ scripts handle environment preparation (installing PHP, Apache, MySQL, and Composer), project creation, and database initialization using the provided \flarum.sql\ dump. They also configure Apache virtual hosts and apply specific Flarum extensions (\michaelbelgium/flarum-discussion-views\ and \fof/byobu\). A \start-flarum.sh\ script is included to launch the Apache server.

bin/forum · high confidence

Added ObjectMapper warmup utility for operator serialization

A new utility class, ObjectMapperUtils, has been added to the common workflow operator utilities. It provides a method to explicitly start a background thread that pre-loads logical operator metadata for serialization and deserialization, which helps prevent initial delays in application logic that relies on these operations.

common/workflow-operator/src/main/scala/org/apache/texera/amber/util · high confidence

Added cluster and single-node status listeners

The Amber engine now includes new listener components to monitor cluster membership and node availability. The ClusterListener subscribes to Pekko cluster events to detect when nodes join or leave; if a node is removed, it triggers workflow recovery (if fault tolerance is enabled) or forcefully stops affected executions, and broadcasts the current worker node count to all sessions. A corresponding SingleNodeListener is provided for environments running a single node, returning its own address when queried. These changes enhance the system's ability to maintain workflow state consistency in distributed deployments.

amber/src/main/scala/org/apache/texera/amber/clustering · high confidence

Added license collection utility for third-party dependencies

A new script, collect-licenses.ts, has been added to the agent-service/bin directory to scan node\_modules and generate a JSON manifest of third-party package names, versions, and licenses. This tool is designed to be executed after installation to produce output compatible with frontend license aggregation plugins, facilitating compliance checks and the generation of NOTICE files in CI pipelines.

agent-service/bin · high confidence

Admin dashboard now exposes user and execution management endpoints

The admin dashboard now includes new REST endpoints for managing users and viewing execution history. Administrators can list users with detailed profile information (including affiliation and joining reason), update user roles (triggering email notifications on role changes), and create placeholder user accounts. Additionally, a new execution list endpoint allows admins to view, filter, and sort workflow executions by status, name, or time, providing visibility into system usage and performance.

amber/src/main/scala/org/apache/texera/web/resource/dashboard/admin · high confidence

Amber operator environment and licensing structure established

The amber module now includes a structured set of dependency and licensing files to support its operators and development workflow. The operator-requirements.txt file defines the runtime Python dependencies for amber operators, including specific versions for pillow (12.3.0), torch (2.13.0), scikit-learn (1.7.2), and transformers (5.5.0), along with a CPU-only pin for torch on Linux x86\_64. A separate dev-requirements.txt file isolates test and development tools like pytest and ruff. Additionally, the module introduces DESCRIPTION for R UDF support, .scalafmt.conf and .scalafix.conf for code formatting and cleanup rules, and comprehensive LICENSE-binary and NOTICE-binary files to track third-party license compliance for both Java and Python components.

amber · high confidence

Centralized health check and resource access control infrastructure

The \common/resource\ module now provides a shared health check endpoint at \/healthcheck\ that returns a JSON status of 'ok', accessible without authentication. Additionally, it introduces a unified resource access control framework (\ResourceAccess\ and \ResourceTables\) that standardizes ownership and privilege checks (read/write) for datasets and models, allowing services to reuse consistent logic for verifying user permissions and listing visible resources.

common/resource · high confidence

File service relocated and initialized with LakeFS-backed storage and model management

The file-service has been moved to a new location and initialized as a standalone Dropwizard application. It now uses RustFS as the default object store (configured via docker-compose.yml) and integrates with LakeFS for dataset and model versioning. The service exposes REST APIs for managing datasets and models, including access control (grant/revoke), contributor metadata, and model file uploads with framework/format validation. It also includes staged file cleanup, health checks, and centralized LakeFS error handling.

file-service · high confidence

Initial Apache Incubator governance and project configuration

The repository is initialized with the standard Apache Software Foundation (ASF) Incubator configuration files, including \.asf.yaml\ for GitHub rulesets and metadata, \LICENSE\, \NOTICE\, and \DISCLAIMER\ files to satisfy ASF licensing requirements, and \SECURITY.md\ to define the project's security model and reporting procedures. Additionally, developer guidance documents (\AGENTS.md\, \CONTRIBUTING.md\) and build configuration files (\.gitignore\, \.dockerignore\, \.scalafmt.conf\, \.licenserc.yaml\) are added to establish the project's development workflow, coding standards, and CI/CD integration.

(repo-wide) · high confidence

Initial release of the agent-service component

This change introduces the new \agent-service\, a backend service designed to manage LLM agents. The addition includes the core configuration files required to run the service: a \bun.lock\ file establishing the dependency graph (including \elysia\, \@ai-sdk/openai\, and \pino\), a \tsconfig.json\ for TypeScript compilation, a \.prettierrc\ for code formatting, a \.dockerignore\ for container builds, and a \.env.example\ file that documents the necessary environment variables for local development and deployment.

agent-service · high confidence

Introduce Hub resource models and unified search support

The Hub dashboard now supports a unified search experience across workflows, datasets, and models by introducing new backend models (ActionType, EntityType, EntityTables) and a centralized HubResource. This change adds the ability to track user actions (view, like, clone, unlike) and manage entity-specific metadata, enabling consistent interaction patterns for all hub resources.

amber/src/main/scala/org/apache/texera/web/resource/dashboard/hub · high confidence

Introduce Iceberg-based result storage for workflow outputs

The workflow result storage layer now uses Apache Iceberg tables to persist and retrieve workflow execution results, including runtime statistics and console messages. This change introduces new schema definitions for these result types and implements read/write operations via IcebergDocument and IcebergTableWriter, enabling structured, scalable storage of workflow outputs backed by the Iceberg catalog instance.

common/workflow-core/src/main/scala/org/apache/texera/amber/core/storage/result · high confidence

Introduce LargeBinary type for S3-backed large objects

The tuple schema now supports a new LargeBinary type, allowing workflows to handle large binary data by storing references to objects in S3 rather than keeping the data in memory. This change adds the LargeBinary class, which validates and manages S3 URIs, and updates the core tuple system (Attribute, AttributeType, Schema, Tuple) to recognize and serialize this new type alongside existing types like STRING, INTEGER, and BINARY.

common/workflow-core/src/main/scala/org/apache/texera/amber/core/tuple · high confidence

Introduce Python LSP server container and Kubernetes deployment

This change adds the infrastructure to run the Python Language Server Protocol (pylsp) as a standalone service. It includes a Dockerfile based on Python 3.12 that installs python-lsp-server (v1.14.0) and configures a startup script to run pylsp on port 3000 with automatic restarts every 5 minutes. Additionally, a Kubernetes Deployment and Service manifest are provided to deploy 8 replicas of this service within the 'texera' namespace, exposing port 3000 for client connections.

bin/pylsp · high confidence

Introduce Python Virtual Environment management and real-time installation streaming

This change adds the backend infrastructure for Python Virtual Environments (PVEs) in the Amber web resource layer. PveManager handles the lifecycle of isolated Python environments—including creation, installation of user-defined packages, and resolution of system packages from requirements.txt—while persisting environment metadata to the database. PveResource exposes REST endpoints for listing, saving, updating, and deleting PVEs, as well as fetching installed packages. Additionally, PveWebsocketResource provides a WebSocket endpoint that streams pip installation logs to the frontend in real time, allowing users to monitor the progress of environment setup and package installation.

amber/src/main/scala/org/apache/texera/web/resource/pythonvirtualenvironment · high confidence

Introduce Python worker and stoppable utility components

This change adds the core Python worker implementation (\PythonWorker\) which manages network communication, heartbeats, and the main processing loop via dedicated threads. It also introduces a new \stoppable\ utility package containing the \Stoppable\ protocol and \StoppableQueueBlockingRunnable\, a helper class designed to safely interrupt runnables blocked on queue operations by injecting a specific stop marker.

amber/src/main/python/core, amber/src/main/python/core/util/stoppable · high confidence

Introduce Python worker architecture managers

The Python worker runtime now includes a set of dedicated manager classes in the \core.architecture.managers\ package to handle core execution concerns. This adds a \Context\ class that orchestrates the worker lifecycle, including a \StateManager\ that enforces a defined state-transition graph (allowing direct READY to COMPLETED transitions) and tracks state versions to prevent stale reports. An \ExecutorManager\ handles UDF code loading into isolated temporary modules with process-unique names to avoid import collisions, and properly cleans up resources on close. Additional managers manage pause/resume logic (\PauseManager\), exception reporting (\ExceptionManager\), embedded control message alignment (\EmbeddedControlMessageManager\), debug command handling (\DebugManager\), and worker statistics (\StatisticsManager\).

amber/src/main/python/core/architecture/managers · high confidence

Introduce Python-based ProxyClient and ProxyServer for Arrow Flight communication

Added new Python modules in the core proxy package that implement a ProxyClient and ProxyServer built on Apache Arrow Flight. The ProxyClient provides methods to call remote actions and send data batches, handling timeouts and optional handshakes, while the ProxyServer exposes endpoints for data ingestion, internal control messages (such as heartbeat and shutdown), and actor messages, including a decorator for standardized acknowledgment responses.

amber/src/main/python/core/proxy · high confidence

Introduce Python-based input and output managers for Amber

New Python modules (\input\_manager.py\ and \output\_manager.py\) have been added to the Amber architecture to handle data flow management. The input manager now processes data payloads, including support for \StateFrame\ objects that carry loop counters, and manages threads for reading materialized input ports. The output manager handles data partitioning (e.g., round-robin, hash-based shuffle, broadcast) and manages background threads for writing output tuples and state to storage, with specific logic to reset storage URIs for loop-end operators.

amber/src/main/python/core/architecture/packaging · high confidence

Introduce Texera Agent for conversational workflow editing

The agent-service now includes a new Texera Agent that manages LLM-driven interactions for building and modifying dataflows. This component introduces an in-memory WorkflowState to track operators, links, and positions, alongside a WorkflowResultState to cache execution results per conversation step. The agent exposes tools for adding, modifying, and deleting operators, as well as executing them, while maintaining a ReAct step tree to handle conversation history and branching. System prompts are dynamically generated based on available operators and include specific guidance for Python and R UDFs, ensuring the agent understands the dataflow context and available tools.

agent-service/src/agent · high confidence

Introduce \`pyb\` macro for safe Python code generation with encodable string support

The PyBuilder module now provides a \pyb"..."\ Scala macro for constructing Python source code templates. This feature distinguishes between standard Python literals and 'encodable' strings (marked with \@EncodableStringAnnotation\), which are base64-encoded at compile time and rendered as runtime decode expressions to prevent syntax errors or injection issues. The macro includes compile-time boundary validation to ensure encodable splices are not placed inside quotes, comments, or adjacent to invalid characters, and adds runtime guards for nested builders to maintain safety when the content is determined dynamically.

common/pybuilder/src/main · high confidence

Introduce agent-service for LLM agent management and workflow execution

The new agent-service provides the backend infrastructure for the Texera Agent, enabling it to manage LLM agents, execute workflows, and edit workflows. This service introduces a REST API for agent lifecycle management (creation, listing, deletion) and a WebSocket protocol (\/agents/:id/react\) for real-time interaction, including sending prompts, stopping execution, and receiving live step and status updates. It handles authentication by delegating to the LLM gateway as the user, configures logging via environment variables (including a new \TEXERA\_SERVICE\_LOG\_LEVEL\), and includes comprehensive test coverage for the server, logging, and WebSocket interactions.

agent-service/src · high confidence

Introduce asynchronous RPC client and server for Amber

The Amber engine now includes a new asynchronous RPC infrastructure consisting of an AsyncRPCClient, AsyncRPCServer, and AsyncRPCHandlerInitializer. This change adds the underlying communication layer that allows the engine to send and receive control invocations asynchronously, replacing or supplementing previous synchronous mechanisms. Users benefit from improved responsiveness and non-blocking control flow within the Amber execution environment.

amber/src/main/python/core/architecture/rpc · high confidence

Introduce centralized authentication and authorization module

The \common/auth\ module now provides a unified authentication stack for all Texera services. It registers a JAX-RS filter that validates Bearer JWTs, injects the authenticated user into the request context, and enforces \@RolesAllowed\ annotations on endpoints. The module also includes a startup-time enforcer that fails if any HTTP endpoint lacks an explicit access-control annotation, ensuring consistent security policies across the platform.

common/auth/src/main/scala/org/apache/texera/auth · high confidence

Introduce centralized bin/ directory with dev, build, and deployment tooling

The repository now includes a \bin/\ directory that serves as the single entry point for development, building, and deployment workflows. This change introduces \local-dev.sh\ and \single-node.sh\ as unified wrappers for the local Docker dev stack and single-node deployment, respectively. It adds \build-images.sh\ and \merge-image-tags.sh\ to handle multi-architecture Docker image builds and manifest merging. Code generation is streamlined via \frontend-proto-gen.sh\ and \python-proto-gen.sh\, while \fix-format.sh\ provides a unified interface for running Scala, frontend, and Python formatters (including Ruff for Python). The directory also contains configuration for a LiteLLM proxy (\litellm-config.yaml\), a Flarum forum setup (\config.php\, \.htaccess\), and shared utility scripts for logging and path resolution.

bin · high confidence

Introduce input port materialization reader and port storage writer runnables

Added new runnable components in the storage layer to handle materialized state and tuple data for input ports. The \InputPortMaterializationReaderRunnable\ reads persisted states and tuples from storage, applies partitioning logic (such as round-robin, hash-based, or broadcast) to batch data for downstream workers, and emits control messages to manage channel alignment. The \PortStorageWriter\ acts as a queue-blocking runnable that receives tuple elements and writes them to a buffered item writer, enabling efficient persistence of input port data.

amber/src/main/python/core/storage/runnables · high confidence

Introduce notebook-migration-service microservice for read-only Jupyter notebook viewing

A new backend microservice has been added to manage the lifecycle of Python notebooks generated from Texera workflows. This service provides REST endpoints to store, retrieve, and delete notebook files and their mappings, while also computing per-request Jupyter iframe URLs for secure, read-only embedding. It enforces JWT authentication and role-based access control, ensuring that only users with write access to a workflow can interact with its notebook. The service integrates with Jupyter (either a shared instance or per-user Kubernetes pods) to provision environments, handle token derivation, and support cross-origin communication for cell highlighting within the Texera application.

notebook-migration-service · high confidence

Introduce pyright-language-service as a standalone language server component

A new \pyright-language-service\ module has been added to provide Python language support via the Pyright language server. This component includes the necessary source files (\main.ts\, \language-server-runner.ts\, \server-commons.ts\) to initialize and run the server, along with configuration (\config.json\) and build settings (\tsconfig.json\). It establishes the infrastructure for connecting to the Pyright server process and handling WebSocket communication, effectively relocating and renaming the previous \core/pyright-language-server\ functionality into this dedicated service location.

pyright-language-service · high confidence

Introduce storage model abstractions for file and buffered I/O operations

The storage model layer now includes new abstractions for handling file-based resources and buffered writes. \VirtualDocument\ and \ReadonlyVirtualDocument\ define the interface for read/write and read-only operations on single resources, while \LakeFSFileDocument\ implements reading versioned files from LakeFS via presigned URLs or direct fetch. \BufferedItemWriter\ provides a trait for buffering items before flushing to storage, and \OnVersionedFileResource\ exposes repository, version, and path details for versioned files. \ReadonlyLocalFileDocument\ adds support for reading local files as read-only documents.

common/workflow-core/src/main/scala/org/apache/texera/amber/core/storage/model · high confidence

Introduce warehouse management REST API with decoupled catalog naming

Adds a new REST endpoint for per-user warehouse management, allowing users to check feature status, list their warehouses, and create or delete them. The implementation introduces a stable, system-generated Lakekeeper catalog name (based on user and warehouse IDs) that is decoupled from the user-facing display name, ensuring that renaming a warehouse does not break existing data URIs. The API response now includes owner details (name and avatar) for each warehouse to support future shared-warehouse scenarios.

amber/src/main/scala/org/apache/texera/web/resource/dashboard/user/warehouse · high confidence

Introduces type definitions for the agent-service

The agent-service now includes a dedicated types module (\agent-service/src/types\) that defines the data structures for agent states, execution results, and workflow content. This adds TypeScript interfaces for \AgentState\, \ReActStep\, \AgentSettings\, and \OperatorResultSummary\, alongside detailed workflow schemas like \WorkflowContent\ and \OperatorPredicate\. These types establish the contract for how the agent service manages LLM agents, handles tool calls, and interacts with workflow execution data.

agent-service/src/types · high confidence

Introduction of Amber workflow executor core components

The workflow execution engine now includes the foundational components for the Amber project, located in the \org.apache.texera.amber.core.executor\ package. This adds the \OperatorExecutor\ trait defining the standard lifecycle and tuple processing interface, the \SourceOperatorExecutor\ trait for source nodes, the \ExecFactory\ for instantiating executors from Java code or class names, and \JavaRuntimeCompilation\ to support dynamic Java UDF compilation at runtime. These changes establish the core execution model for the Amber workflow system.

common/workflow-core/src/main/scala/org/apache/texera/amber/core/executor · high confidence

Introduction of database access layer with connection pooling and centralized settings management

This change introduces the core database access components in the common/dao module, including a new SqlServer class that manages a PostgreSQL connection via a HikariCP pool to reduce connection overhead, and a SiteSettings accessor that centralizes reading and writing of key-value configuration entries. It also adds SQLSTATE constants for error handling and a jOOQ configuration file to generate data access objects from the texera\_db schema.

common/dao/src/main · high confidence

Introduction of materialized execution mode and operator output port caching

The workflow core now supports a new execution mode, allowing workflows to run in either pipelined or materialized mode (configurable via WorkflowSettings). In materialized mode, the system computes deterministic cache keys for operator output ports based on their upstream sub-DAGs, enabling the reuse of previously computed results and reducing redundant computation.

common/workflow-core/src/main/scala/org/apache/texera/amber/core/workflow · high confidence

Introduction of timed buffer for console message batching

The buffer utility module now includes a TimedBuffer implementation that batches ConsoleMessage objects. This buffer flushes messages either when a configurable maximum number of messages is reached or when a specified time interval elapses, helping to reduce the frequency of output operations for console logging.

amber/src/main/python/core/util/buffer · high confidence

LakeFS storage client and storage utilities added

The system now includes a new LakeFS storage client implementation and associated storage utility functions. The client provides methods for initializing repositories, writing and retrieving files, generating presigned URLs, and handling multipart uploads, while the utilities offer helper functions for managing read/write locks.

common/workflow-core/src/main/scala/org/apache/texera/amber/core/storage/util · high confidence

New API client library for agent-service backend interactions

The agent-service now includes a dedicated API client layer in \agent-service/src/api\ to communicate with backend services. This adds support for authenticating requests via JWT bearer tokens (with user extraction and validation), persisting and retrieving workflows with proper serialization, compiling logical plans into physical plans, and fetching operator metadata. The implementation ensures that authentication headers are correctly passed for workflow operations and that configuration endpoints are safely exposed to callers.

agent-service/src/api · high confidence

New Amber engine coordinator architecture

The Amber workflow engine introduces a new coordinator architecture to manage execution, recovery, and scheduling. This change adds a \Coordinator\ actor that orchestrates workflow execution via a \CoordinatorProcessor\ and \WorkflowScheduler\, which generates schedules using a \CostBasedScheduleGenerator\. The new system tracks execution state through a hierarchy of \WorkflowExecution\, \RegionExecution\, \OperatorExecution\, and \WorkerExecution\ objects, enabling detailed metrics aggregation and state tracking. It also introduces a \GlobalReplayManager\ to handle workflow recovery status and a \CoordinatorTimerService\ for periodic status updates and statistics persistence. Communication is standardized through a new \ClientEvent\ trait and its implementations (e.g., \ExecutionStateUpdate\, \OperatorPortResultUriAvailable\), and RPC handling is managed by \CoordinatorAsyncRPCHandlerInitializer\ with specific handlers like \AdvanceRegionExecutionsHandler\.

amber/src/main/scala/org/apache/texera/amber/engine/architecture/coordinator · high confidence

New Amber execution state and command handling components

The Amber module introduces new infrastructure for managing execution state and processing internal commands. In the web layer, Scala classes (StateStore, ExecutionStateStore, WorkflowStateStore, ExecutionReconfigurationStore) provide reactive, thread-safe storage for execution metadata, statistics, console output, breakpoints, and reconfiguration data, exposing diffs as WebSocket events. In the Python core, new actor command handlers (ActorCommandHandler, BackpressureHandler, CreditUpdateHandler) define how the engine processes backpressure and credit update signals, with the backpressure handler directly controlling data flow on internal queues.

amber/src/main/python/core/architecture/handlers/actorcommand, amber/src/main/scala/org/apache/texera/web/storage · high confidence

New Dockerfiles for all services and a privileged mounter for LakeFS storage

The repository now provides dedicated Dockerfiles in bin/dockerfiles/ for every Texera service, including the access-control, agent, config, file, notebook-migration, workflow-compiling, and workflow-computing-unit-managing services, as well as the computing-unit master and worker images. The texera-web-application image now builds the frontend with Yarn 4 and merges frontend license artifacts into the final image. A new mounter service (bin/mounter/mounter.py) and its Dockerfile introduce a per-node privileged DaemonSet that performs GeeseFS FUSE mounts on behalf of unprivileged computing-unit pods, improving security by keeping user code unprivileged while still providing access to LakeFS-backed storage. Tests for the mounter are also included.

bin/dockerfiles · high confidence

New Kubernetes Helm chart for Texera deployment

The \bin/k8s\ directory now contains a complete Helm chart (\texera-helm\) for deploying Texera on Kubernetes. This chart introduces RustFS as the default in-cluster object store (replacing the previous MinIO setup) and integrates Lakekeeper as the Iceberg catalog. It supports external S3 storage via configurable overlays (e.g., \values-aws.yaml\) and includes dependencies for PostgreSQL (using the groonga/pgroonga image), Envoy Gateway, and Lakekeeper. The chart also provides SQL scripts for database initialization and configuration files for development and AWS environments.

bin/k8s · high confidence

New WebSocket event models for execution monitoring and pagination

The Amber module now includes a set of new Scala case classes in the \org.apache.texera.web.model.websocket.event\ package to support richer real-time updates via WebSocket. These new event types allow the frontend to receive detailed execution statistics (\OperatorStatisticsUpdateEvent\), cache status changes (\CacheStatusUpdateEvent\), execution duration and running state (\ExecutionDurationUpdateEvent\), worker assignment updates (\WorkerAssignmentUpdateEvent\), and region state changes (\RegionStateEvent\). Additionally, support for paginated query results (\PaginatedResultEvent\) and Python console output (\ConsoleUpdateEvent\) has been added, with the base \TexeraWebSocketEvent\ trait updated to register these new types for JSON serialization.

amber/src/main/scala/org/apache/texera/web/model/websocket/event · high confidence

New WebSocket request models for workflow control and pagination

The Amber module introduces a set of new Scala case classes for handling WebSocket requests, including \ResultPaginationRequest\ to support wide-column table pagination and search, \WorkflowExecuteRequest\ with optional warehouse targeting, and control requests for workflow lifecycle (execute, pause, resume, kill, checkpoint) and debugging (Python expression evaluation, debug commands). These models extend \TexeraWebSocketRequest\ and are registered for Jackson polymorphic deserialization, enabling the web layer to process these specific client commands.

amber/src/main/scala/org/apache/texera/web/model/websocket/request · high confidence

New access control and repository mounting endpoints for computing units

The access-control-service now exposes two new JAX-RS resource endpoints to manage access and data integration for computing units. AccessControlResource handles authorization for specific paths, including routing requests to per-user JupyterLab pods and validating access to Python Virtual Environment (PVE) routes and execution stats. Additionally, ComputingUnitMountResource introduces a new /mounts API that allows users with WRITE access to a computing unit and READ access to a repository to mount that repository (dataset or model) onto the unit, enforcing strict validation of repository existence and commit hashes before forwarding the request to the mounter service.

access-control-service/src/main/scala/org/apache/texera/service/resource · high confidence

New agent tools for workflow CRUD and execution with structured result formatting

The agent-service now exposes tools that allow the LLM agent to create, modify, and delete operators within a workflow, as well as execute them. The \addOperator\ and \modifyOperator\ tools validate operator properties against the schema and manage input/output port connections, while \deleteOperator\ removes operators and their links. Execution is handled by \executeOperator\, which validates the workflow state, serializes access via a per-workflow mutex, and returns structured results. Operator results are formatted into a consistent text summary including table shapes, input port metadata, and warnings, with internal keys like \\_\_is\visualization\\\ and \\\_row\index\\_\ filtered from the visible output.

agent-service/src/agent/tools · high confidence

New agent-service utility modules for workflow layout, context assembly, and metadata handling

The agent-service now includes a suite of utility modules in \src/agent/util\ that support LLM agent interactions with workflows. \auto-layout.ts\ uses the dagre library to automatically position workflow operators in a left-to-right layout, mirroring frontend spacing settings. \context-utils.ts\ assembles structured Markdown context for the agent, grouping ReAct steps into completed or ongoing tasks and serializing the current dataflow (operators, links, and schemas) in topological order. \workflow-system-metadata.ts\ provides a singleton cache for operator metadata fetched from the backend, including JSON schema validation via Ajv and compact schema formatting for error messages. \workflow-utils.ts\ offers helpers to extract input port schemas from upstream links and a service to instantiate new operator predicates with default properties. Comprehensive unit tests have been added for all these utilities to ensure correct behavior.

agent-service/src/agent/util · high confidence

New collaboration, feedback, and Hugging Face integration endpoints

The web resource layer now exposes several new capabilities: real-time workflow collaboration via a WebSocket endpoint at /wsapi/collab that handles session management, command broadcasting, and lock contention; a new /feedback REST API allowing users to submit feedback and admins to review it; and a /huggingface API for browsing models, uploading audio, and proxying media. Additionally, the Gmail email service now surfaces SMTP failures to the UI, and the WebSocket payload limit has been increased to prevent dropping large result frames.

amber/src/main/scala/org/apache/texera/web/resource · high confidence

New computing-unit-managing-service for administering computing units

A new standalone service has been introduced to manage computing units, providing dedicated REST endpoints for administrators to list all active units and terminate any unit, as well as for users to manage access privileges (grant/revoke) and create units with support for curated images. The service is built on Dropwizard, uses a configurable log level via TEXERA\_SERVICE\_LOG\_LEVEL, and enforces role-based access control to ensure only authorized users can perform administrative actions or modify sharing settings.

computing-unit-managing-service · high confidence

New core utility modules for Amber

The \amber/src/main/python/core/util\ package now includes several new modules: \atomic.py\ provides a thread-safe \AtomicInteger\ class; \virtual\_identity.py\ adds helpers for parsing worker IDs and serializing/deserializing \GlobalPortIdentity\ objects; \expression\_evaluator.py\ introduces an \ExpressionEvaluator\ for evaluating Python expressions with context; \base\protocols.py\ defines protocols like \Putable\, \Getable\, and \Runnable\; and \runnable.py\ defines the \Runnable\ protocol. The \\\init\\_.py\ file exposes key classes and functions from these modules.

amber/src/main/python/core/util · high confidence

The single-node example environment now includes two new sample datasets to support testing and demonstration workflows. The \iris-species\ dataset provides 150 records of iris measurements (sepal and petal dimensions) categorized by species (setosa, versicolor, virginica), sourced from Kaggle. Additionally, the \popular-movies-of-imdb\ dataset adds 1000 records of popular movies including titles, overviews, languages, and vote averages, sourced from TMDb. These files are located in \bin/single-node/examples/datasets/\ and are intended for use with the example data loader.

bin/single-node/examples · high confidence

New licensing audit and compliance tooling in bin/licensing

This change introduces a suite of Python scripts in the bin/licensing directory to automate and enforce binary distribution compliance. The new audit\_jar\_licenses.py script scans bundled JARs to identify licenses that require per-dependency attribution files, while generate\_notice\_binary.py creates standardized NOTICE-binary files by extracting and deduplicating META-INF/NOTICE content from JARs. The check\_binary\_deps.py script validates that the actual bundled dependencies match the declared LICENSE-binary records, supporting multiple ecosystems (Java, Python, npm) and distinguishing between direct and transitive dependencies. Additionally, concat\_license\_binary.py merges per-module license files into a unified distribution-wide license record, and comprehensive unit tests have been added to verify the correctness of these new compliance tools.

bin/licensing · high confidence

New local development stack manager with TUI and Linux support

A new \bin/local-dev.sh\ entry point and \bin/local-dev/\ tooling directory have been introduced to manage the Texera local development environment. This replaces previous ad-hoc setup methods with a unified interface that supports both macOS and Linux (requiring \iproute2\ and Docker Compose v2). Key capabilities include an interactive Textual TUI dashboard (\-i\), automatic SQL schema migration before builds, JSON status output for scripting, and the ability to deploy from sibling git worktrees. The tooling also handles host LAN IP detection for RustFS connectivity, manages the Jupyter image rebuild process, and automatically offers to install missing toolchains (JDK, Node, Python) on first run.

bin/local-dev · high confidence

New partitioner implementations for data distribution semantics

The \amber/src/main/python/core/architecture/sendsemantics\ module now includes concrete implementations for distributing data tuples across workers: \BroadcastPartitioner\ (sends all data to all receivers), \OneToOnePartitioner\ (directs data to a single specific worker), \RoundRobinPartitioner\ (cycles through receivers), \HashBasedShufflePartitioner\ (routes based on tuple hash codes), and \RangeBasedShufflePartitioner\ (routes based on value ranges). These classes implement the \Partitioner\ base class to handle batching, flushing, and state management for their respective send semantics.

amber/src/main/python/core/architecture/sendsemantics · high confidence

New serialization helpers for workflow port identities

Added three new Scala classes in the serde utility package to handle serialization and deserialization of workflow port identities. GlobalPortIdentitySerde provides a custom string format for GlobalPortIdentity objects that avoids characters like underscores and slashes to prevent conflicts with VFS URI parsing, while PortIdentityKeySerializer and PortIdentityKeyDeserializer handle Jackson-based serialization of PortIdentity objects using an underscore-separated string format.

common/workflow-core/src/main/scala/org/apache/texera/amber/util/serde · high confidence

New single-node deployment with RustFS, Lakekeeper, and AI features

The \bin/single-node\ directory now provides a complete, self-contained Docker Compose deployment for running Texera on a single machine. This setup replaces the previous MinIO object store with RustFS and uses Lakekeeper as the Iceberg REST catalog. It also introduces built-in support for the Texera Agent (powered by LiteLLM and Claude Haiku) and a Python notebook migration tool (powered by JupyterLab). The deployment includes a \single-node.sh\ wrapper script for easy management, a pre-configured \.env\ file for customization, and an optional \--with-examples\ flag to launch with demo workflows and datasets.

bin/single-node · high confidence

New storage layer for Iceberg tables, datasets, and models

The storage subsystem now supports persisting workflow results, state, and runtime statistics as Iceberg tables via a new \DocumentFactory\ and \IcebergCatalogInstance\. This introduces a bounded, per-warehouse catalog cache that automatically closes idle REST catalogs to prevent resource leaks. Additionally, the system now resolves logical paths for versioned datasets and models (e.g., \/dataset/...\ or \/model/...\) into physical URIs, while maintaining backward compatibility with legacy unprefixed dataset paths.

common/workflow-core/src/main/scala/org/apache/texera/amber/core/storage · high confidence

New utility modules for avatar sanitization, email validation, and retry logic

This change introduces three new utility components in the common/util package. AvatarUtil now stores and validates full identity-provider avatar URLs, restricting accepted sources to the googleusercontent.com domain to prevent open-redirect risks. EmailUtil provides functions to validate email format and normalize addresses by trimming whitespace and lowercasing. RetryUtil offers a shared, blocking exponential-backoff retry mechanism with configurable delays, attempt limits, and interrupt handling, allowing other modules to avoid implementing custom retry loops.

common/util · high confidence

New web service layer for execution management and notifications

The web service layer in amber now includes dedicated services for handling execution console output, result persistence and export, runtime statistics, and workflow reconfiguration. A new EmailNotificationService allows users to receive email alerts when workflows complete or fail. ExecutionResultService handles the conversion of execution results to JSON, including optimized binary preview handling. ExecutionConsoleService manages real-time console updates and Python debugging interactions. ExecutionReconfigurationService enables dynamic operator logic changes during paused executions. ExecutionStatsService provides real-time operator metrics and worker assignment updates. LakekeeperClient manages per-user warehouse creation and deletion with proper async purge handling. These services collectively improve execution monitoring, result management, and workflow control capabilities.

amber/src/main/scala/org/apache/texera/web/service · high confidence

New websocket response models for cluster, region, and logic updates

The Amber module now includes new Scala case classes for websocket responses, enabling the frontend to receive structured updates for cluster status (worker count), region layer changes, heartbeat signals, logic modification validation/completion, and Python expression evaluation results.

amber/src/main/scala/org/apache/texera/web/model/websocket/response · high confidence

New workflow management REST API endpoints in Amber

The Amber module now exposes a complete set of REST resources for managing user workflows, including access control, execution history, versioning, and core workflow operations. Users can now manage workflow sharing permissions (granting/revoking read/write access), view and export execution results (with dataset access checks), browse and restore workflow versions, and update workflow metadata such as cover images and default view modes (Canvas vs. Form). These endpoints handle authentication, privilege checks, and database persistence for workflow lifecycle management.

amber/src/main/scala/org/apache/texera/web/resource/dashboard/user/workflow · high confidence

Python Iceberg storage backend with REST catalog and Windows support

The Python worker in Amber now supports storing results in Apache Iceberg tables using a new Python-based storage implementation. This adds support for REST catalogs, allowing multiple warehouses to be managed within a single process, and fixes local storage paths on Windows by normalizing drive letters to file URIs. It also introduces support for large binary data types by encoding them with a specific suffix in the Iceberg schema.

amber/src/main/python/core/storage/iceberg · high confidence

Python SDK exposes loop operators and large binary I/O streams

The pytexera Python SDK now exports LoopStartOperator and LoopEndOperator, enabling users to construct iterative workflows directly in Python. Additionally, the SDK exposes LargeBinaryInputStream and LargeBinaryOutputStream along with the largebinary type, allowing for the handling of large binary data inputs and outputs within Python-based operators.

amber/src/main/python/pytexera · high confidence

Python UDFs now support configurable UI parameters

Developers can now define and configure parameters for Python User-Defined Functions directly from the Amber UI. This change introduces the \UDFOperatorV2\ base class and a \UiParameter\ API that allows operators to declare typed parameters (e.g., integers, strings) which are automatically injected and parsed at runtime. The implementation includes validation to prevent type conflicts and warnings for unused parameters, ensuring that user inputs from the UI are correctly applied to operator logic.

amber/src/main/python/pytexera/udf · high confidence

Python storage layer for datasets and large binaries

The Python storage module now provides classes to manage dataset files and large binary objects. DatasetFileDocument parses logical paths (supporting both new resource-type prefixes and legacy unprefixed paths for backward compatibility) and retrieves files via a presigned URL endpoint with configurable timeouts and retries. LargeBinaryInputStream and LargeBinaryOutputStream enable lazy, streaming access to large binary data stored in S3, with background multipart uploads and proper resource cleanup.

amber/src/main/python/pytexera/storage · high confidence

Python storage layer now uses Iceberg with warehouse-scoped URIs and S3 binary support

The Python storage module in \amber/src/main/python/core/storage\ has been implemented to manage data persistence via Iceberg tables, replacing or supplementing previous mechanisms. This change introduces a \DocumentFactory\ that creates, opens, and checks for the existence of documents by mapping workflow and execution IDs to specific Iceberg namespaces (for results and state) and storage keys. URIs are now structured with optional warehouse prefixes (\/wh/\<name\>\) to support per-warehouse catalog routing, ensuring that data writes are correctly scoped. Additionally, the \StorageConfig\ class now includes configuration for S3 endpoints and a dedicated base URI for large binary storage, enabling the system to handle large binary data types by writing them to execution-scoped S3 paths.

amber/src/main/python/core/storage · high confidence

Support for storing and streaming large binary data via S3

The workflow core now includes utilities to handle data objects larger than 2GB by storing them in an S3-compatible bucket. This change introduces a LargeBinaryManager to organize these objects by execution ID, along with LargeBinaryInputStream and LargeBinaryOutputStream classes that enable lazy reading and background multipart uploading of large binary payloads. An S3StorageClient abstraction is also added to manage bucket creation, object uploads, and paginated directory deletions, ensuring that large binary storage is scoped and cleaned up correctly per execution.

common/workflow-core/src/main/scala/org/apache/texera/service · high confidence

Texera Web server restructure and new AI assistant feature

The web server application has been reorganized into distinct entry points for the main web interface (TexeraWebApplication) and the computing unit master (ComputingUnitMaster), introducing a new StaticAssetCacheFilter to optimize frontend performance by setting immutable cache headers on fingerprinted assets. Additionally, a new AI Assistant feature is available, allowing users to request Python type annotations via an integrated OpenAI API endpoint, supported by a new Python-based type annotation visitor and its corresponding tests.

amber/src/main/scala/org/apache/texera/web · high confidence

Unified dashboard search across workflows, datasets, and models

The dashboard now provides a single search endpoint that queries workflows, datasets, and models together. Users can filter results by resource type, keywords, creation/modification dates, owners, and specific IDs, and sort by name, creation time, edit time, or execution time. The search supports full-text matching (using PGroonga or standard text search) and includes owner email information in the results.

amber/src/main/scala/org/apache/texera/web/resource/dashboard · high confidence

Workflow Compiling Service introduces standalone Python script export

The workflow-compiling-service now includes a new endpoint that allows users to export a workflow as a standalone Python script. This feature is implemented via the \WorkflowToPythonResource\ and \WorkflowToPythonTranslator\, which convert the workflow's logical plan into executable Python code using pandas. The service also adds a new compilation endpoint (\WorkflowCompilationResource\) and enforces role-based access control on these resources. Additionally, the service configuration now supports log level control via the \TEXERA\_SERVICE\_LOG\_LEVEL\ environment variable.

workflow-compiling-service · high confidence

Workflow export to standalone Python scripts

The workflow operator now supports exporting workflows as standalone Python scripts. This change introduces the \StandaloneCodeGenerator\ trait and \StandaloneHelpers\ object, which allow operators to produce self-contained Python code with shared type-casting and data-handling utilities. The \LogicalOp\ file registers the new \FileScanOp\ and \FileScanSourceOp\ operators, and the \PortDescriptor\ includes backward-compatibility handling for multi-input ports, ensuring exported scripts can run independently of the Texera engine.

common/workflow-operator/src/main/scala/org/apache/texera/amber/operator · high confidence

Architecture

Amber Python control-plane handlers relocated and restructured

The control-plane message handlers for the Amber Python engine have been moved into a new \amber/src/main/python/core/architecture/handlers/control\ directory. This change reorganizes the codebase by splitting the previous monolithic handler logic into distinct, single-responsibility classes (such as \StartWorkerHandler\, \EndWorkerHandler\, and \DebugCommandHandler\) that inherit from a new \ControlHandler\ base. For users, this represents an internal architectural cleanup that improves code maintainability and clarity without altering the external behavior of the data processing pipeline.

amber/src/main/python/core/architecture/handlers/control · high confidence

Amber worker promise handlers restructured into traits

The worker promise handlers in the Amber engine have been refactored from monolithic classes into separate Scala traits (e.g., AddInputChannelHandler, StartHandler, EndHandler). This change organizes the worker's lifecycle and control logic into modular components, improving code maintainability and separation of concerns within the engine's architecture.

amber/src/main/scala/org/apache/texera/amber/engine/architecture/worker/promisehandlers · high confidence

Behavioural changes

Access control service relocated and expanded with LakeFS mount support

The access-control-service has been moved to the org.apache.texera.service package and now includes new capabilities for LakeFS repository mounts. This change introduces the ComputingUnitMountResource and supporting utilities (ComputingUnitNodeLocator, MounterClient) that allow authorized users to mount repositories by locating the correct Kubernetes node and authenticating with the privileged node mounter. Additionally, the service now enforces role annotations at startup, records user activity via a dedicated event listener, and configures request logging through the TEXERA\_SERVICE\_LOG\_LEVEL environment variable.

access-control-service/src/main/scala/org/apache/texera/service · high confidence

Aligns Amber's JWT authentication with microservices claim parsing

Amber now uses a shared JWT claim parser to construct the SessionUser object, ensuring that authentication and authorization logic (including role checks) behaves identically to the microservices. This change standardizes how user identity and roles are extracted from tokens across the platform.

amber/src/main/scala/org/apache/texera/web/auth · high confidence

Amber engine RPC protocol and state definitions migrated to Apache Texera

The Amber engine's internal communication contracts have been updated to the \org.apache.texera.amber\ namespace, introducing new Protobuf definitions for the \CoordinatorService\ and \WorkerService\. These changes add RPC methods for workflow reconfiguration, jump-to-operator debugging, and region execution advancement, while also introducing a monotonic \state\_version\ field to worker state messages to ensure causal ordering of state updates by the coordinator.

amber/src/main/protobuf · high confidence

Amber engine client interface migrated to Apache Texera package

The Amber engine's client-side components, specifically the AmberClient and ClientActor, have been moved from the org.apache.amber namespace to org.apache.texera.amber. This change updates the package declarations and internal imports to reflect the project's rebranding, ensuring the client actor correctly initializes the Coordinator and handles control requests within the new organizational structure.

amber/src/main/scala/org/apache/texera/amber/engine/common/client · high confidence

Amber engine common utilities and runtime initialization

The Amber engine's common module now includes foundational utilities and runtime initialization logic. This adds a centralized configuration object for the master node address, a Kryo initializer that registers specific serializers for lambda and closure handling, and a logging trait that formats log messages with actor identity. The runtime module provides the core mechanism to start and manage the Pekko actor system for both master and worker nodes, including cluster seed node configuration. Additionally, new checkpoint support classes enable the serialization and deserialization of operator states, while a Fries reconfiguration algorithm is introduced to handle workflow updates by computing minimal cut sets. Utility functions for exponential-backoff retries, path resolution, and workflow state mapping are also included.

amber/src/main/scala/org/apache/texera/amber/engine/common · high confidence

Amber engine migrates actor infrastructure from Akka to Apache Pekko

The Amber engine's core architecture components—including the actor service, message transfer, and reference mapping services—have been migrated from Akka to Apache Pekko. This change updates the underlying actor framework used for workflow execution, requiring the engine to use Pekko-specific imports and APIs for actor lifecycle, scheduling, and remote communication.

amber/src/main/scala/org/apache/texera/amber/engine/architecture/common · high confidence

Build system relocated to project/ with updated tooling and licensing automation

The root build configuration has been moved into a dedicated project/ directory, introducing new SBT plugins (scalafmt, scalafix, jacoco, license-report, native-packager, fs2-grpc) and updating core build dependencies: sbt to 1.12.9, jOOQ codegen to 3.19.36, Typesafe Config to 1.4.9, and the PostgreSQL driver to 42.7.13. This relocation includes automated license handling that generates per-module META-INF/LICENSE files (including specific MIT attribution for the workflow-operator module) and ships LICENSE-binary, NOTICE-binary, DISCLAIMER, and the licenses/ directory in distribution zips. Additionally, a new test filtering mechanism (TestFilters) enables CI jobs to split tests into unit and integration subsets using environment variables and ScalaTest tags.

project · high confidence

Centralized configuration management for platform services

The platform now uses a unified set of Scala configuration objects in the common config module to manage settings for authentication, storage, Kubernetes computing units, and the GUI. This change introduces support for new login providers (ORCID, Apple) and features like per-user warehouses and out-of-pod dataset mounting, while also replacing the previous manual locking mechanism for the JWT secret with a thread-safe lazy initialization.

common/config/src/main/scala/org/apache/texera/common · high confidence

Centralized console message reporting for errors and print statements

The Amber engine now uses a unified console message system for operator-facing output. Uncaught exceptions are captured and reported as structured ERROR messages including source location and full tracebacks, while Python print statements are intercepted via a context manager to be logged as structured PRINT messages with caller metadata. All console messages are timestamped using the system's local timezone to ensure consistent chronological ordering in the UI.

_amber/src/main/python/core/util/console\message · high confidence

Centralized default configuration for Texera services

The system now ships with a unified set of default configuration files in the common config module, establishing baseline settings for the engine, GUI, authentication, storage, and Kubernetes computing units. Key defaults include enabling the Form View by default, disabling the AI Copilot and Python notebook migration features, and configuring the Iceberg catalog to use the REST type. Authentication defaults allow local, Google, ORCID, and Apple logins (with ORCID and Apple off by default), and the cluster runtime has been switched to use Apache Pekko.

common/config/src/main/resources · high confidence

Config service relocated and restructured with new API endpoints

The config-service has been moved to a new location and restructured as a standalone Dropwizard application. It now exposes configuration via specific API endpoints: /config/pre-login for unauthenticated access (including login provider flags and invite-only status), /config/gui for authenticated GUI settings, /config/amber for engine configs, and /config/settings for admin management of site settings. The service initializes by preloading default settings into the database and enforces role-based access control on all endpoints.

config-service · high confidence

Configurable service logging and new resource files for Amber

The Amber service resources now include configuration files for the computing-unit-master, compiling-service-web, and web components, alongside a new cache configuration and logback setup. Log levels for these services are now controllable via the TEXERA\_SERVICE\_LOG\_LEVEL environment variable, defaulting to INFO, and all log output is consistently directed to the logs/ directory.

amber/src/main/resources · high confidence

Database schema migration and Liquibase integration

The SQL layer now uses Liquibase for version-controlled schema management, introducing a \changelog.xml\ that orchestrates 27 database updates (IDs 23–50). These changes add support for new features including ORCID and Apple authentication providers, per-user warehouse registrations, model hub engagement tracking, and custom workflow cover images, while also refactoring the authentication model to store full avatar URLs and decoupling the Lakekeeper catalog name from the user-facing name.

sql · high confidence

Helm templates reorganized into base, aws, and on-prem directories with configurable S3 storage

The Kubernetes Helm templates have been restructured into a modular layout: base templates for shared resources, aws/ for AWS-specific configurations (like external S3 credentials), and on-prem/ for local deployment resources. This change introduces configurable object storage via the storage.s3 values, allowing deployments to use external S3-compatible stores instead of the default in-cluster RustFS. The access-control-service now supports a dedicated service account for secure repository mounting, and the gateway configuration has been updated to support dynamic routing with proper traffic policies for agent services and Jupyter notebooks.

bin/k8s/templates · high confidence

Improved Windows compatibility and Arrow timestamp handling in workflow utilities

The workflow-core utilities now support Iceberg REST catalogs for result storage and handle Arrow timestamps as wall-clock times to ensure consistent behavior across different server time zones. On Windows, local Iceberg storage no longer requires a winutils installation by using a custom file system that skips unsupported POSIX permission operations. Additionally, JSON flattening now correctly emits array-of-primitive entries, and Arrow integer parsing rejects non-standard bit widths to prevent silent data corruption.

common/workflow-core/src/main/scala/org/apache/texera/amber/util · high confidence

Introduce LinkedBlockingMultiQueue with keyed removal support

The \amber\ module now includes a new \LinkedBlockingMultiQueue\ implementation in \core.util.customized\_queue\. This queue supports keyed items and includes a fix for removal paths, ensuring that items can be correctly removed from the queue without leaving stale references or causing synchronization issues. The change also introduces helper classes for inner class handling and queue base interfaces to support this functionality.

_amber/src/main/python/core/util/customized\queue · high confidence

Introduce Python-based Amber worker architecture

The Amber engine now uses a new Python worker implementation located in \core/runnables\, replacing the previous internal structure. This change introduces a multi-threaded execution model where a \MainLoop\ coordinates a dedicated \DataProcessor\ thread for handling tuples, states, and internal markers, while separate \NetworkReceiver\ and \NetworkSender\ components manage external communication via a proxy server. A new \Heartbeat\ component monitors the JVM parent process health, and the \MainLoop\ includes specific logic to read loop input tables from materialized storage rather than shipping them in state, which resolves concurrency issues with object store operations.

amber/src/main/python/core/runnables · high confidence

Introduce cost-based scheduling with materialized execution support

The Amber engine's scheduling layer has been refactored to support cost-based region scheduling and materialized execution modes. The new \CostBasedScheduleGenerator\ partitions the physical plan into regions and assigns storage URIs in two passes to handle potential directed cycles in the region graph. A \DefaultCostEstimator\ now estimates region costs using past execution statistics (longest-running operator time) when available, falling back to the number of materialized ports for first-time runs. The \RegionExecutionManager\ enforces a two-phase execution scheme (dependee ports followed by non-dependee ports) to correctly handle input-port dependencies and materialized state, while \WorkflowExecutionManager\ coordinates region advancement and loop bookkeeping via materialized URIs.

amber/src/main/scala/org/apache/texera/amber/engine/architecture/scheduling · high confidence

Introduce execution-scoped large binary storage and Python Virtual Environment support

The scheduling configuration now supports executing Python UDFs inside Python Virtual Environments (PVEs) and scopes large binary storage to specific executions. Worker configurations include a \pveName\ field to identify the virtual environment and an \largeBinaryBaseUri\ that is generated per execution ID, ensuring that large binary artifacts are isolated and cleaned up by execution rather than globally.

amber/src/main/scala/org/apache/texera/amber/engine/architecture/scheduling/config · high confidence

Introduces asynchronous RPC client and server for control messaging

The engine now uses a new asynchronous RPC layer in the \org.apache.texera.amber.engine.common.rpc\ package to handle control messages. This replaces the previous manual request-response handling with a dynamic proxy-based client (\AsyncRPCClient\) that multiplexes requests over the network output gateway, and a reflective server (\AsyncRPCServer\) that routes incoming control invocations to registered handlers. This change simplifies the implementation of control commands by allowing developers to call remote methods as if they were local, with responses returned via promises.

amber/src/main/scala/org/apache/texera/amber/engine/common/rpc · high confidence

Migrate workflow core protobuf definitions to Apache Texera Amber namespace

The workflow core protocol buffer definitions (executor, virtual identity, workflow, and runtime state) have been moved and renamed from the \apache.amber\ package to \org.apache.texera.amber.core\. This change updates the protobuf package declarations and file paths to reflect the project's rebranding from Amber to Texera Amber, ensuring that generated code and internal messaging contracts align with the new namespace.

common/workflow-core/src/main/protobuf · high confidence

New HTTP request and response models for authentication and result export

This change introduces new Scala case classes that define the structure of HTTP payloads for the Amber web service. For authentication, it adds \UserLoginRequest\ for username/password login, \UserRegistrationRequest\ (which includes an optional \code\ field for email verification), \RegistrationResponse\, and \TokenIssueResponse\. For data handling, it adds \ResultExportRequest\ (supporting various export types like CSV and Google Sheets) and \ResultExportResponse\. These models serve as the contract for the API endpoints handling user signup, login, and workflow result exports.

amber/src/main/scala/org/apache/texera/web/model/http · high confidence

New materialization and storage writer threads for checkpoint state round-trips

The worker managers now include dedicated threads to handle state materialization and storage writing, ensuring that checkpoint state round-trips correctly. InputPortMaterializationReaderThread reads persisted state and tuples from storage, rebuilding loop envelopes (loop counter and start ID) so that JVM operators inside loops carry the correct context through to LoopEnd. OutputPortStorageWriterThread writes tuples to storage and captures failures to prevent silent worker completion when results are missing. StatisticsManager uses plain maps without withDefaultValue to ensure its statistics survive Kryo serialization during checkpointing.

amber/src/main/scala/org/apache/texera/amber/engine/architecture/worker/managers · high confidence

New messaging layer with credit-based flow control and congestion management

The Amber engine's messaging layer has been replaced with a new architecture that introduces credit-based flow control and congestion control to prevent network saturation. This change adds a \DeadLetterMonitorActor\ to surface dropped messages, \AmberFIFOChannel\ to guarantee exactly-once delivery semantics, and a suite of partitioners (RoundRobin, Hash-based, Range-based, Broadcast, One-to-One) to manage how data tuples are distributed to downstream workers. Additionally, an \InputManager\ now handles materialized input ports via dedicated reader threads, and a \WorkerTimerService\ enables adaptive network buffer flushing to optimize throughput.

amber/src/main/scala/org/apache/texera/amber/engine/architecture/messaginglayer · high confidence

New utility objects for computing unit access checks and HTTP header constants

This change introduces two new utility files in the auth module. ComputingUnitAccess.scala provides a centralized method to determine a user's privilege level (owner/write/none) for a specific computing unit by performing a single SQL join between the workflow and access tables, replacing potential N+1 query patterns. HeaderField.scala defines constants for standard HTTP header names used to pass user identity and access information (such as user ID, name, email, and computing unit access) in requests.

common/auth/src/main/scala/org/apache/texera/auth/util · high confidence

Python core models restructured with loop-control bookkeeping and channel markers

The Python core models in amber have been reorganized into a new package structure, introducing explicit support for loop execution and channel lifecycle management. State objects now carry loop-control bookkeeping fields (loop\_counter and loop\_start\_id) alongside user content, enabling the runtime to track loop iterations without interfering with user-defined state. New InternalMarker classes (StartChannel, EndChannel) and an enhanced InternalQueue implementation preserve input port provenance and ensure that data sub-queues registered after a disable\_data call remain disabled, preventing data leaks during pause or backpressure scenarios. Additionally, the models now support large binary types via URI-to-largebinary conversion in ArrowTableTupleProvider and provide a PythonTemplateDecoder for base64-decoded code templates in operators.

amber/src/main/python/core/models · high confidence

Python schema model supports LARGE\_BINARY and fixes numeric coercion

The Python schema model now includes a LARGE\_BINARY type, enabling the system to handle large binary data via PyArrow metadata preservation and serialization as URI strings. Additionally, the schema utilities now correctly coerce integral floats and NumPy scalar types to integers for INT and LONG fields, preventing precision loss or type errors when processing numeric data.

amber/src/main/python/core/models/schema · high confidence

Python worker communication migrated to Apache Arrow Flight

The Python worker proxy layer has been rewritten to use Apache Arrow Flight for data and control communication between the JVM and Python processes. This change introduces new proxy components (PythonProxyClient and PythonProxyServer) that handle data serialization via Arrow schemas, support state frames with loop-envelopes, and transmit embedded control messages. Additionally, the Python worker startup configuration is now Base64-encoded to ensure safe transmission as a command-line argument across platforms, particularly resolving Windows argv quoting issues.

amber/src/main/scala/org/apache/texera/amber/engine/architecture/pythonworker · high confidence

Python worker startup config now uses Base64-encoded JSON with strict key validation

The Python worker entry point (texera\_run\_python\_worker.py) now receives its startup configuration as a Base64-encoded JSON string rather than raw positional arguments. This change ensures the configuration survives Windows command-line argv quoting issues and enforces strict validation: the worker now checks for an exact set of expected keys (including new fields for Iceberg REST catalog, S3 large binary storage, and R path) and rejects any missing, unexpected, or non-string values. This improves reliability and configurability for users running Python workers across different platforms and storage backends.

amber/src/main/python · high confidence

Rebranding to Apache Texera Amber

The project has been rebranded from 'Apache Amber' to 'Apache Texera Amber', reflected in the new package namespace \org.apache.texera.amber\ across the engine, configuration, and workflow core modules. This change includes the introduction of new storage implementations (\HDFSRecordStorage\, \VFSRecordStorage\, \EmptyRecordStorage\) and a \SequentialRecordStorage\ abstraction for handling record I/O, alongside utility classes like \ConfigParserUtil\ and \WorkflowRuntimeException\ to support the updated architecture.

amber/src/main/scala/org/apache/texera/amber/engine/common/storage, common/config/src/main/scala/org/apache/texera/amber, common/workflow-core/src/main/scala/org/apache/texera/amber/core · high confidence

Refactor Amber message types and introduce StateFrame loop envelope

The Amber engine's message-passing layer has been restructured to improve clarity and support loop execution. New message payload types have been introduced, including \DataPayload\ (with \StateFrame\ and \DataFrame\), \DirectControlMessagePayload\, \RecoveryPayload\, and \WorkflowFIFOMessagePayload\, alongside the \WorkflowMessage\ envelope. A key behavioral change is the introduction of the \StateFrame\ case class, which carries loop metadata (\loopCounter\ and \loopStartId\) alongside the actual state. This ensures that when JVM operators are nested inside Python loop bodies, the loop envelope is preserved and passed through unchanged, preventing collisions with user state and ensuring correct loop context propagation.

amber/src/main/scala/org/apache/texera/amber/engine/common/ambermessage · high confidence

Refactored log replay architecture with new logging and ordering components

The log replay subsystem in the Amber engine has been restructured to improve fault tolerance and replay reliability. A new \AsyncReplayLogWriter\ now handles log persistence in a separate thread, decoupling I/O from the main processing flow. The \ReplayLogger\ and \ReplayLogManager\ classes manage the recording of processing steps and workflow messages, while the \ReplayOrderEnforcer\ and its \ReplayOrderEnforcer\ implementation ensure that replayed operations occur in the correct sequence relative to the original execution. Additionally, a \ReplayLogGenerator\ was introduced to parse existing logs and extract processing steps and messages for replay.

amber/src/main/scala/org/apache/texera/amber/engine/architecture/logreplay · high confidence

Reintroduce Amber worker architecture with operator reconfiguration support

The Amber engine's worker architecture has been re-enabled, restoring the core components (WorkflowWorker, DataProcessor, DPThread) and wiring the ExecutionReconfigurationService back into the engine. This change brings back operator reconfiguration capabilities, allowing workflows to be modified and reconfigured at runtime. It also introduces scoped large binary storage and cleanup by execution ID, improves state ordering by version rather than timestamp, and carries loop StateFrame envelopes through JVM operators to ensure correct state handling within loop bodies.

amber/src/main/scala/org/apache/texera/amber/engine/architecture/worker · high confidence

Relocate Amber module to new directory structure

The Amber module has been relocated to a new directory structure, moving core Python architecture and storage model components (such as BufferedItemWriter and VirtualDocument abstractions) and Scala collaborative WebSocket event/request models into their new locations under the amber source tree.

(repo-wide) · high confidence

Standardize frontend development environment and tooling configuration

The frontend project now includes explicit configuration files to standardize the development environment and code quality checks. An \.editorconfig\ enforces consistent indentation (2 spaces) and line endings, while \.prettierrc.json\ and \.eslintrc.json\ define formatting and linting rules for TypeScript and HTML. A \.gitignore\ file excludes build artifacts and dependencies, and a \.nvmrc\ file pins the Node.js version to 24. Additionally, the project now bundles the Yarn 4.14.1 release binary to ensure consistent package management across environments.

frontend · high confidence

State materialization now includes loop bookkeeping columns

The State representation in the workflow core has been updated to include two new columns, loop\_counter and loop\_start\_id, alongside the existing content column. These fields allow the system to carry loop control information through JVM operators without modifying the user's state JSON, ensuring that loop context is preserved correctly when state is materialized and transported across the pipeline.

common/workflow-core/src/main/scala/org/apache/texera/amber/core/state · high confidence

Unified workflow compiler with configurable error handling and improved JSON serialization

The workflow compiler has been consolidated into a single shared component that supports two distinct modes: Lenient mode for editing-time validation, which accumulates per-operator errors and returns a null physical plan, and Strict mode for pre-execution checks, which fails fast on the first error. This change also introduces robust JSON serialization for LogicalLink, ensuring OperatorIdentity round-trips correctly whether the frontend sends plain strings or object shapes, and adds specific error handling for Python-based operators during code generation to capture and report configuration issues.

common/workflow-compiler · high confidence

Worker state transitions now use a logical version clock instead of timestamps

The StateManager in the Amber engine now tracks state changes using a monotonically increasing logical version rather than wall-clock timestamps. This change ensures that state reports from workers are ordered causally, allowing the controller to correctly reject stale or out-of-order updates without relying on synchronized clocks across distributed processes. The WorkerStateManager now exposes a state version alongside the current state, improving reliability in distributed execution scenarios.

amber/src/main/scala/org/apache/texera/amber/engine/common/statetransition · high confidence

Fixes

Fix resource leak in dataset repository deletion

The \deleteRepo\ method in \GitVersionControlLocalFileStorage\ now properly closes the \Files.walk\ stream using a try-with-resources block, preventing potential file descriptor leaks when deleting entire dataset repositories.

common/workflow-core/src/main/scala/org/apache/texera/amber/core/storage/util/dataset · high confidence

Test coverage

Added NonParallelTest annotation to serialize flaky shared-resource tests; Added comprehensive test coverage for the web module; Added end-to-end tests for Amber engine batch size propagation, data processing, pause/resume, and reconfiguration; Added fault-tolerance tests for checkpointing and log replay; Added integration test infrastructure and sample data for the workflow operator; Added integration test suite for loop, multi-region, and reconfiguration workflows; Added integration tests for AccessControlService bootstrap and initialization; Added integration tests for Iceberg storage components; Added smoke and unit tests for the local-dev shell script and TUI; Added smoke tests for the single-node deployment wrapper; Added test coverage for Amber worker promise handlers; Added test coverage for Arrow schema utilities, ObjectMapper warmup, and Python worker pool; Added test coverage for GlobalPortIdentity and PortIdentity serialization; Added test coverage for Loop Start and Loop End operator descriptions; Added test coverage for SiteSettings, SqlServer transactions, and UserWarehouse schema; Added test coverage for dashboard search and resource logic; Added test coverage for dataset and workflow resource endpoints; Added test coverage for storage utilities and LakeFS client; Added test coverage for the authentication resource module; Added test coverage for workflow resource endpoints; Added test data resources for workflow-core; Added test suite for VirtualDocument storage model; Added tests for AI assistant initialization logic; Added tests for ActorCommandHandler base class behavior; Added tests for FileResolver path resolution; Added tests for Git version control storage and readonly local file documents; Added tests for LargeBinary I/O and S3 storage integration; Added tests for LargeBinary model validation and behavior; Added tests for Python proxy client and server components; Added tests for Python storage components; Added tests for State JSON serialization and loop bookkeeping; Added tests for Time Series Plot operator descriptor; Added tests for UserActivityEventListener; Added tests for WebSocket protocol serialization and deserialization; Added tests for access control resource routing and privilege checks; Added tests for access-control-service resource and utility components; Added tests for computing unit access resolution and auth header fields; Added tests for console message error reporting and print replacement utilities; Added tests for the Amber control-plane architecture; Added tests for the File Lister operator and dataset path parsing; Added tests for the Python worker startup configuration parsing; Added tests for the Warehouse REST API resource; Added tests for user joining reason endpoints; Added unit and integration tests for Iceberg storage components; Added unit and integration tests for InputManager, OutputManager, and state materialization; Added unit tests for AddressInfo and WorkerExecution components; Added unit tests for Amber scheduling configuration classes; Added unit tests for AmberClient and ClientActor; Added unit tests for AmberMessage envelopes and DataFrame behavior; Added unit tests for Arrow and CSVOld file scan operators; Added unit tests for AsyncRPCClient and AsyncRPCServer; Added unit tests for CSV scan source operators; Added unit tests for ConfigParserUtil; Added unit tests for Distinct and Difference operators; Added unit tests for ErrorUtils; Added unit tests for File Scan operator descriptors and utilities; Added unit tests for Hash Join and Cartesian Product operators; Added unit tests for Hugging Face codegen operators; Added unit tests for Hugging Face operator descriptors; Added unit tests for JSONL scan source operator; Added unit tests for JWT authentication and role-based authorization; Added unit tests for Java and R UDF operator descriptors; Added unit tests for Keyword Search operator and its case-sensitive analyzer; Added unit tests for ML Scorer operator descriptors and metrics; Added unit tests for Map operator description and execution; Added unit tests for NetworkOutputBuffer and non-range partitioners; Added unit tests for PyBuilder validation and rendering logic; Added unit tests for Python AsyncRPC components; Added unit tests for Python UDF Source Operator descriptor; Added unit tests for Python UDF example operators and UI parameter support; Added unit tests for Python UDF operator descriptors and UI parameter injection; Added unit tests for Python Virtual Environment resource and websocket endpoints; Added unit tests for Python core utilities and worker wiring; Added unit tests for Python operator descriptor validation; Added unit tests for Random K Sampling operator; Added unit tests for Regex and Substring Search operators; Added unit tests for Reservoir Sampling operator; Added unit tests for SQL source operator schema introspection and execution; Added unit tests for Sklearn training operator descriptors; Added unit tests for Sort Partitions and Symmetric Difference operators; Added unit tests for StateManager and WorkerStateManager; Added unit tests for StoppableQueueBlockingRunnable lifecycle and error handling; Added unit tests for Texera configuration objects; Added unit tests for TimedBuffer; Added unit tests for Type Casting and Unnest String operators; Added unit tests for Union operator descriptor and execution; Added unit tests for UserQuotaResource; Added unit tests for admin execution and user resources; Added unit tests for control-plane worker handlers; Added unit tests for coordinator promise handlers; Added unit tests for core data models and operators; Added unit tests for core executor traits and factory; Added unit tests for customized queue utilities and LinkedBlockingMultiQueue; Added unit tests for dataset statistics utilities; Added unit tests for dataset storage and version control utilities; Added unit tests for external API source operators; Added unit tests for operator metadata generation and validation; Added unit tests for resource scheduling policies; Added unit tests for schema model and Arrow conversion utilities; Added unit tests for sklearn advanced operator descriptors; Added unit tests for storage model and runnables; Added unit tests for storage model classes; Added unit tests for the Aggregate operator and related utilities; Added unit tests for the Amber messaging layer components; Added unit tests for the Amber record-storage subsystem; Added unit tests for the Dictionary Matcher operator; Added unit tests for the Hub dashboard resource and entity models; Added unit tests for the If operator descriptor and execution logic; Added unit tests for the Interval Join operator; Added unit tests for the Limit operator descriptor and executor; Added unit tests for the Projection operator; Added unit tests for the Python worker bridge and internal queue; Added unit tests for the Text Input operator's schema inference and row scanning logic; Added unit tests for the URL Fetcher operator and its utilities; Added unit tests for the cluster membership listener; Added unit tests for the common auth module; Added unit tests for the filter operator components; Added unit tests for the log-replay subsystem; Added unit tests for the tuple and schema core library; Added unit tests for visualization operators and utilities; Added unit tests for worker manager components; Added unit tests for workflow operator descriptors and metadata; Added unit tests for workflow result storage models and Iceberg integration; Added unit tests for workflow-core core types and utilities; Added unit tests for workflow-core utility classes; Added validation to ensure Helm template values are defined in values.yaml.

Dependencies

Dependency updates across build manifests

This change updates dependency versions in the build manifests for the access-control-service, agent-service, amber, and common modules, as well as the root build configuration. Specific updates include bumping the agent-service's @ai-sdk/openai to 4.0.36 and ai to 7.0.58, upgrading amber's pyarrow to 23.0.1 and numpy to 2.1.0, and updating the root build's Jackson version to 2.18.8 and Netty to 4.2.15.Final. The access-control-service now uses Dropwizard 4.0.7, while the amber module continues with Dropwizard 1.3.23 and Pekko 1.7.0. These changes ensure the services use the specified library versions for their respective functionalities.

(dependencies) · high confidence

Housekeeping

Updated license and notice files for access-control-service

The access-control-service now includes regenerated LICENSE-binary and NOTICE-binary files. These files consolidate the Apache License 2.0 terms and provide detailed attribution notices for third-party dependencies, including Eclipse Jetty and Jersey, ensuring compliance with open-source licensing requirements.

access-control-service · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

Baseline

  • First survey — no prior run to compare against. CAI 49.

Lenses

  • Code Health 79
  • Architecture 94
  • Maturity 61
  • Readiness 41
  • Security 59
  • Accessibility 45

Changes since last survey

  • 300 commits — 252 feature/other, 48 fixes

By area

  • frontend/src — 121 commits
  • amber/src — 72 commits
  • common/workflow-operator — 31 commits
  • file-service/src — 10 commits
  • common/workflow-core — 7 commits
  • agent-service/src — 6 commits
  • .github/workflows — 5 commits
  • bin/k8s — 5 commits
  • (root) — 3 commits
  • access-control-service/src — 3 commits
  • bin/local-dev — 3 commits
  • bin/single-node — 3 commits
  • common/config — 3 commits
  • computing-unit-managing-service/src — 3 commits
  • amber/LICENSE-binary-python — 2 commits
  • common/pybuilder — 2 commits
  • common/workflow-compiler — 2 commits
  • frontend/vitest.config.ts — 2 commits
  • notebook-migration-service/src — 2 commits
  • pyright-language-service/package.json — 2 commits

Notable commits

  • fix: chore: add 1.3.0-incubating and 1.4.0-incubating-SNAPSHOT to bug-report options (#8309)
  • fix: fix(WorkflowExecutionService): shutdown console writer thread on unsubscribe (#7914)
  • fix: fix(agent-service): delete links targeting removed input ports when shrinking input ports (#7349)
  • fix: fix(agent-service): handle falsy model throws (#7638)
  • fix: fix(agent-service): reject maxSteps 0 in settings (#7821)
  • fix: fix(amber): declare cloudpickle in LICENSE-binary-python (#8293)
  • fix: fix(amber): enable checkpoint serialization (#7996)
  • fix: fix(amber): guard resends by destination (#8002)
  • fix: fix(amber): refresh typing-extensions in LICENSE-binary-python (#8375)
  • fix: fix(amber): reject duplicate worker initialization (#8083)
  • fix: fix(amber): report a malformed PVE websocket handshake (#7852)
  • fix: fix(amber, operator): keep error information on three failure paths (#7804)
  • fix: fix(ci, test, frontend): stop the frontend matrix manufacturing false reds (#7623)
  • fix: fix(computing-unit): repair the owner-avatar accessor in the spec (#7633)
  • fix: fix(deploy): enable CORS on RustFS so browsers can fetch presigned URLs (#8562)
  • fix: fix(deps, ci): bump sbt/setup-sbt to v1.5.7 to restore CI (#7710)
  • fix: fix(deps, frontend): deduplicate Angular toolchain under @angular-builders/custom-webpack (#8127)
  • fix: fix(deps, frontend): update dependency @angular/core to v21.2.20 (#8494)
  • fix: fix(frontend): give unit tests timeout headroom for loaded macOS runners (#7717)
  • fix: fix(frontend): keep the warehouse run-button label inside the button, and name the picker in its tooltip (#8589)
  • …and 280 more

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

apache/texera was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 20 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit 28594efe6cf7628be81bd9ff8939199417773ecd — the exact code this score is about.
  • Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-b51f968c9b10.