Skip to content
CAI
Software that uses CAICheck a score

locationtech/geomesa

66.1

Adequate · 27 September 2026

130k

lines of production code

Scala

with Java

3

measurements over time

CAI band scale
CAI trend line
CAI lens gauges

What this system is

This system is a geospatial data management framework that provides a unified API for storing, querying, and transforming spatial features across diverse backends including Accumulo, Cassandra, HBase, Kafka, Redis, and file systems. It supports high-performance spatial indexing, distributed processing via Spark and MapReduce, and flexible data ingestion through a comprehensive converter framework for formats like Avro, Parquet, and JSON. The system also integrates with standard geospatial tools via GeoServer plugins and offers CLI utilities for schema management, data export, and cluster configuration.

How it got here

2014–2016 — Multi-store expansion and core refactoring

83 changes.

This period focused on extending GeoMesa's backend support by introducing initial implementations for Cassandra, HBase, and Spark SQL integration, alongside a new in-memory CQEngine index. Concurrently, the codebase underwent a major architectural overhaul, removing legacy Accumulo components and refactoring core serialization, indexing, and data store abstractions to support these new targets.

2017–2018 — Arrow integration and converter v2 API

55 changes.

This period focused on introducing Apache Arrow support for high-performance geospatial data storage and processing, alongside a comprehensive rewrite of the data conversion framework to version 2. It also expanded ecosystem integrations by adding PySpark bindings, HBase Spark support, and a new Lambda data store combining Kafka and Accumulo.

2019–2024 — multi-store expansion and infrastructure modernization

60 changes.

This period focused on significantly expanding GeoMesa's storage capabilities by introducing native support for Redis, Confluent Kafka, and Parquet, while modernizing the Accumulo backend with Iceberg metadata and advanced RPC coprocessors. The work also standardized the conversion and metrics subsystems through Micrometer integration and a new validation framework, alongside comprehensive CLI tooling and distribution packaging improvements for all supported data stores.

2025–2026 — Trino integration and test expansion

15 changes.

This period focused on introducing a new Trino datastore and plugin, enabling secure, read-only querying of Iceberg tables with row-level visibility and structural JSON support. Concurrently, the team significantly expanded test coverage across Accumulo, HBase, Kryo, and FileSystem modules, while also adding new converter-based storage capabilities for Hadoop batch processing.

Features

Accumulo DataStore audit logging implementation

The Accumulo datastore now includes a complete audit trail for query operations. This change introduces new components to capture, store, and retrieve query events: an audit writer that asynchronously persists query metadata (including user, filter, timing, and execution plans) to a dedicated Accumulo table, an audit reader to retrieve these logs, and the necessary event transformation logic to map between internal audit structures and Accumulo mutations. A service provider file is also added to register the data store factory, ensuring the audit functionality is integrated into the standard GeoMesa data store lifecycle.

geomesa-accumulo/geomesa-accumulo-datastore/src/main · high confidence

Accumulo DataStore implementation and query planning

This change introduces the core Accumulo DataStore implementation, including the factory for creating connections, the query planning engine for executing scans, and the writer components for handling conditional writes and column family mappings. It also adds a SchemaCopier utility for bulk-moving data between Accumulo clusters. For users, this provides the foundational support for storing and querying geospatial data in Apache Accumulo.

repository · high confidence

Add Arrow export command to CLI tools

Users can now export features from a GeoMesa Arrow data store using the command-line interface. This change introduces the \ArrowExportCommand\ class, which integrates the Arrow data store into the existing export framework, allowing users to specify the data store URL and feature type name to perform exports.

geomesa-arrow/geomesa-arrow-tools/src/main/scala/org/locationtech/geomesa/arrow/tools/export · high confidence

Add Avro Schema Registry converter factory

Introduces the AvroSchemaRegistryConverterFactory, enabling the Geomesa convert system to register and instantiate converters that utilize the Confluent Schema Registry for Avro schema resolution. This change adds the necessary service provider configuration and factory implementation to support this specific conversion path.

geomesa-convert/geomesa-convert-avro-schema-registry/src/main · high confidence

Add CLI configuration files and dependency resolution script

The distribution now includes a Maven assembly component that packages environment configuration files (hadoop-env.sh, log4j.properties) into the conf directory. Additionally, a new dependencies.sh script is provided to automatically resolve and manage Hadoop client dependencies for the command-line tools, including logic to handle specific version constraints and exclusions. Example configuration files for PostGIS and Shapefile data sources are also added to assist with initial setup.

geomesa-gt · high confidence

Add CLI tools for the Redis data store

This change introduces the command-line interface (CLI) tools for interacting with the GeoMesa Redis data store. It adds a new \RedisRunner\ entry point and a suite of commands for schema management (create, describe, update, remove, manage partitions), data ingestion and export, feature deletion, and statistics analysis (count, bounds, histogram, top-K, analyze). The tools support connecting to Redis instances, including those in cluster mode, and allow users to specify authorizations for data access.

geomesa-redis/geomesa-redis-tools · high confidence

Add Leaflet map viewer resource to exporters

The feature-exporters module now includes a Leaflet-based HTML map viewer (index.html) as a bundled resource. This allows users to visualize exported geospatial data directly in a browser with interactive features like layer toggling (points vs. heatmap) and base map selection (Greyscale/OSM).

geomesa-features/geomesa-feature-exporters · high confidence

Add Simple Feature-to-Feature converter

Users can now transform one Simple Feature type into another using the new FeatureToFeatureConverter. This component, registered via the META-INF service loader, allows direct conversion of existing SimpleFeature objects into the internal array format required by the conversion pipeline, enabling workflows where features need to be re-typed or re-structured without parsing from raw input streams.

geomesa-convert/geomesa-convert-simplefeature/src/main · high confidence

Add command-line tools for Trino data store operations

Users can now manage and inspect GeoMesa data stored in Trino via the \geomesa-trino\ CLI. This release introduces commands to export features (\export\), describe schema attributes (\describe-schema\), list available feature types (\get-type-names\), and retrieve SimpleFeatureType definitions (\get-sft-config\). These tools connect to a Trino coordinator using host, port, schema, and authentication parameters, enabling programmatic interaction with Trino-backed geospatial data.

geomesa-trino/geomesa-trino-tools · high confidence

Add delimited text converter factory for the new converter API

The text converter module now registers a \DelimitedTextConverterFactory\ via the Java Service Provider Interface, enabling the ingestion of delimited text files (CSV/TSV) using the updated v2 converter API. This factory supports automatic schema inference from input data, handles both standard delimited formats and GeoMesa-specific exported formats (identified by an 'id,' header), and allows configuration of parsing options such as error modes and character sets.

geomesa-convert/geomesa-convert-text/src/main · high confidence

Added SerializationOptions helper class for feature serialization configuration

A new SerializationOptions class has been added to the interop package to provide static factory methods for building serialization option sets. This class exposes convenient Scala-compatible methods (withUserData and withoutId) that delegate to the SerializationOption builder, simplifying the creation of specific serialization configurations for users of the Geomesa feature interop layer.

geomesa-features/geomesa-feature-common/src/main/java/org/locationtech/geomesa/features/interop · high confidence

Added Trino command-line tool distribution and dependency management

This change introduces the packaging and configuration for the GeoMesa Trino command-line tools. It adds a Maven assembly descriptor (\component.xml\) to bundle the Trino plugin zip and necessary configuration files (like \hadoop-env.sh\ and \log4j.properties\) into the distribution. Additionally, it provides a shell script (\dependencies.sh\) that dynamically resolves and lists the required runtime dependencies, including Hadoop, AWS SDKs (v1 or v2 depending on Hadoop version), and other libraries, ensuring the tools can locate the correct JARs from the environment.

geomesa-trino · high confidence

Added interop sample jobs for MapReduce feature I/O

New sample MapReduce jobs have been added to the interop package to demonstrate reading and writing features via GeoMesa Accumulo. FeatureCountJob shows how to count features and output geometry to HDFS using map/reduce counters, while FeatureWriterJob demonstrates copying features from one GeoMesa store to another by transforming them into a new SimpleFeatureType. These samples rely on the new GeoMesaAccumuloInputFormat wrapper, which delegates to the existing Accumulo input format to simplify configuration for CQL-based queries.

geomesa-accumulo/geomesa-accumulo-jobs/src/main/java · high confidence

Arrow-specific status and schema inspection commands

The Arrow command-line tools now include dedicated commands for inspecting Arrow data stores: \describe-schema\ shows feature type attributes and, with the new \--show-dictionaries\ flag, displays dictionary values; \get-sft-config\ retrieves the SimpleFeatureType definition; and \get-type-names\ lists available feature types. These commands extend the existing status tooling to support the Arrow data store backend.

geomesa-arrow/geomesa-arrow-tools/src/main/scala/org/locationtech/geomesa/arrow/tools/status · high confidence

Attribute-level join index support for Accumulo

The Accumulo indices module now supports creating join indices at the attribute level. This change introduces the \AccumuloFeatureIndexFactory\ and registers it via the \GeoMesaFeatureIndexFactory\ service provider, enabling the system to recognize and instantiate \JoinIndex\ variants (versions 2 through 8) when an attribute is configured with the \IndexCoverage.JOIN\ option. The implementation includes the \AttributeJoinIndex\ mixin to handle query planning and filtering for these join indices, allowing users to optimize queries that rely on attribute-based joins within the Accumulo data store.

geomesa-accumulo/geomesa-accumulo-indices · high confidence

Automated generation of the Geomesa BOM dependency list

A new shell script has been added to automatically generate the \geomesa-bom/pom.xml\ file by scanning installed Maven artifacts. This script builds the project to collect artifact metadata, then constructs the dependency management section with appropriate classifiers, types, and exclusions for bundled runtime JARs, ensuring the Bill of Materials remains synchronized with the actual built artifacts without manual maintenance.

geomesa-utils-parent · high confidence

Cassandra CLI command implementations

Adds Cassandra-specific command-line tool implementations for schema and data management, including create, remove, and update schema operations, feature ingestion, feature deletion, and schema description/explanation commands. The ingest command explicitly restricts distributed file ingestion, throwing an error if remote files are detected.

geomesa-cassandra/geomesa-cassandra-tools/src/main/scala/org/locationtech/geomesa/cassandra/tools/commands · high confidence

Composite converter supports conditional routing of input lines to multiple delegates

The composite converter now allows a single input stream to be processed by multiple underlying converters based on configurable predicates. Each line of input is evaluated against a sequence of conditions; if a condition matches, the line is passed to the corresponding delegate converter. This enables users to route different types of records within a single file to different processing pipelines without needing separate conversion steps.

geomesa-convert/geomesa-convert-common/src/main/scala/org/locationtech/geomesa/convert2/composite · high confidence

Enable Spark integration for HBase data stores

Users can now read from and write to HBase data stores using Spark RDDs. This change introduces the HBaseSpatialRDDProvider implementation and registers it via the META-INF/services file, allowing Spark applications to leverage GeoMesa's spatial capabilities directly with HBase-backed data.

geomesa-hbase/geomesa-hbase-spark · high confidence

HBase Datastore implementation and configuration

The HBase datastore module now provides a complete implementation for connecting to and managing data in HBase. This includes the core data store and index adapter classes, a metadata layer backed by HBase tables, and support for various aggregation scans (Arrow, Bin, Density, Stats, Version) via coprocessors. Configuration is handled through a dedicated parameter set, allowing users to specify connection details, Zookeeper ensemble servers, coprocessor settings, and security options. A service provider file registers the HBase data store factory, making it available for use within the GeoMesa ecosystem.

geomesa-hbase/geomesa-hbase-datastore · high confidence

Initial Cassandra DataStore implementation

This release introduces a new Cassandra-backed data store for GeoMesa, enabling users to store and query geospatial features using Apache Cassandra. The implementation includes the core data store and index adapter classes, along with specific column mappers for Id, Z2, Z3, and Attribute indexes to handle schema mapping and query planning. It also provides a metadata layer backed by Cassandra tables and registers the new store via the GeoTools service provider mechanism.

geomesa-cassandra/geomesa-cassandra-datastore/src/main · high confidence

Initial implementation of Arrow attribute readers and writers

This change introduces the core \ArrowAttributeReader\ and \ArrowAttributeWriter\ components in the \geomesa-arrow-gt\ module, enabling the conversion between GeoTools SimpleFeatures and Apache Arrow vectors. The new code provides factory methods to create specific readers and writers for various data types (such as geometry, strings, dates, and UUIDs) and supports dictionary encoding for optimized storage and retrieval of repeated attribute values.

geomesa-arrow/geomesa-arrow-gt/src/main/scala/org/locationtech/geomesa/arrow/vector · high confidence

Initial release of Cassandra command-line tools

This change introduces the core infrastructure for the GeoMesa Cassandra CLI, including the \geomesa-cassandra\ runner and essential connection parameter handling. It enables users to manage Cassandra-specific geospatial data through commands for schema creation and removal, feature ingestion and deletion, data export, and playback, along with utilities for describing schemas and explaining queries.

geomesa-cassandra/geomesa-cassandra-tools/src/main/scala/org/locationtech/geomesa/cassandra/tools · high confidence

Initial release of GeoMesa Arrow integration components

This change introduces the initial structure for the GeoMesa Arrow module, adding a distribution component configuration that packages Guava and filtered log4j properties, a README describing the new Arrow JTS Geometry extension for encoding/decoding JTS geometries, and a shell script defining dependency handling for the command-line tools.

geomesa-arrow · high confidence

Initial release of Geomesa Cassandra CLI tools and configuration

This change introduces the command-line tools for Geomesa Cassandra, including the necessary configuration files to run them. It adds a Maven assembly descriptor for the distribution, shell scripts to manage the Cassandra environment and dependencies (such as the DataStax driver and Netty libraries), a logback configuration for structured logging, and example schema and converter definitions for CSV, XML, JSON, and Avro data formats.

geomesa-cassandra · high confidence

Introduce Arrow-backed SimpleFeature implementation

Added ArrowSimpleFeature, a new class that implements the GeoTools SimpleFeature interface using Apache Arrow vectors. This enables lazy evaluation of attributes, allowing filters to examine only relevant vectors for optimized reads, while ensuring features remain tied to their underlying Arrow data structures.

geomesa-arrow/geomesa-arrow-gt/src/main/scala/org/locationtech/geomesa/arrow/features · high confidence

Introduce GeoCQEngine in-memory datastore

Adds a new GeoTools-compatible in-memory datastore backed by CQEngine, enabling users to store and query spatial features in memory with optional geometry indexing. The change registers the \GeoCQEngineDataStoreFactory\ via the standard service loader mechanism and provides configuration parameters to enable geometry indexing and specify a namespace.

geomesa-memory/geomesa-cqengine-datastore · high confidence

Introduce GeoTools-based CLI commands for PostGIS data stores

This change adds a new set of command-line tools (accessible via the \geomesa-gt\ runner) that allow users to manage schemas and data in GeoTools-backed data stores, specifically targeting PostGIS. The new commands include \create-schema\, \delete-features\, \describe-schema\, \get-sft-config\, \get-type-names\, \remove-schema\, \update-schema\, \partition-upgrade\, \export\, \playback\, and \ingest\. The \update-schema\ command specifically supports updating user data and partitioning configurations (such as interval and cron settings) for partitioned PostGIS stores, while \partition-upgrade\ handles upgrading partitioning functions. A new resource file \gt-libjars.list\ is added to manage required libraries (commons-dbcp, postgresql, spatial4j, systems-common) for these tools.

geomesa-gt/geomesa-gt-tools · high confidence

Introduce HBase RPC coprocessor framework and protocol buffers

This change adds the core infrastructure for executing GeoMesa queries via HBase coprocessors, including the \GeoMesaProto\ protocol buffer definitions, the \GeoMesaCoprocessor\ client-side execution logic, and the \GeoMesaHBaseCallBack\ for handling RPC responses. It also introduces a suite of HBase-specific filters (\CqlTransformFilter\, \Z2HBaseFilter\, \Z3HBaseFilter\, \S2HBaseFilter\, \S3HBaseFilter\) that push down spatial and CQL predicates to the HBase layer, along with a deprecated \JSimpleFeatureFilter\ for compatibility mode 2.3. This enables more efficient query processing by filtering data at the source rather than in the client.

geomesa-hbase/geomesa-hbase-rpc · high confidence

Introduce Hadoop file-system abstraction and distributed copy utilities

This change adds a new Hadoop utilities module that provides a file-system delegate layer for Hadoop-based storage operations. It introduces HadoopDelegate to abstract file handling (including support for TAR, ZIP, and JAR archives) and HadoopUtils to manage configuration resource loading and Kerberos ticket renewal. Additionally, it adds DistributedCopyOptions to handle Hadoop DistCp configuration across different Hadoop versions, enabling more robust bulk data copying capabilities within the Geomesa ecosystem.

geomesa-utils-parent/geomesa-hadoop-utils · high confidence

Introduce Iceberg-backed FileSystemRDDProvider for Spark

A new FileSystemRDDProvider implementation is added to the Spark module, enabling spatial data processing via Spark RDDs backed by Iceberg metadata. This provider reads feature data from Parquet files located via Iceberg's location references and writes features using the existing FileSystemDataStore, effectively bridging the Spark runtime with the Iceberg-based file system data store.

geomesa-fs/geomesa-fs-spark/src/main/scala · high confidence

Introduce Kafka DataStore with Zookeeper-less metadata and Kafka Streams integration

The Kafka DataStore now supports a Zookeeper-less architecture by storing metadata in a dedicated Kafka topic (default 'geomesa-catalog') instead of Zookeeper, while retaining backward compatibility for existing Zookeeper-based deployments via an automatic migration path. This change introduces a new \GeoMesaStreamsBuilder\ for Kafka Streams integration, allowing users to create KStreams and KTables of geospatial features directly. The datastore also adds configurable topic management options, including the ability to clear topics on startup or truncate them upon schema deletion, and standardizes configuration parameters for consumers, producers, and serialization types.

geomesa-kafka/geomesa-kafka-datastore · high confidence

Introduce Lambda Data Store for hybrid Kafka and Accumulo storage

Adds a new Lambda Data Store that combines Kafka for transient, real-time event storage with Accumulo for long-term persistence. This component registers a new GeoTools DataStore factory ("Kafka/Accumulo Lambda (GeoMesa)") and provides configuration parameters for Kafka brokers, Zookeeper offsets, and feature expiry. It includes implementations for managing distributed offsets via Zookeeper, handling transient feature writes and deletes through Kafka topics, and aggregating statistics across both the transient and persistent layers.

geomesa-lambda/geomesa-lambda-datastore · high confidence

Introduce Micrometer-based metrics with Prometheus and CloudWatch support

Geotools users can now configure and expose Geomesa metrics using the Micrometer library, replacing the previous implementation. This change adds support for Prometheus (with HTTP and push-gateway export) and Amazon CloudWatch, allowing metrics to be customized via configuration files (cloudwatch.conf, prometheus.conf, instrumentations.conf). It also includes a custom MetricsDataSource for Apache Commons DBCP2 to ensure connection pool gauges are correctly registered with Micrometer, and provides utility classes for managing gauge and tag registration to prevent conflicts.

geomesa-metrics/geomesa-metrics-micrometer · high confidence

Introduce PySpark bindings for GeoMesa spatial operations

This release adds the \geomesa\_pyspark\ Python package, enabling users to run GeoMesa spatial functions directly within PySpark. The new module provides a \configure\ helper to set up the Spark environment with the necessary JARs and Python dependencies, and exposes a comprehensive set of spatial SQL functions (such as \st\_boundary\, \st\_makePoint\, and \st\_geomFromWKT\) that wrap the underlying Scala UDFs. It also introduces a \GeoMesaSpark\ API for creating spatial RDDs from GeoJSON, dictionaries, or tuples, and implements custom PySpark User Defined Types (UDTs) to seamlessly serialize and deserialize Shapely geometry objects (Points, LineStrings, Polygons, etc.) for use in Spark DataFrames.

_geomesa-spark/geomesa\pyspark · high confidence

Introduce Trino datastore with connection pooling and query translation

This change adds the core Trino datastore implementation for GeoMesa, enabling users to query Trino tables as GeoTools feature sources. It introduces a JDBC connection pool keyed by normalized authorization sets to support secure, multi-tenant access, and implements a feature reader that maps Trino types (including structural JSON, nested objects, and decimals) to GeoTools features. The datastore also includes a filter translator that pushes down spatial, temporal, and attribute predicates to Trino SQL, along with support for client-side filtering and JSON-path query projections.

_geomesa-trino-datastore\2.12 · high confidence

Introduce dedicated CLI tools for GeoMesa FileSystem Data Store (FSDS)

This change adds a new set of command-line tools for managing the GeoMesa FileSystem Data Store, accessible via the \geomesa-fs\ runner. The new tools include commands for creating schemas (\create-schema\), ingesting data (\ingest\), exporting and playback (\export\, \playback\), compacting partitions (\compact\), and managing metadata such as registering or unregistering data files (\manage-metadata\). The implementation introduces a new dependency resolution script (\dependencies.sh\) that dynamically detects Hadoop and AWS SDK versions from the environment and includes specific libraries (like Zstandard and Guava) to ensure compatibility, alongside a list of required libjars for distributed execution.

geomesa-fs/geomesa-fs-tools · high confidence

Introduce read-only Trino/Iceberg GeoMesa datastore with structural JSON support

This change adds a new read-only GeoTools DataStore implementation for Trino/Iceberg tables, enabling GeoMesa to query Trino catalogs with Z2/XZ2 partition pruning and row-level visibility. The datastore introduces support for Trino structural types (arrays, maps, and rows) by mapping them to JSON strings with optional Avro schema derivation, allowing nested data to be preserved and queried. It also implements a connection pool keyed by authorization sets to efficiently manage per-request authentication credentials forwarded to Trino, and includes schema discovery from Iceberg table metadata to automatically map column types and visibility columns.

geomesa-trino/geomesa-trino-datastore · high confidence

Introduces Arrow memory management and configuration utilities

Adds a new Scala package object that provides core infrastructure for Arrow integration, including a global Arrow memory allocator with shutdown hooks to prevent leaks, methods to monitor allocated memory, and a configurable batch size property for controlling scan performance.

geomesa-arrow/geomesa-arrow-gt/src/main/scala/org/locationtech/geomesa/arrow · high confidence

Introduces core index API abstractions and merged data store view support

This change establishes the foundational index API in the \geomesa-index-api\ module by introducing key interfaces and classes such as \TableSplitter\ for configuring table splits, \AtomicWriteTransaction\ to enforce atomic feature writes, and \IndexKeySpace\ to manage index key conversions. It also adds the \MergedViewConfigLoader\ interface and registers \MergedDataStoreViewFactory\ and \RoutedDataStoreViewFactory\ via service providers, enabling the system to load and route configurations for merged data store views.

geomesa-index-api · high confidence

Introduction of optimized filter evaluation and new CQL functions

The filter module now uses a new FastFilterFactory to create optimized filter implementations, including FastComparisonOperator, FastTemporalOperator, and FastDWithin, which improve query performance by avoiding repeated literal evaluation. Additionally, new CQL filter functions such as BucketHash, Convert2Viewer, CurrentDate, DateToLong, FastProperty, MurmurHash, ProxyId, XZ2, and Z2 are registered and available for use in queries.

geomesa-filter/src/main · high confidence

JDBC converter factory registration for Geomesa Convert v2

The JDBC converter is now explicitly registered as a factory for the Geomesa Convert v2 service provider interface. This change introduces the JdbcConverterFactory class and its corresponding configuration decoder, enabling the system to discover and instantiate JDBC-based feature converters via the standard Java service loader mechanism.

geomesa-convert/geomesa-convert-jdbc/src/main · high confidence

Kryo serialization now supports projecting features to different schemas

The Kryo feature serializer now includes dedicated serializers and deserializers that allow features to be projected to a different schema during serialization or deserialization. This enables users to serialize data according to a target schema (ProjectingKryoFeatureSerializer) or deserialize stored data into a different, projected schema (ProjectingKryoFeatureDeserializer), facilitating schema evolution and flexible data access without requiring manual attribute mapping.

geomesa-features/geomesa-feature-kryo/src/main/scala/org/locationtech/geomesa/features/kryo · high confidence

New Accumulo iterators for spatial indexing, aggregation, and feature transformation

This release introduces a comprehensive suite of new Accumulo iterators in the \geomesa-accumulo-iterators\ module to enhance query performance and data handling. New iterators include \AgeOffIterator\ and \DtgAgeOffIterator\ for automatic data expiration based on time, and \ArrowIterator\, \BinAggregatingIterator\, \DensityIterator\, and \StatsIterator\ for efficient spatial aggregation and statistical analysis. The \FilterTransformIterator\ enables server-side filtering and feature transformation, while \AttributeKeyValueIterator\ supports attribute-based indexing. Spatial indexing is supported by \S2Iterator\, \S3Iterator\, \Z2Iterator\, and \Z3Iterator\, along with a \RowFilterIterator\ base class. Additional utilities include \KryoVisibilityRowEncoder\ for visibility-aware serialization and \ProjectVersionIterator\ for version detection.

geomesa-accumulo/geomesa-accumulo-iterators/src/main/scala/org/locationtech/geomesa/accumulo/iterators · high confidence

New Accumulo job utility and argument parsing infrastructure

This change introduces new Scala source files in the Accumulo jobs module to support job configuration and execution. AccumuloJobUtils provides helper methods to automatically discover and set required Hadoop library JARs by searching environment variables (GEOMESA\_ACCUMULO\_HOME, ACCUMULO\_HOME) and the classpath, ensuring jobs have access to necessary dependencies. GeoMesaArgs introduces a structured argument parsing framework using JCommander, defining traits for input and output data store parameters (such as user, password, instance ID, Zookeepers, and table name) and feature/CQL arguments, enabling consistent and reusable command-line interface handling for Accumulo-based GeoMesa jobs.

geomesa-accumulo/geomesa-accumulo-jobs/src/main/scala/org/locationtech/geomesa/accumulo/jobs · high confidence

New Accumulo-specific Spark RDD provider

Added the AccumuloSpatialRDDProvider class, which enables reading spatial data from Accumulo into Spark RDDs and writing RDDs back to Accumulo tables. This implementation handles query plan execution, manages connection lifecycle, and supports feature writes using the Accumulo data store.

geomesa-accumulo/geomesa-accumulo-spark/src/main/scala · high confidence

New Apache Arrow file-based data store

Users can now read and write geospatial data stored in Apache Arrow format files. This change introduces the ArrowDataStore and its factory, enabling GeoTools to process .arrow files via the 'arrow.url' parameter. The implementation supports streaming or caching-based reads and append-only writes, with caching disabled by default to allow writing.

geomesa-arrow/geomesa-arrow-datastore/src/main · high confidence

New Arrow I/O components for batch writing, sorting, and dictionary management

The \geomesa-arrow-gt\ module introduces a new set of classes in the \org.locationtech.geomesa.arrow.io\ package to handle Arrow file operations. \BatchWriter\ and its internal \BatchSortingIterator\ enable global sorting of Arrow record batches by a specified field, supporting both ascending and descending orders. \ConcatenatedFileWriter\ provides a mechanism to merge separate Arrow files into a single stream, including handling empty inputs. \DictionaryBuildingWriter\ allows for the dynamic construction of Arrow dictionaries as features are added, with configurable maximum sizes. \SimpleFeatureArrowFileWriter\ and \SimpleFeatureArrowFileReader\ manage the low-level serialization and deserialization of Simple Features to and from Arrow streaming format, supporting both caching and streaming read modes. Supporting utilities in \package.scala\ and \records/\ handle metadata for sort fields, legacy IPC format options, and record batch loading/unloading.

geomesa-arrow/geomesa-arrow-gt/src/main/scala/org/locationtech/geomesa/arrow/io · high confidence

New Avro data file I/O components for SimpleFeature storage

Added three new Scala classes in the Avro IO package to handle binary Avro files for SimpleFeature storage: AvroDataFile defines the metadata schema (including version 3 support for Bytes types) and validation logic; AvroDataFileReader provides a stream-based reader that validates file version and schema compatibility before reading; and AvroDataFileWriter enables writing SimpleFeatures with configurable compression and serialization options. These components work together to provide self-describing, long-term binary storage for spatial features with embedded schema information.

geomesa-features/geomesa-feature-avro/src/main/scala/org/locationtech/geomesa/features/avro/io · high confidence

New CLI tools distribution with automated dependency management and sample data scripts

The geomesa-tools distribution now includes a new shell entry point (geomesa-tools) that automatically checks for and downloads missing Maven dependencies on first run, configurable via GEOMESA\_CHECK\_DEPENDENCIES. It also provides a download-data.sh script to fetch sample datasets (GDELT, GeoLife, OSM-GPX, T-Drive, GeoNames) and ships with pre-configured Simple Feature Types and converters for various data sources (ADSBx, GDELT, etc.) in the conf directory.

geomesa-tools · high confidence

New CQEngine-backed in-memory spatial index with multiple indexing strategies

The geomesa-cqengine module now provides a new in-memory spatial index implementation backed by CQEngine, replacing the previous approach. This change introduces support for three distinct spatial indexing strategies: Bucket, STRtree, and QuadTree, allowing users to select the most appropriate index type for their data distribution and query patterns. The implementation includes new attribute handling classes (SimpleFeatureAttribute, SimpleFeatureFidAttribute) to bridge GeoTools SimpleFeatures with CQEngine's attribute model, and a comprehensive query visitor (CQEngineQueryVisitor) that translates GeoTools filters into CQEngine queries. The index supports various attribute types including geometries, strings, numbers, dates, and UUIDs with appropriate indexing strategies (RadixTree, Navigable, Hash, Unique). A fallback GeoToolsFilterQuery is provided for queries that cannot be optimized by CQEngine indexes. The module also includes parameter classes for configuring index behavior (BucketIndexParam, STRtreeIndexParam) and a factory (GeoIndexFactory) for creating the appropriate index type based on configuration.

geomesa-memory/geomesa-cqengine · high confidence

New CQEngine-backed in-memory spatial indexing for GeoTools

This change introduces a new in-memory spatial index implementation for GeoTools features, leveraging the CQEngine library to provide high-performance spatial queries. The diff adds an abstract base index class (\AbstractGeoIndex\) that wraps a spatial index structure (such as STRtree or Quadtree) and a custom \Intersects\ query type, enabling efficient spatial lookups for features stored in memory. This allows users to perform spatial operations on in-memory feature collections with improved performance compared to previous approaches.

_geomesa-cqengine\2.12 · high confidence

New Cassandra export and playback CLI commands

The Cassandra tools now include two new command-line interfaces: an export command to export features from a GeoMesa data store, and a playback command to replay features based on their date. These commands are implemented as \CassandraExportCommand\ and \CassandraPlaybackCommand\, extending the base \ExportCommand\ and \PlaybackCommand\ respectively, and integrate with the existing Cassandra connection and catalog parameters.

geomesa-cassandra/geomesa-cassandra-tools/src/main/scala/org/locationtech/geomesa/cassandra/tools/export · high confidence

New Confluent Kafka Data Store for Avro Schema Registry integration

This change introduces a new Confluent Kafka data store implementation that enables reading from and writing to Kafka topics using Confluent Avro serialization and the Confluent Schema Registry. Users can now connect to Confluent Kafka clusters by providing a schema registry URL, allowing GeoMesa to automatically resolve feature schemas from the registry. The implementation includes a new DataStore factory, metadata handling via the Schema Registry, and a custom message serializer that supports null message keys by generating random UUIDs. It also allows overriding schemas via a configuration parameter and explicitly prevents schema creation, update, or deletion within the data store, requiring schema management to be handled externally in the registry.

geomesa-kafka/geomesa-kafka-confluent · high confidence

New Converter-based SpatialRDDProvider for Spark

Users can now ingest data into Spark using the existing GeoMesa converter framework via a new \ConverterSpatialRDDProvider\. This provider allows reading from input files by specifying converter definitions and Simple Feature Types either as inline configuration strings or by looking them up by name from the classpath. It supports standard Spark SQL features such as attribute projections and ECQL filtering during the conversion process, enabling flexible data transformation directly within the Spark RDD layer.

geomesa-spark/geomesa-spark-converter · high confidence

New ConverterConfigProvider interface for loading configuration

A new ConverterConfigProvider interface has been added to the convert-common module, defining a loadConfigs method that returns a map of configuration objects. This interface provides a standardized way to load feature type and SFT configurations from providers, enabling external configuration sources to be integrated into the conversion process.

geomesa-convert/geomesa-convert-common/src/main/java/org/locationtech/geomesa/convert · high confidence

New HBase Spark jobs for Kerberos-authenticated feature I/O

The geomesa-hbase-jobs module now includes new Scala components that enable Spark jobs to read and write SimpleFeatures to HBase with Kerberos authentication. GeoMesaHBaseInputFormat provides a MapReduce input format that wraps HBase scans to deliver features to Spark, while HBaseIndexFileMapper converts features into HBase mutations for bulk loading. HBaseJobUtils supplies query-plan helpers to validate single-table scans, and the Security module handles Kerberos login via keytab/principal configuration, ensuring secure execution in distributed environments.

geomesa-hbase/geomesa-hbase-jobs · high confidence

New HBase bootstrap and configuration scripts for AWS and local deployments

The HBase module now includes a new AWS bootstrap script (bootstrap-geomesa-hbase-aws.sh) that automates the installation of GeoMesa HBase on AWS EMR, including setting up environment variables, deploying the distributed runtime JAR to HBase's root directory (supporting both S3 and HDFS storage), and configuring coprocessor auto-registration. Additionally, new configuration files (hbase-env.sh, dependencies.sh) and a Maven assembly component definition are added to manage HBase classpaths, dependency resolution (including HBase, Hadoop, Zookeeper, and Netty versions), and environment setup for CLI tools.

geomesa-hbase · high confidence

New JSON converter and composite JSON support

The JSON converter module now registers factory classes for both standard and composite JSON processing via service provider files, enabling the system to automatically discover and use the new \JsonConverter\ and \JsonCompositeConverter\ implementations. This change introduces support for composite JSON configurations (type \composite-json\), allowing users to define multiple delegate converters with predicates to process different parts of a JSON structure, alongside the existing single-converter JSON ingestion capabilities.

geomesa-convert/geomesa-convert-json/src/main · high confidence

New JTS-based Arrow geometry vector implementations

The \geomesa-arrow-jts\ module now includes a complete set of new classes to handle JTS geometry types (Point, LineString, Polygon, and their multi-variants) as Apache Arrow vectors. This adds support for both single-precision (float) and double-precision storage, along with a WKB fallback vector, enabling more efficient memory usage and native Arrow integration for spatial data.

geomesa-arrow/geomesa-arrow-jts/src/main · high confidence

New Java API for loading SimpleFeatureConverters

A new \SimpleFeatureConverterLoader\ class has been added to the \geomesa-convert-common\ module, providing a static \load\ method that accepts a \SimpleFeatureType\ and a Typesafe Config object to instantiate a \SimpleFeatureConverter\. This change introduces a dedicated Java-friendly API for converter initialization, bridging the Scala-based converter implementation with Java callers.

geomesa-convert/geomesa-convert-common/src/main/java/org/locationtech/geomesa/convert2 · high confidence

New Kafka CLI tools and Zookeeper migration support

The geomesa-kafka CLI now includes a full suite of commands for managing Kafka-backed GeoMesa stores, including schema creation, removal, and updates, as well as ingest, export, and a new playback command for time-based replay of features. A new 'migrate-zookeeper-metadata' command allows users to migrate feature schemas from Zookeeper to Kafka, supporting a Zookeeper-less architecture. The tools also support configuring Kafka producers and consumers via properties files, integrating with Confluent Schema Registry, and handling authorizations.

geomesa-kafka/geomesa-kafka-tools · high confidence

New Lambda data store CLI commands

The \geomesa-lambda\ tool now includes a suite of commands for managing and analyzing data stored in the Lambda data store, which combines Accumulo and Kafka. Users can create, remove, and delete features from schemas, export data, and playback historical records. Additionally, statistical analysis commands are available to calculate bounds, counts, histograms, and top-K values for feature types. These commands require connection parameters for Accumulo (instance, zookeepers, user, password, keytab, catalog, auths) and Kafka (brokers, optional zookeepers, partitions).

geomesa-lambda/geomesa-lambda-tools · high confidence

New MapReduce I/O formats and serialization for GeoMesa jobs

The geomesa-jobs module now includes new MapReduce input and output formats to simplify data ingestion and export. Users can ingest data using ConverterInputFormat (supporting converters, filters, and re-typing) and AvroFileInputFormat, while GeoMesaOutputFormat enables writing features directly to GeoMesa data stores with optional index targeting. A new SimpleFeatureSerialization class handles Kryo-based serialization of features for Hadoop jobs, and JobUtils provides helper methods to configure libjars for distributed execution.

geomesa-jobs · high confidence

New MapReduce input and output formats for Accumulo RFiles

Added new MapReduce components to handle Accumulo RFiles directly: GeoMesaAccumuloInputFormat groups tablet splits to optimize mapper counts for queries, while GeoMesaAccumuloFileOutputFormat writes RFiles to HDFS instead of using batch writers. These changes enable more efficient bulk data ingestion and export workflows by bypassing the overhead of traditional batch writing and reducing the number of mappers spawned during scans.

geomesa-accumulo/geomesa-accumulo-jobs/src/main/scala/org/locationtech/geomesa/accumulo/jobs/mapreduce · high confidence

New MapReduce jobs for attribute indexing, schema migration, and index back-filling

Three new distributed jobs are now available in the Accumulo jobs module to manage data and indexing operations. The AttributeIndexJob allows users to create or rebuild attribute and join indexes on specific schema attributes via a MapReduce process. The SchemaCopyJob enables copying data between data stores while rewriting features to the latest serialization format, which is useful for migrating older data to leverage new performance improvements. The WriteIndexJob allows users to back-fill data into specific indexes without re-writing unchanged indices, providing a more efficient way to add new indexing capabilities to existing datasets.

geomesa-accumulo/geomesa-accumulo-jobs/src/main/scala/org/locationtech/geomesa/accumulo/jobs/index · high confidence

New Redis-backed GeoMesa data store

GeoMesa now supports storing and querying geospatial features in Redis. This new data store uses Redis sorted sets for indexing and supports both standalone Redis instances and Redis clusters. Key capabilities include automatic feature expiration (TTL), distributed locking for concurrent access, and metadata storage within Redis. Users can configure connection parameters such as the Redis URL, cluster mode, connection pool size, and socket timeout via standard GeoMesa data store parameters.

geomesa-redis/geomesa-redis-datastore · high confidence

New Shapefile converter for Geomesa ingest

This change introduces a new Shapefile converter module that enables users to ingest .shp files directly into Geomesa. The implementation includes the core converter logic, a factory for configuration and schema inference, and specific transformer functions (shp, shpFeatureId) to access shapefile attributes during conversion. It also handles automatic charset inference from .cpg files and reprojection to EPSG:4326 when the source shapefile uses a different coordinate reference system.

geomesa-convert/geomesa-convert-shp/src/main · high confidence

New Spark SQL integration with Apache Sedona and expanded geospatial functions

This release introduces a new Spark SQL module that integrates with Apache Sedona, enabling users to leverage Sedona's spatial capabilities within GeoMesa. The changes include a new DataSource implementation for reading and writing GeoMesa data via Spark SQL, along with a suite of DataFrame DSL functions for geospatial operations such as distance calculations, geometry transformations, and spatial relations. Additionally, the module provides PySpark-compatible UDFs for common geometric operations and integrates Sedona's optimization rules and UDFs into the Spark SQL execution plan, allowing for seamless interoperability between GeoMesa and Sedona's spatial engine.

geomesa-spark/geomesa-spark-sql · high confidence

New build and release automation scripts

The build system now includes a suite of new shell scripts to streamline development and release workflows. A new \do-release.sh\ script automates the release process, handling version detection, GitHub tagging, artifact signing with GPG, publishing to Maven Central, and creating GitHub releases. Supporting utilities include \bloop-export.sh\ for initializing Bloop IDE integration, \change-scala-version.sh\ for updating Scala versions across POMs, \calculate-cqs.sh\ for analyzing compile-time dependencies, \update-copyright-year.sh\ for annual license updates, \update-maven-plugins.sh\ for upgrading Maven plugins, and \update-maven-toolchains.sh\ for generating Maven toolchain configurations based on detected JDKs.

build/scripts · high confidence

New composite XML converter for nested data structures

The XML converter module now includes a composite converter (XmlCompositeConverter) and its factory (XmlCompositeConverterFactory) that allows processing nested XML structures by delegating to multiple sub-converters. This enables users to map complex, hierarchical XML documents to simple features by applying different conversion rules to different parts of the XML tree, controlled via configuration predicates.

geomesa-convert/geomesa-convert-xml/src/main/scala · high confidence

New converter configuration loading and error handling modes

The converter framework now introduces a pluggable configuration loading system via \ConverterConfigLoader\ and \ConverterConfigProvider\, allowing configs to be loaded from the classpath or external URLs. It also adds new \ErrorMode\, \ParseMode\, and \LineMode\ enumerations to manage conversion behavior, with \ErrorMode\ supporting 'raise-errors', 'log-errors', and 'return-errors' options, and \ParseMode\ supporting 'incremental' and 'batch' processing. Additionally, an \EnrichmentCache\ system is introduced with support for simple in-memory and resource-based (CSV) caches to optimize data enrichment during conversion.

geomesa-convert/geomesa-convert-common/src/main/scala/org/locationtech/geomesa/convert · high confidence

New converter transform functions for type casting, collections, and geometry parsing

The converter API now includes a comprehensive set of new transform functions to handle data conversion and manipulation. Users can cast values to integers, longs, floats, doubles, booleans, and strings using both function calls (e.g., \toInt\, \stringToBoolean\) and type-cast syntax (e.g., \::int\, \::boolean\). New collection functions allow parsing delimited strings into lists and maps (\parseList\, \parseMap\), accessing list items, and transforming list elements. Geometry parsing has been expanded to support points, lines, polygons, and collections from WKT, WKB, or coordinate arrays, including support for Z and M coordinates. Additional functions cover date parsing with various ISO formats, math operations, base64 encoding/decoding, and ID generation including MurmurHash and UUID variants.

geomesa-convert/geomesa-convert-common/src/main/scala/org/locationtech/geomesa/convert2/transforms · high confidence

New converter-based read-only storage and Hadoop job support

This change introduces a new 'converter' storage type that allows reading data from external file paths by applying a SimpleFeatureConverter to transform raw files into features, enabling read-only access to data without writing to the underlying Iceberg table. It also adds Hadoop MapReduce input and output formats (ParquetPartition and ParquetSimpleFeature) to support batch processing jobs like compaction, including configuration for partition schemes, file sizes, and compression.

geomesa-fs/geomesa-fs-storage · high confidence

New dependency installation scripts for Confluent and Parquet support

The geomesa-kafka tools now include dedicated shell scripts to install required dependencies for Confluent Schema Registry and Parquet formats. Users can run \install-confluent-support.sh\ to fetch Confluent libraries (such as the schema registry client and Avro serializer) and \install-parquet-support.sh\ for Parquet dependencies, simplifying the setup process for these integrations.

geomesa-kafka · high confidence

New feature-exporters module with multi-format export support

A new \geomesa-feature-exporters\ module has been introduced, providing a unified \FeatureExporter\ interface and concrete implementations for exporting spatial features into various formats. Users can now export data to Arrow (with support for dictionary encoding, batch sizes, axis-order flipping, and flattened struct output), Avro (with configurable compression), binary (with label/geometry/datetime/track field configuration), delimited text (CSV/TSV with optional headers and IDs), GeoJSON, GeoParquet, GML (GML2/GML3 with asynchronous encoding), Shapefile, and interactive Leaflet maps (including heatmap generation). The module also includes a \NullExporter\ for discarding output and utility streams for byte-counting and optional GZIP compression.

geomesa-features/geomesa-feature-exporters/src/main/scala/org/locationtech/geomesa/features/exporters · high confidence

New geomesa-arrow CLI runner and base command infrastructure

The geomesa-arrow command-line interface is now bootstrapped with a dedicated runner (geomesa-arrow) that registers a suite of tools including schema description, export, ingestion, type name retrieval, SFT configuration, and various statistics commands (bounds, count, top-K, histogram). This change introduces the base ArrowDataStoreCommand trait and URL parameter handling that these commands rely on to connect to Arrow resources.

geomesa-arrow/geomesa-arrow-tools/src/main/scala/org/locationtech/geomesa/arrow/tools · high confidence

New geometry serialization infrastructure and GeoJSON export support

The serialization module now includes a complete rewrite of geometry encoding with new \WkbSerialization\ and \TwkbSerialization\ traits that handle X, Y, Z, and M dimensions, alongside a new \DimensionalBounds\ utility for extracting coordinate bounds. A new \GeoJsonSerializer\ has been added to enable exporting SimpleFeatures to GeoJSON, including support for JSON-type attributes, while \HintKeySerialization\ provides mappings for GeoTools hints.

geomesa-features/geomesa-feature-common/src/main/scala/org/locationtech/geomesa/features/serialization · high confidence

New pluggable security providers and visibility filter functions

The geomesa-security module now includes a pluggable AuthorizationsProvider interface and a SpringAuditProvider to manage user authorizations and audit information, alongside new GeoTools filter functions (isVisible, visibility, getVisibilities) that enforce feature-level visibility rules using the Accumulo access evaluator.

geomesa-security · high confidence

New setup-namespace.sh script for Accumulo namespace configuration

A new shell script, setup-namespace.sh, has been added to the geomesa-accumulo-tools/bin directory to automate the setup of Accumulo namespaces for GeoMesa. This tool installs the GeoMesa distributed runtime JAR into HDFS and configures the corresponding Accumulo namespace, including setting up the necessary classpath contexts and permissions. It supports both password-based and Kerberos token authentication, allowing users to easily configure their Accumulo instance to use GeoMesa's distributed runtime capabilities.

geomesa-accumulo/geomesa-accumulo-tools/bin · high confidence

New streaming and caching readers for Arrow input streams

Added three new reader implementations in the \geomesa-arrow-gt\ module to handle Arrow data from \InputStream\ sources: \StreamingSimpleFeatureArrowFileReader\ for single-pass, low-memory feature iteration; \CachingSimpleFeatureArrowFileReader\ for scenarios requiring random access or multiple passes by caching batches in memory; and \MultiStreamSimpleFeatureArrowFileReader\ to support re-reading the input stream for multiple queries. These classes provide the concrete I/O layer for reading Arrow files, complementing existing file-based readers.

geomesa-arrow/geomesa-arrow-gt/src/main/scala/org/locationtech/geomesa/arrow/io/reader · high confidence

New utility classes and service registrations for binary encoding, caching, and concurrency

This change introduces a suite of new utility components within geomesa-utils to support binary feature encoding, thread-local caching, and controlled concurrency. It adds the BinaryOutputEncoder and related callback interfaces to serialize spatial features into compact binary formats (16-byte and 24-byte records) with optional sorting and label support. It also introduces ThreadLocalCache and SoftThreadLocal for managing per-thread cached values with expiration and garbage-collection safety, alongside CachedThreadPool and ExitingExecutor to manage thread lifecycles and concurrency limits. Additionally, it registers new GeoTools service providers for JSON path filtering, property access, and type conversion factories, and provides Java interoperability wrappers for SimpleFeatureTypes and WKTUtils.

geomesa-utils-parent/geomesa-utils · high confidence

Parquet file ingestion support added to converters

Users can now ingest data from Parquet files using the Geomesa Convert framework. This change introduces the ParquetConverter and ParquetConverterFactory, which handle reading Parquet files (including GeoParquet) and inferring schema types. It also registers Parquet-specific transformer functions (such as parquetPoint, parquetLineString, etc.) to correctly parse geometry data stored in Parquet format.

geomesa-convert/geomesa-convert-parquet/src/main · high confidence

Partitioned PostGIS data store with row-level security and metrics

The partitioned PostGIS data store is now registered as a discoverable DataStore factory and includes support for per-row visibility filtering via a PostgreSQL session variable that applies user authorizations to every connection. The store also exposes PostgreSQL database metrics through Micrometer and provides a converter to map SQL arrays to Java lists for correct handling of list-type attributes.

geomesa-gt/geomesa-gt-partitioning · high confidence

Redis-backed enrichment cache for GeoMesa converters

Users can now use Redis as a backing store for enrichment caches within the GeoMesa converter framework. This change introduces a new Redis-specific implementation that allows converters to fetch and cache attribute data from a Redis instance during processing. The feature is configured via a service provider file that registers the Redis factory, enabling users to specify Redis connection details and caching behaviors (such as expiration and local caching) in their converter configurations.

geomesa-convert/geomesa-convert-redis-cache/src/main · high confidence

Register Accumulo Spark provider via service loader

A new META-INF/services file is added to register the Accumulo-specific SpatialRDDProvider implementation, enabling the Spark module to automatically discover and use the Accumulo backend for spatial RDD operations.

geomesa-accumulo/geomesa-accumulo-spark/src/main/resources · high confidence

Row-level visibility security for the spatial Iceberg connector

The spatial Iceberg connector now enforces row-level visibility filters for all SQL consumers (direct SQL, JDBC, BI tools). This is achieved by injecting a per-row filter via the Trino \ConnectorAccessControl\ interface, which resolves user authorizations using pluggable \AuthorizationResolver\ implementations (file-based or extra-credential-based) and evaluates them against table visibility columns using a new \is\_visible\ SQL UDF. The system ensures fail-closed security by hiding rows that do not match the user's authorizations, while maintaining an allow-all baseline for other access control aspects.

_geomesa-trino-plugin\2.12, geomesa-trino/geomesa-trino-plugin · high confidence

Support for reading and writing arbitrary GeoTools data stores via Spark

Users can now use GeoMesa Spark to read from and write to any data store supported by GeoTools (such as PostGIS) by setting the 'geotools' parameter to 'true' in the data store configuration. This change introduces the GeoToolsSpatialRDDProvider, which acts as a generic bridge for GeoTools-backed stores, allowing Spark RDDs to be loaded from and saved to these databases alongside existing GeoMesa-specific providers.

geomesa-gt/geomesa-gt-spark · high confidence

Vector process implementations and service registration

The geomesa-process-vector module now provides the concrete implementations for the GeoMesa vector processing capabilities, including analytic operations (Density, MinMax, Point2Point, Sampling, Stats, TrackLabel, Unique), query operations (Join, ProximitySearch, Query, RouteSearch, KNearestNeighborSearch), and transform operations (ArrowConversion, BinConversion, DateOffset, HashAttribute, TubeSelect). These classes are registered via the standard Java ServiceLoader mechanism in META-INF/services, making them discoverable by the GeoMesaProcessFactory for execution within the GeoServer Web Processing Service (WPS) and other GeoTools-based environments.

geomesa-process/geomesa-process-vector · high confidence

Removals

Removal of core.scala constants file

The file geomesa-core/src/main/scala/geomesa/core/core.scala has been deleted. This removal eliminates a package object that previously defined default configuration constants and property names, such as DEFAULT\_GEOMETRY\_PROPERTY\_NAME, DEFAULT\_FEATURE\_TYPE, and various iterator-related settings. Users relying on these specific default values or the package object structure will need to adjust their configuration or code accordingly.

geomesa-core/src/main/scala/geomesa/core · high confidence

Removal of legacy Accumulo data store implementation

The legacy Accumulo data store implementation in geomesa-core has been removed. This includes the deletion of the core data store classes (AccumuloDataStore, AccumuloFeatureStore, AccumuloFeatureReader, AccumuloFeatureWriter, AccumuloFeatureSource), the MapReduce ingestion components (MapReduceAccumuloDataStore, FeatureIngestMapper), and the configuration utilities (AccGeoConfiguration). This change eliminates the old data access layer and its associated MapReduce-based ingestion path from the core module.

geomesa-core/src/main/scala/geomesa/core/data · high confidence

Removal of legacy GeoMesa Accumulo plugin components

The geomesa-plugin module has removed its legacy GeoServer integration files, including the Accumulo data store and coverage store edit panels, the WMS coverage format factory, and the associated Spring configuration and service provider entries. This cleanup eliminates the deprecated UI panels and factory registrations that previously allowed GeoServer to connect to and serve raster data from Accumulo tables using the old plugin architecture.

geomesa-plugin · high confidence

Removal of legacy core utility and index constant classes

The \Constants\ class (Java) and \BoundingBoxUtil\ and \Ingest\ objects (Scala) in \geomesa-core\ have been removed. This eliminates the Java-side wrapper for Scala index constants, the utility for generating Accumulo query ranges from bounding boxes, and the MapReduce-based ingestion framework that relied on these components.

geomesa-core/src/main/java, geomesa-core/src/main/scala/geomesa/core/util · high confidence

Removal of legacy geomesa-utils source files and license

This change removes a significant set of legacy Scala and Java source files from the geomesa-utils module, including GeoHash iteration utilities (BoundingBoxGeoHashIterator, RadialGeoHashIterator, etc.), distance calculation helpers (VincentyModel, GeomDistance), and GeoTools integration code (ShapefileIngest, JodaConverterFactory). The module's LICENSE.txt file is also deleted. These deletions indicate a cleanup or migration away from the old geomesa.utils package structure and dependencies (such as Joda Time) within this utility library.

geomesa-utils · high confidence

Removal of legacy iterator stack

The \geomesa.core.iterators\ package has been removed, deleting the legacy iterator implementations including \AggregatingCombiner\, \AggregatingKeyIterator\, \ColumnQualifierAggregatingIterator\, \ConsistencyCheckingIterator\, \DeDuplicatingIterator\, \RowOnlyIterator\, \SimpleFeatureFilteringIterator\, \SpatioTemporalIntersectingIterator\, \SurfaceAggregatingIterator\, \TimestampRangeIterator\, and \TimestampSetIterator\. This change eliminates the previous in-memory aggregation and filtering logic used during scans, replacing it with a new distributed aggregation approach for capabilities such as heatmaps.

geomesa-core/src/main/scala/geomesa/core/iterators · high confidence

Removal of the geomesa-dist sub-project and its documentation

The geomesa-dist module has been removed from the codebase. This change deletes the distribution assembly configuration, the license file, the DocBook documentation source, and the Scala REPL utility script. As a result, the local distribution packaging and bundled documentation previously provided by this module are no longer available in this location.

geomesa-dist · high confidence

Behavioural changes

Accumulo CLI tools restructured with new runner and connection handling

The Accumulo command-line interface has been reorganized to improve connection management and library resolution. A new AccumuloRunner registers the full suite of Accumulo-specific commands (such as bulk-copy, compact, and age-off configuration) and automatically resolves instance, ZooKeeper, and authentication details from an accumulo-client.properties file. The tools now use a dedicated library list to ensure required Accumulo and Hadoop JARs are available on the classpath, and shared parameter traits standardize connection options across all commands.

geomesa-accumulo/geomesa-accumulo-tools/src/main · high confidence

Added placeholder to ensure Scala profile activation

A .gitkeep file was added to the main Scala source directory to ensure the Scala profile is activated during the build process.

geomesa-features/geomesa-feature-all/src/main · high confidence

Arrow JTS vectors now support configurable axis order

The \geomesa-arrow-jts\ module introduces new abstract vector implementations for JTS geometry types (LineString, MultiLineString, MultiPoint, MultiPolygon, Polygon) that respect an axis-order configuration. By checking \isFlipAxisOrder()\, these vectors now write and read X/Y coordinates in either the standard order or a flipped order, allowing users to control coordinate axis orientation when serializing and deserializing geometries to/from Apache Arrow formats.

geomesa-arrow-jts · high confidence

Arrow-specific statistics commands now support optional feature type names

The Arrow stats CLI tools (bounds, count, histogram, and top-K) have been updated to no longer require a feature type name to be explicitly provided. These commands now inherit the ProvidedTypeNameParam, allowing users to run statistics without specifying the typename argument, while still enforcing exact calculation modes for accuracy.

geomesa-arrow/geomesa-arrow-tools/src/main/scala/org/locationtech/geomesa/arrow/tools/stats · high confidence

Avro converter now supports schema inference for arbitrary Avro files

The Avro converter in the convert module has been updated to automatically infer the SimpleFeatureType from the schema of arbitrary Avro input files, rather than requiring a pre-defined schema. This change introduces an \AvroConverterFactory\ and \AvroFunctionFactory\ that handle type inference for standard Avro types (strings, numbers, booleans, lists, maps) and derive geometry fields where possible. Users can now ingest Avro files without explicitly providing a schema configuration, as the converter will analyze the file's embedded schema to determine the feature structure.

geomesa-convert/geomesa-convert-avro/src/main · high confidence

Avro serialization now supports native list and map types

The Avro feature serializer has been updated to support native Avro list and map encodings (serialization version 6), replacing the previous opaque binary representation. This change introduces a new \FieldNameEncoder\ to handle schema field name encoding and updates the serialization version constants, allowing users to store complex collection attributes more efficiently and compatibly within Avro schemas.

geomesa-features/geomesa-feature-avro/src/main/scala/org/locationtech/geomesa/features/avro · high confidence

Configure Maven repository filtering and Scala version defaults

Maven builds now enforce explicit group IDs for third-party repositories via a new remote repository filter configuration, restricting access to specific groups such as io.confluent and various GeoTools-related artifacts. Additionally, default Scala binary and patch versions (2.12 and 21, respectively) are now set globally for the build.

.mvn · high confidence

Default converter configuration for validation and parsing

A new default configuration file (base-converter-defaults.conf) establishes baseline settings for converters, specifically setting the default validator to 'index', the parse mode to 'incremental', and the error mode to 'log-errors'.

geomesa-convert/geomesa-convert-common/src/main/resources · high confidence

Enables FileSystemRDDProvider via Spark service discovery

The Spark module now registers the FileSystemRDDProvider through the standard Java ServiceLoader mechanism. By adding the provider class to the META-INF/services configuration, the system can automatically discover and use the file-system-based spatial RDD implementation without requiring explicit manual wiring.

geomesa-fs/geomesa-fs-spark/src/main/resources · high confidence

Explicit Accumulo job dependency list added

A new resource file, accumulo-libjars.list, has been added to define the specific JAR dependencies required for Accumulo-based jobs. This list explicitly enumerates libraries such as accumulo-core, geomesa, and various supporting utilities, ensuring that the correct classpath is provided when running these tools.

geomesa-accumulo/geomesa-accumulo-jobs/src/main/resources · high confidence

File System Data Store now uses Iceberg for metadata by default

The GeoMesa File System Data Store has switched its default metadata implementation from the legacy ConverterCatalog to Apache Iceberg. This change improves performance for statistics queries (such as counts and min/max calculations) by leveraging Iceberg's file-level bounds, and introduces new configuration parameters like 'fs.catalog.type' to explicitly select the catalog implementation (Iceberg or Converter). Users can still opt into the legacy behavior by setting the catalog type to 'converter', but the default path now relies on Iceberg for metadata management.

geomesa-fs/geomesa-fs-datastore/src/main · high confidence

Fixed-width converter factory registration for Geomesa Convert 2

The fixed-width converter is now explicitly registered as a factory for the Geomesa Convert 2 API. A new service provider file declares the FixedWidthConverterFactory, and the factory class itself implements the AbstractConverterFactory interface, enabling the system to discover and instantiate fixed-width converters using the updated configuration model (BasicConfig, BasicOptions) and field decoding logic.

geomesa-convert/geomesa-convert-fixedwidth/src/main · high confidence

Format-aware converter inference for data conversion

The conversion system now uses a new TypeAwareInference mechanism to select the appropriate data converter based on the input file format. By mapping specific formats (such as Avro, JSON, CSV, TSV, Parquet, Shapefile, and XML) to their corresponding factory classes, the system can prioritize the correct converter when inferring the schema from a data sample, ensuring more accurate handling of diverse input types.

geomesa-convert/geomesa-convert-all/src/main · high confidence

HBase CLI tools now support distributed execution with remote filtering controls

The HBase command-line tools have been updated to support distributed ingestion and export jobs. This change introduces a --no-remote-filters flag to commands like ingest, export, and stats, allowing users to disable HBase remote filtering and coprocessors when needed. Additionally, distributed commands now automatically resolve the Zookeeper quorum from the HBase configuration if not explicitly provided, and they correctly bundle the necessary HBase library JARs for MapReduce jobs.

geomesa-hbase/geomesa-hbase-tools · high confidence

Introduce v2 converter API with new core abstractions

The converter-common module now provides a new v2 converter API, introducing the \SimpleFeatureConverter\ trait for stream-based processing, the \AbstractConverterFactory\ for configuration-driven instantiation, and the \ParsingConverter\ trait for separating parsing from feature creation. This change also adds the \ErrorHandlingIterator\ to manage parsing failures via configurable error modes (log, return, or raise) and defines core configuration traits (\ConverterConfig\, \Field\, \ConverterOptions\) to standardize how converters are configured and validated.

geomesa-convert/geomesa-convert-common/src/main/scala/org/locationtech/geomesa/convert2 · high confidence

Lambda distribution now bundles Accumulo, Kafka, and Hadoop CLI configuration and dependency scripts

The geomesa-lambda distribution package has been updated to include specific configuration files and dependency resolution scripts for its command-line tools. Users will now find bundled shell scripts (setup-namespace.sh, accumulo-env.sh, kafka-env.sh, hadoop-env.sh, and log4j.properties) in the bin and conf directories, alongside a new dependencies.sh script that manages classpath resolution for Accumulo, Hadoop, Zookeeper, and Kafka libraries. This ensures the lambda tools correctly locate and version their underlying infrastructure dependencies, including specific handling for Accumulo 2.1+ telemetry libraries and Hadoop version compatibility.

geomesa-lambda · high confidence

Local-only ingestion for Arrow data

The new ArrowIngestCommand restricts file ingestion to local paths only. If any specified file is remote, the command now fails with an error instead of attempting to process it, ensuring that only locally accessible files are ingested into the Arrow data store.

geomesa-arrow/geomesa-arrow-tools/src/main/scala/org/locationtech/geomesa/arrow/tools/ingest · high confidence

New Accumulo CLI environment and dependency configuration scripts

The CLI tools now include new shell scripts (accumulo-env.sh and dependencies.sh) to manage the Accumulo runtime environment and dependency classpaths. These scripts automatically detect and configure Accumulo, Hadoop, and Zookeeper versions, handling specific version requirements such as Hadoop 3.2.4 compatibility, Accumulo 2.1+ telemetry libraries (Micrometer, OpenTelemetry), and Zookeeper Jute inclusion. This ensures the CLI tools correctly resolve dependencies for the target Accumulo version without manual classpath configuration.

geomesa-accumulo/geomesa-accumulo-tools/conf-filtered · high confidence

New Avro user-data serialization with union types and legacy V4 support

The Avro serialization module now includes dedicated classes for encoding and decoding user data. The new AvroUserDataSerialization object implements a union-type encoding for standard primitives (strings, numbers, booleans, bytes) and explicitly drops the USE\_PROVIDED\_FID hint to reduce noise. A separate AvroUserDataSerializationV4 object provides backward-compatible serialization for older data formats, supporting additional types like Date, UUID, Geometry, and Lists. CollectionSerialization handles encoding of lists and maps as opaque byte buffers, while SimpleFeatureDatumWriter integrates these components to write SimpleFeatures with the new user-data strategy.

geomesa-features/geomesa-feature-avro/src/main/scala/org/locationtech/geomesa/features/avro/serialization · high confidence

New Kryo deserialization implementation with optimized bit-set storage

The Kryo feature serialization module now includes a new \ActiveDeserialization\ implementation that fully deserializes SimpleFeatures before returning them, supporting both mutable and immutable feature types. This change introduces a custom \IntBitSet\ implementation to reduce memory footprint for tracking null attributes, and adds \NonMutatingInput\ to prevent buffer mutation during string reading. The deserializer handles both V2 and V3 serialization formats, ensuring backward compatibility while improving performance and memory efficiency for deserialized features.

geomesa-features/geomesa-feature-kryo/src/main/scala/org/locationtech/geomesa/features/kryo/impl · high confidence

New Kryo serialization components for geometries, index values, and user data

The Kryo serialization module now includes dedicated serializers for JTS geometries, attribute join index values, and SimpleFeature user data. The new GeometrySerializer handles geometry encoding using the TWKB format, while IndexValueSerializer provides a mechanism to serialize specific attribute subsets for index lookups. Additionally, KryoUserDataSerialization introduces a custom encoding scheme for user data maps, supporting complex types like UUIDs, lists, and geometries, and ensuring backward compatibility with older GeoTools hint keys.

geomesa-features/geomesa-feature-kryo/src/main/scala/org/locationtech/geomesa/features/kryo/serialization · high confidence

New Kryo-based JSON serialization for GeoMesa features

The \geomesa-feature-kryo\ module now includes a new \KryoJsonSerialization\ implementation that serializes JSON attributes into a compact, BSON-like binary format using Jackson for parsing. This change replaces previous JSON handling with a more efficient serialization strategy that supports top-level arrays, handles null values explicitly, and includes error handling for corrupt data during deserialization, improving performance and robustness for features containing JSON fields.

geomesa-features/geomesa-feature-kryo/src/main/scala/org/locationtech/geomesa/features/kryo/json · high confidence

New converter validator framework with Micrometer metrics

The converter module now uses a new validation API where validators are created via factories and report success/failure counts using Micrometer metrics. New built-in validators include 'cql' (filtering by CQL expression), 'has-dtg' and 'has-geo' (checking for non-null date/geometry), 'id' (ensuring feature IDs are present), 'index' (validating geometry bounds and date ranges for Z2/Z3 indices), and 'none' (no validation). The default validation behavior remains 'index', and existing 'z-index' validator is deprecated in favor of 'index'.

geomesa-convert/geomesa-convert-common/src/main/scala/org/locationtech/geomesa/convert2/validators · high confidence

Redis distribution packaging and tool configuration

The geomesa-redis distribution now includes a Maven assembly component that packages Guava into the lib directory and filters the log4j.properties configuration file from the common-env module. Additionally, a new dependencies.sh script is provided for the CLI tools, defining empty functions to manage classpath dependencies and exclusions.

geomesa-redis · high confidence

Refactored Kafka consumer infrastructure with version compatibility and batch processing

The Kafka consumer implementation in geomesa-kafka-utils has been refactored to introduce a new threaded consumer architecture and support for batch processing. A new BaseThreadedConsumer class manages consumer threads, while ThreadedConsumer provides single-record processing with configurable offset commit intervals to reduce commit frequency. A new BatchConsumer class enables batched message processing with at-least-once guarantees, allowing consumers to commit, continue, or pause based on processing results. Additionally, new version-compatibility wrappers (KafkaConsumerVersions, KafkaAdminVersions, RecordVersions) use reflection to support multiple Kafka client versions (0.9 through 2.1), handling differences in method signatures for polling, offset management, and message headers.

geomesa-kafka/geomesa-kafka-utils · high confidence

Refactored SimpleFeature implementation with immutable and mutable base classes

The core feature implementation in geomesa-feature-common has been restructured to provide distinct base classes for immutable and mutable SimpleFeatures. AbstractImmutableSimpleFeature now throws UnsupportedOperationException on all set operations, while AbstractMutableSimpleFeature handles attribute setting with automatic type conversion via FastConverter. A new FastSettableFeature trait allows bypassing type conversion for performance when types are already validated. The ScalaSimpleFeature class now uses an internal array for values and supports both mutable and immutable variants, with the immutable version using Collections.unmodifiableMap for user data. Serialization options have been reorganized into a Builder pattern with implicit conversions, adding support for lazy evaluation and native collections. A new TransformSimpleFeature class enables on-the-fly attribute transformation by wrapping an underlying feature and applying transform evaluations.

geomesa-features/geomesa-feature-common/src/main · high confidence

Refactored Spark module with new SpatialRDDProvider architecture and Python serialization support

The geomesa-spark-core module has been decomposed and refactored to introduce a pluggable SpatialRDDProvider architecture, allowing different data stores to be integrated via a ServiceLoader-based provider pattern. This change includes a new RPC endpoint for distributing Kryo schema registrations across the Spark cluster to improve serialization performance, and adds specific serialization utilities (GeoMesaSeDerUtil) to support PySpark by mapping JTS Geometry types to WKB for Python deserialization. The public API is also updated with Java and Python bindings (JavaGeoMesaSpark, PythonGeoMesaSpark) that expose these new capabilities.

geomesa-spark/geomesa-spark-core · high confidence

Refactored Spark-JTS module with Spark 4 compatibility and Sedona integration

The geomesa-spark-jts module has been restructured to support Spark 4.2 and integrate with Apache Sedona. A new \AbstractGeometryUDT\ base class now delegates serialization to either GeoMesa's native WKB encoder or Apache Sedona's UDT, selectable via the \geomesa.use.sedona\ system property. To maintain compatibility with Spark 4's internal API changes, a new \ColumnUtils\ object bridges public \Column\ types to Catalyst internals, and the \initJTS\ initialization routine sets \spark.sql.functionResolution.sessionOrder\ to 'first' to ensure session-registered UDFs take precedence over Spark 4's native ST\\ builtins. The module also introduces \GeometryLiteralRules\ to handle serialization differences across Spark versions (GenericInternalRow, UnsafeRow, and primitive byte arrays) and provides comprehensive DataFrame DSL functions for geometric construction, conversion, and spatial analysis.

geomesa-spark/geomesa-spark-jts · high confidence

Refactored Z-curve indexing to use a unified SpaceFillingCurve interface and new Z2/Z3 implementations

The Z3 module's space-filling curve logic has been restructured to rely on a new \SpaceFillingCurve\ trait and dedicated \Z2\ and \Z3\ value classes for index generation, replacing the previous monolithic implementation. This change introduces specific curve implementations including \Z2SFC\, \S2SFC\, and \XZ2SFC\, alongside a \LegacyYearXZ3SFC\ to maintain backward compatibility with existing index keys that used an incorrect 52-week time bound. The refactoring also adds \XZ2Reversal\ utilities to decode index values back to bounding boxes and updates the \BinnedTime\ logic to support configurable time periods (weeks, days, months, years) with corresponding test coverage.

geomesa-z3 · high confidence

Register GeoMesa Convert2 service implementations via META-INF

The converter framework now automatically discovers and registers its core components through standard Java service-provider files in META-INF/services. This includes enabling CQL-based validation (CqlValidatorFactory), a suite of transformer functions (such as Cast, Collection, Date, Encoding, Enrichment, Geometry, Math, Scripting, and String functions), and specific feature-type validators (HasDtg, HasGeo, Id, Index, None). Additionally, it registers configuration providers (ClassPathConfigProvider, URLConfigProvider), enrichment cache factories, a composite feature converter, and metrics reporters (Console, Slf4j), ensuring these capabilities are available without manual wiring.

geomesa-convert/geomesa-convert-common/src/main/resources/META-INF · high confidence

Removal of legacy Accumulo DataStore factory registration

The service provider configuration file for GeoTools DataStore factories has been removed, specifically deleting the entry for \geomesa.core.data.AccumuloDataStoreFactory\. This change eliminates the automatic registration of the Accumulo data store implementation via the standard Java service loader mechanism in this module, likely as part of the broader namespace migration and index refactoring efforts.

geomesa-core/src/main/resources · high confidence

Removal of legacy index encoding and decoding infrastructure

The core index package has removed the legacy text-based encoding and decoding infrastructure, including the \Decoders\, \Extractors\, \Formatters\, and \IndexFormat\ files. This eliminates the previous mechanism for mapping spatio-temporal data to Accumulo keys using string-based formatters (such as \GeoHashTextFormatter\ and \DateTextFormatter\) and abstracts away the low-level key extraction logic, signaling a shift to a new internal representation for index entries.

geomesa-core/src/main/scala/geomesa/core/index · high confidence

Standardized metric naming and tagging for converters

The converter module now uses Micrometer for metrics, introducing a standardized naming convention where all converter metrics are prefixed with 'geomesa.converter' (configurable via the 'geomesa.convert.metrics.prefix' system property) and include a 'converter.name' tag to identify the specific converter instance.

geomesa-convert/geomesa-convert-common/src/main/scala/org/locationtech/geomesa/convert2/metrics · high confidence

StatsCombiner updated for cross-version compatibility

The StatsCombiner implementation has been updated to ensure compatibility across different versions of the underlying metadata storage. The code now includes a fallback mechanism in the reduce method to handle legacy row key formats (SingleRowMetadata) when decoding feature type names, preventing errors during stat combination if the new KeyValueStoreMetadata decoding fails. This change supports the broader upgrade to Accumulo 2.1 by ensuring the combiner can correctly process stats from tables that may still use older metadata serialization patterns.

geomesa-accumulo/geomesa-accumulo-iterators/src/main/scala/org/locationtech/geomesa/accumulo/combiners · high confidence

XML converter service registration and default configuration

The XML converter module now registers its factory implementations (XmlConverterFactory and XmlCompositeConverterFactory) and transformer functions via Java Service Provider Interface files, ensuring they are automatically discovered by the conversion framework. Additionally, a new default configuration file (xml-converter-defaults.conf) establishes baseline settings for XML parsing, including the use of Saxon's XPath factory, incremental parse mode, and index-based validation.

geomesa-convert/geomesa-convert-xml/src/main/resources · high confidence

Test coverage

6 commits adding/updating tests in geomesa-convert/geomesa-convert-parquet/src/test/resources; Add Spark SQL integration tests for Apache Sedona; Add integration tests for FileSystemRDDProvider with SeaweedFS; Added comprehensive test coverage for Kryo serialization; Added integration tests for Accumulo Spark provider; Added integration tests for JDBC converter with PostgreSQL; Added serialization performance benchmarking tool; Added test SPI configuration for MergedDataStoreView; Added test configuration for SimpleFeature conversion; Added test configuration for fixed-width converter SFT; Added test coverage for Accumulo index strategies and query planning; Added test coverage for Accumulo iterator behaviors; Added test coverage for Arrow IO writers and readers; Added test coverage for Avro feature serialization and data file I/O; Added test coverage for StatsCombiner configuration; Added test coverage for converter transform functions and predicates; Added test data for density iterator; Added test fixtures and updated Hadoop configuration for Accumulo datastore tests; Added test fixtures for static JavaScript functions; Added test for FeatureToFeatureConverter Java API; Added test for Kryo serialization with user data; Added test infrastructure for Accumulo data store validation; Added test infrastructure for Iceberg and SeaweedFS integration; Added test resource for shapefile encoding support; Added test resources for Avro schema registry conversion; Added test resources for Cassandra datastore integration; Added test resources for Geonames and Renegades data ingestion; Added test resources for XML feature parsing and BOM handling; Added test resources for converter validation; Added test resources for the FS datastore; Added test service registrations for converter transforms and validators; Added test to verify Java API for atomic writes; Added tests for Accumulo DataStore schema and security features; Added tests for Accumulo GeoServer plugin processes; Added tests for Accumulo audit query event transformation; Added tests for Accumulo export command and feature exporters; Added tests for Accumulo filter processing; Added tests for Accumulo ingest and update features commands; Added tests for Accumulo schema builder, batch multi-scanner, and batch writer configuration; Added tests for ArrowDataStore write, read, and filtering operations; Added tests for Avro converter path and union type handling; Added tests for Avro serialization backward compatibility; Added tests for Avro serialization of collections and SimpleFeatures; Added tests for AvroSchemaRegistryConverter; Added tests for CQL and custom validator implementations; Added tests for Converter and FileSystem DataStores; Added tests for FastFilterFactory optimization and behavior; Added tests for FeatureToFeatureConverter circular reference handling; Added tests for FixedWidthConverter validation and user data handling; Added tests for GeoJSON serialization of features and attributes; Added tests for HBase bin aggregation and WPS process integration; Added tests for JSON converter configuration and composite parsing logic; Added tests for Kryo JSON serialization and JSON path access; Added tests for Kryo serialization components; Added tests for Leaflet map rendering and CDN link inclusion; Added tests for MergedDataStoreView and RoutedDataStoreView; Added tests for Parquet converter functionality; Added tests for Redis enrichment cache functionality; Added tests for ScalaSimpleFeature attribute handling and created shared test resource symlink; Added tests for ShapefileConverter functionality; Added tests for SimpleFeatureVector encoding and wrapping; Added tests for XML converter composite and multi-feature parsing; Added tests for composite converters, enrichment caches, and scripting functions; Added tests for converter API, discovery, and URL configuration; Added tests for delimited text converter edge cases and validation; Added tests for the Java API for converters; Added tests for type inference logic; Added unit tests for AccumuloJobUtils; Added unit tests for Arrow geometry vector implementations; Added unit tests for Cassandra and ScyllaDB data stores; Added unit tests for feature export functionality; Added unit tests for filter bounds, helper, and logic utilities; Added unit tests for filter functions; Removal of legacy iterator and ingest test utilities; Removed legacy Accumulo data store and filter extraction tests; Removed legacy test configuration files; Removed obsolete DocumentationTest class; Removed obsolete index query planner and schema tests; Standardized test logging configuration; Test resources consolidated via symlinks; Test resources now use a symlink to a shared directory.

Dependencies

New Maven modules for documentation and Accumulo distribution packaging

This change introduces new Maven module definitions for the project's documentation build and the Accumulo binary distribution. The \docs/pom.xml\ configures the Sphinx-based documentation build using the \maven-antrun-plugin\ and \maven-resources-plugin\, with profiles for HTML, single-page HTML, and LaTeX output, while \docs/requirements.txt\ pins specific versions for Sphinx and its extensions. Additionally, new POMs for the Accumulo subsystem (\geomesa-accumulo-dist\, \geomesa-accumulo-distributed-runtime\, \geomesa-accumulo-gs-plugin\, \geomesa-accumulo-indices\, \geomesa-accumulo-iterators\) define the dependency graph, shading rules for Jackson and Netty, and assembly configurations required to package the Accumulo-specific GeoMesa artifacts and GeoServer plugin.

(dependencies) · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

This is the PUBLIC form of this artifact. Findings are listed in full, but the details of SECURITY findings — which rule fired, in which file, on which line, and how to fix it — are deliberately withheld, and any secret-scanner results are excluded entirely. Where detail is absent here it was REMOVED FOR PUBLICATION; it is not missing from the analysis. The complete artifact is available from the repository owner.

Score

  • CAI 52 → 66 (+14.2)
  • Rubric changed (rubric-2026.08.15 → rubric-2026.09.15) — scores are not directly comparable.

Lenses

  • Code Health 95 → 87 (-7.3)
  • Architecture 96 → 77 (-18.9)
  • Maturity 55 → 61 (+5.9)
  • Readiness 38 → 64 (+26.7)
  • Security 56 → 74 (+17.3)

Resolved (67)

  • Coverage not measured — test suite did not build
  • Dimension evaluation failed
  • Duplicated block (124 lines × 2) (geomesa-hbase/geomesa-hbase-rpc/src/main/java/org/locationtech/geomesa/hbase/proto/GeoMesaProto.java)
  • Duplicated block (16 lines × 2) (geomesa-trino/geomesa-trino-datastore/src/main/java/org/locationtech/geomesa/trino/datastore/TrinoSchemaDiscovery.java)
  • Duplicated block (171 lines × 2) (geomesa-trino/geomesa-trino-plugin/src/main/java/org/locationtech/geomesa/trino/spatial/iceberg/connector/SpatialConnectorMetadata.java)
  • Duplicated block (242 lines × 2) (geomesa-hbase/geomesa-hbase-rpc/src/main/java/org/locationtech/geomesa/hbase/proto/GeoMesaProto.java)
  • Duplicated block (34 lines × 2) (geomesa-hbase/geomesa-hbase-rpc/src/main/java/org/locationtech/geomesa/hbase/proto/GeoMesaProto.java)
  • Duplicated block (48 lines × 2) (geomesa-hbase/geomesa-hbase-rpc/src/main/java/org/locationtech/geomesa/hbase/proto/GeoMesaProto.java)
  • Duplicated block (5 lines × 2) (geomesa-trino/geomesa-trino-plugin/src/main/java/org/locationtech/geomesa/trino/spatial/iceberg/connector/BboxFilteringPageSource.java)
  • Duplicated block (7 lines × 2) (geomesa-accumulo/geomesa-accumulo-jobs/src/main/java/org/locationtech/geomesa/accumulo/jobs/mapreduce/interop/FeatureWriterJob.java)
  • Duplicated block (7 lines × 2) (geomesa-kafka/geomesa-kafka-datastore/src/main/java/org/locationtech/geomesa/kafka/utils/interop/GeoMessageProcessor.java)
  • Duplicated block (8 lines × 2) (geomesa-trino/geomesa-trino-plugin/src/main/java/org/locationtech/geomesa/trino/spatial/iceberg/connector/SpatialConnectorMetadata.java)
  • High: security finding (details withheld)
  • High: security finding (details withheld)
  • High: security finding (details withheld)
  • High: security finding (details withheld)
  • High: security finding (details withheld)
  • High: security finding (details withheld)
  • High: security finding (details withheld)
  • High: security finding (details withheld)
  • …and 47 more

New (628)

  • AbstractMultiPolygonVector.get (cognitive 22) (geomesa-arrow/geomesa-arrow-jts/src/main/java/org/locationtech/geomesa/arrow/jts/impl/AbstractMultiPolygonVector.java)
  • AbstractMultiPolygonVector.set (cognitive 22) (geomesa-arrow/geomesa-arrow-jts/src/main/java/org/locationtech/geomesa/arrow/jts/impl/AbstractMultiPolygonVector.java)
  • AbstractPolygonVector.get (cognitive 16) (geomesa-arrow/geomesa-arrow-jts/src/main/java/org/locationtech/geomesa/arrow/jts/impl/AbstractPolygonVector.java)
  • AbstractPolygonVector.set (cognitive 16) (geomesa-arrow/geomesa-arrow-jts/src/main/java/org/locationtech/geomesa/arrow/jts/impl/AbstractPolygonVector.java)
  • AccumuloAuditWriter.writeQueuedEvents (cognitive 17) (geomesa-accumulo/geomesa-accumulo-datastore/src/main/scala/org/locationtech/geomesa/accumulo/audit/AccumuloAuditWriter.scala)
  • AccumuloBulkIngestCommand.startIngest (cognitive 16) (geomesa-accumulo/geomesa-accumulo-tools/src/main/scala/org/locationtech/geomesa/accumulo/tools/ingest/AccumuloBulkIngestCommand.scala)
  • AccumuloCompactCommand.execute (cognitive 57) (geomesa-accumulo/geomesa-accumulo-tools/src/main/scala/org/locationtech/geomesa/accumulo/tools/data/AccumuloCompactCommand.scala)
  • AccumuloCompactCommand.execute (cyclomatic 29) (geomesa-accumulo/geomesa-accumulo-tools/src/main/scala/org/locationtech/geomesa/accumulo/tools/data/AccumuloCompactCommand.scala)
  • AccumuloDataStore.preSchemaCreate (cognitive 16) (geomesa-accumulo/geomesa-accumulo-datastore/src/main/scala/org/locationtech/geomesa/accumulo/data/AccumuloDataStore.scala)
  • AccumuloIndexAdapter.createQueryPlan (cognitive 33) (geomesa-accumulo/geomesa-accumulo-datastore/src/main/scala/org/locationtech/geomesa/accumulo/data/AccumuloIndexAdapter.scala)
  • AccumuloIndexAdapter.createTable (cognitive 19) (geomesa-accumulo/geomesa-accumulo-datastore/src/main/scala/org/locationtech/geomesa/accumulo/data/AccumuloIndexAdapter.scala)
  • AccumuloJoinIndexAdapter.createQueryPlan (cognitive 56) (geomesa-accumulo/geomesa-accumulo-datastore/src/main/scala/org/locationtech/geomesa/accumulo/data/AccumuloJoinIndexAdapter.scala)
  • AccumuloJoinIndexAdapter.createQueryPlan (cyclomatic 24) (geomesa-accumulo/geomesa-accumulo-datastore/src/main/scala/org/locationtech/geomesa/accumulo/data/AccumuloJoinIndexAdapter.scala)
  • AddIndexCommandExecutor.addIndex (cognitive 20) (geomesa-accumulo/geomesa-accumulo-tools/src/main/scala/org/locationtech/geomesa/accumulo/tools/data/AddIndexCommand.scala)
  • ArrowAttributeReader.apply (cyclomatic 22) (geomesa-arrow/geomesa-arrow-gt/src/main/scala/org/locationtech/geomesa/arrow/vector/ArrowAttributeReader.scala)
  • ArrowGeometryReader.apply (cognitive 27) (geomesa-arrow/geomesa-arrow-gt/src/main/scala/org/locationtech/geomesa/arrow/vector/ArrowAttributeReader.scala)
  • ArrowGeometryReader.apply (cyclomatic 21) (geomesa-arrow/geomesa-arrow-gt/src/main/scala/org/locationtech/geomesa/arrow/vector/ArrowAttributeReader.scala)
  • ArrowLineStringBBox.evaluate (cognitive 16) (geomesa-arrow/geomesa-arrow-gt/src/main/scala/org/locationtech/geomesa/arrow/filter/ArrowFilterOptimizer.scala)
  • AttributeIndexKeySpace.getTieredRangeBytes (cyclomatic 17) (geomesa-index-api/src/main/scala/org/locationtech/geomesa/index/index/attribute/AttributeIndexKeySpace.scala)
  • AttributeIndexKeySpace.toIndexKey (cognitive 20) (geomesa-index-api/src/main/scala/org/locationtech/geomesa/index/index/attribute/AttributeIndexKeySpace.scala)
  • …and 608 more

Changes since last survey

  • 48 commits — 38 feature/other, 10 fixes

By area

  • .github/workflows — 12 commits
  • (root) — 10 commits
  • geomesa-fs/geomesa-fs-storage — 6 commits
  • geomesa-gt/geomesa-gt-partitioning — 6 commits
  • geomesa-trino/geomesa-trino-datastore — 4 commits
  • geomesa-convert/geomesa-convert-avro — 2 commits
  • geomesa-trino/geomesa-trino-plugin — 2 commits
  • .mvn/rrf — 1 commit
  • build/scripts — 1 commit
  • docs/user — 1 commit
  • geomesa-cassandra/geomesa-cassandra-datastore — 1 commit
  • geomesa-kafka/geomesa-kafka-confluent — 1 commit
  • geomesa-utils-parent/geomesa-utils — 1 commit

Notable commits

  • fix: Docs - Fix ScalaDoc for implicit executor services (#10783)
  • fix: Docs - Fix required Maven version (#10782)
  • fix: FS/Trino - Fix iceberg field ids to support schema evolution (#10808)
  • fix: Fix closing resources on InterruptedException, not swallowing exceptions in close methods (#10803)
  • fix: Fix missing Jackson 3 in binary distributions (#10805)
  • fix: Fix spark integration tests not running after failsafe upgrade (#10835)
  • fix: Fix surefire not failing build when tests fail (#10800)
  • fix: Fix update-maven-toolchains script jdk version detection (#10784)
  • fix: Postgis - fix concurrency locks when creating schemas in parallel (#10822)
  • fix: Test - Fix potential arbitrary file access during zip extraction (#10844)
  • change: Action to upload SBOMs to Eclipse on tags (#10841)
  • change: Add code of conduct (#10842)
  • change: Add explicit groupIds for 3rd party Maven repos (#10804)
  • change: Add link to reporting vulnerabilities on GitHub (#10819)
  • change: Add release process to CONTRIBUTING.md (#10821)
  • change: Add zizmor scanning workflow (#10820)
  • change: Bump jackson to 2.22.3/3.2.3, httpcore5 to 5.4.4 (#10843)
  • change: Bump maven-surefire-plugin to 3.6.0 (#10831)
  • change: Bump parquet to 1.18.0 (#10795)
  • change: Bump to geotools 35.1 (#10790)
  • …and 28 more

Architecture

  • Containers 0 added · 0 removed · contexts 5 added · 0 removed · edges 0 added · 0 removed

Added bounded contexts (5)

  • geomesa-arrow-jts
  • geomesa-cqengine_2.12
  • geomesa-trino-datastore_2.12
  • geomesa-trino-plugin_2.12
  • repository

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

locationtech/geomesa was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 27 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit 10446803fd1712b29e4e52465d40b091726b10cb — the exact code this score is about.
  • Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-d00c643c3f66.