Skip to content
CAI
Software that uses CAICheck a score

gnieh/fs2-data

56.1

Adequate · 20 September 2026

23.3k

lines of production code

Scala

primary language

1

measurement over time

CAI band scale
CAI lens gauges

What this system is

This system is a functional streaming data processing library for Scala that provides parsers, serializers, and query engines for JSON, XML, CSV, CBOR, and MessagePack. It enables efficient, memory-conscious transformation of data streams through low-level token processing and high-level AST manipulation, supporting both Scala 2 and Scala 3. The library includes specialized tools for compile-time validated queries (JSONPath, XPath, jq), generic type derivation for CSV, and integration with popular ecosystem libraries like Circe and Play JSON.

How it got here

2019–2021 — CBOR module and CSV refactoring

28 changes.

This period focused on introducing a new CBOR module with both low-level streaming and high-level structured APIs, alongside a comprehensive refactoring of the CSV library's internal architecture. The CSV changes included migrating to a unified type class hierarchy, adding generic derivation support for Scala 2 and 3, and removing legacy parsing APIs. The work was supported by migrating the build system to sbt and upgrading dependencies to Scala 3.

2022 — Streaming query engines and format integrations

39 changes.

This period focused on implementing streaming XPath and JSONPath query capabilities using finite-state automata, alongside significant internal refactoring of XML and JSON parsers. It also expanded the library's ecosystem by adding integration modules for Circe, Play JSON, and Scala-XML, while introducing CBOR conversion support.

2023–2025 — jq engine overhaul and documentation migration

22 changes.

This period focused on rewriting the JSON jq query execution engine using an Event-Structure Parsing strategy and introducing a Scala 3 string interpolator for compile-time validation. Concurrently, the project migrated its entire documentation site to the Laika generator and expanded support for MessagePack serialization with high-level APIs and platform-specific type handling.

Features

Add CBOR diagnostic and debugging utilities

The CBOR module now includes a diagnostic pipe that converts CBOR data items into human-readable strings following the diagnostic notation defined in RFC 8949. This feature, along with a dedicated debug pipe for logging, allows users to inspect and troubleshoot CBOR streams during development. The update also introduces specific exception types for parsing, tag decoding, and validation errors, as well as support for half-precision floating-point values and standard CBOR tag constants.

cbor/shared/src/main/scala/fs2/data/cbor · high confidence

Add CBOR to JSON transformation

The \fs2.data.cbor\ package now includes a \decodeItems\ pipe that transforms a stream of CBOR items into a stream of JSON tokens. This implementation follows section 6.1 of RFC 8949, handling specific CBOR tags such as positive and negative big numbers (converting them to decimal strings), and expected encoding tags (base64, base64url, base16) by converting their byte contents into the appropriate string representations. It also supports indefinite-length text and byte strings, as well as maps and arrays, ensuring that keys and values are correctly tokenized for downstream JSON processing.

cbor-json/shared/src/main/scala/fs2/data/cbor · high confidence

Add Circe integration for JSON parsing and serialization

This change introduces a new integration module for the Circe JSON library within the fs2-data project. It provides the core \CirceBuilder\, \CirceTokenizer\, and implicit \Serializer\/\Deserializer\ instances that allow users to parse, serialize, and tokenize JSON data using Circe's \Json\ type. Additionally, it includes a comprehensive suite of tests (covering parsing, tokenization, JSONPath, selectors, and merge patches) that leverage the Jawn parser to ensure consistent cross-platform behavior, specifically working around fatal exceptions on Scala.js.

json/circe · high confidence

Add JSON data manipulation cookbook

Introduces a new cookbook demonstrating how to use fs2-data to build a jq-like CLI tool for reading, transforming, and writing JSON data. The documentation includes a high-level overview of the parsing and rendering pipeline, a basic implementation example, and sample data files (sample.json and nested.jsonl) to illustrate handling standard JSON arrays and JSON Lines with nested structures.

site/cookbooks · high confidence

Add JSON to CBOR streaming encoder

Introduces a new \encodeItems\ pipe in the \fs2.data.json.cbor\ package that converts JSON tokens into CBOR items. This encoder handles booleans, nulls, strings, and arrays/objects using indefinite-length encoding, while optimizing number representation by selecting the smallest exact CBOR integer or float type (avoiding half-floats) and using tagged big-integer/decimal formats for values that exceed standard size limits.

cbor-json/shared/src/main/scala/fs2/data/json · high confidence

Add Play JSON integration module

Introduces a new \json/play\ module that enables fs2-data to work with Play Framework's \JsValue\ type. This addition provides the necessary builder, tokenizer, serializer, and deserializer implementations to parse, serialize, and manipulate JSON data using Play's native JSON library, along with a comprehensive suite of tests covering tokenization, path selection, merge patches, and exception handling.

json/play · high confidence

Add Scala 2 JSON selector string interpolator

Introduces a new \selector\ string interpolator for Scala 2 users, enabling compile-time validation of JSON selector expressions. The implementation adds \SelectorInterpolator.scala\ and \package.scala\ in the \scala-2\ source directory, leveraging the \Literally\ macro to parse and validate selector strings at compile time, ensuring that invalid selectors are caught early rather than at runtime.

json/interpolators/src/main/scala-2, json/interpolators/src/main/scala-3 · high confidence

Add Scala 3 jq string interpolator

Introduces a new \jq\ string interpolator for Scala 3, enabling users to write JSON queries directly in code using compile-time validation and expression conversion. This adds support for Scala 3-specific syntax and macros in the JSON jq module, ensuring compatibility and fixing compilation issues on Scala 3 platforms.

json/src/main/scala-3/fs2/data/json/jq · high confidence

Add Scala 3 semiauto CSV derivation and tuple support

The CSV generic module now provides Scala 3-specific semiauto derivation capabilities via the new \semiauto\ object, enabling users to derive \RowDecoder\, \RowEncoder\, \CsvRowDecoder\, and \CsvRowEncoder\ instances for case classes using shapeless3. This includes support for named column mapping through the \StaticHeaders\ type class, allowing decoders to handle missing columns gracefully. Additionally, a new \tuples\ object provides automatic derivation of row decoders and encoders for Scala 3 tuple types, extending generic CSV support to tuple structures.

csv/generic/shared/src/main/scala-3/fs2/data/csv/generic · high confidence

Add Scala-XML integration for parsing and building documents

Users can now parse and construct XML documents using the standard scala-xml library. This change introduces a DocumentBuilder that converts fs2-data XML events into scala-xml Node/Elem structures and a DocumentEventifier that streams scala-xml nodes back into fs2-data XmlEvents, enabling seamless interoperability with existing scala-xml codebases.

xml/scala-xml/src/main · high confidence

Add abstract Table type class for lookup operations

A new \Table\ type class has been introduced in the \fs2.data.matching\ package to provide a unified abstraction for lookup tables, which is useful for implementing finite state automata or transducers. This addition includes implicit instances for \Map\, \Function\, and \PartialFunction\, along with extension methods to simplify lookup syntax, enabling more flexible and composable pattern matching logic within the library.

finite-state/shared/src/main/scala/fs2/data/matching · high confidence

Add documentation for the fs2-data-xml module

New documentation pages have been added to the site for the fs2-data-xml module, covering basic usage, XPath operators, library bindings, and XML renderers. The documentation is structured using Laika and includes examples for parsing XML streams, resolving namespaces and entities, building DOMs, and filtering data using XPath expressions.

site/documentation/xml · high confidence

Add jq-like JSON query example

A new example application demonstrating how to use fs2-data's JSON jq capabilities has been added to the examples directory. This tool allows users to execute jq-style queries on JSON data, supporting input from either a string or a file, and outputting results to stdout or a specified file.

examples/jqlike · high confidence

Add low-level CBOR data item parser

Introduces a new low-level CBOR parser in the \fs2.data.cbor.low.internal\ package, providing the foundational \ItemParser\ and \MajorType\ components required to decode CBOR streams into discrete data items. This change establishes the internal parsing logic for handling various CBOR major types (such as integers, byte strings, and text strings) and indefinite-length sequences, serving as the base layer for higher-level serialization and validation APIs.

cbor/shared/src/main/scala/fs2/data/cbor/low/internal · high confidence

Add platform-specific Date serialization and deserialization for JavaScript

The msgpack library for JavaScript now includes built-in support for serializing and deserializing \js.Date\ objects. This change introduces platform-specific instances that map Msgpack timestamp formats (32-bit, 64-bit, and 96-bit) to and from JavaScript Date objects, ensuring accurate round-trip conversion of date values in the high-level API.

msgpack/js · high confidence

Add platform-specific serialization and deserialization for Java Instant in msgpack high module

The \fs2.data.msgpack.high\ module now includes platform-specific instances for serializing and deserializing \java.time.Instant\ objects. This change introduces \PlatformSerializerInstances\ and \PlatformDeserializerInstances\ in the \jvm-native\ source directory, enabling the high-level API to correctly handle Instant types by mapping them to Msgpack timestamp formats (32, 64, or 96-bit) and back.

msgpack/jvm-native · high confidence

Add scalafix rule to migrate JSON parsing syntax

A new scalafix rule named 'json-parse' is introduced to automatically migrate code using the legacy \tokens\ and \values\ stream combinators to the new \ast.parse\ method. This change simplifies JSON parsing logic by replacing verbose \.through(tokens).through(values)\ chains with a single \.through(ast.parse)\ call, requiring the \fs2.data.json.\_\ import.

scalafix/rules · high confidence

Add search box to the navigation bar

The main navigation template now includes a search input field integrated into the sidebar. This feature is powered by Pagefind, which is initialized via a script to render the search interface within the designated element, allowing users to search the site content directly from the navigation menu.

site/helium · high confidence

Added CSV benchmark dataset

A new CSV file containing 113 sample records has been added to the benchmark resources to support CSV parsing performance tests.

benchmarks/src/main/resources · high confidence

Added JSON Merge Patch support for streaming JSON

The \fs2-data-json\ library now includes a new \mergepatch\ package that implements RFC 7396 JSON Merge Patch. This addition provides a streaming \patch\ function, allowing users to apply JSON Merge Patch documents to fs2 streams of JSON tokens. The implementation handles recursive object patching, value replacement, and key deletion (via null values) directly on the token stream, enabling efficient, memory-conscious patching operations without requiring full deserialization into intermediate data structures.

json/diffson/src/main · high confidence

Added literal and enum cell decoders/encoders

New source files have been added to the CSV module to support decoding and encoding of Scala literal types (String, Char, Byte, Short, Int, Long, Float, Double, Boolean) and Scala Enumeration values. The \LiteralCellDecoders\ and \LiteralCellEncoders\ traits (in the scala-2.13+ directory) provide implicit decoders/encoders that match specific constant values using \ValueOf\, while \EnumDecoders\ and \EnumEncoders\ (in the scala-2 directory) provide decoders/encoders for \Enumeration\ values by matching string representations.

csv/shared/src/main/scala-2.13+ · high confidence

Added scalafix rules for CSV, JSON, and XML stream patterns

New scalafix rules have been added to automatically refactor stream processing code for CSV, JSON, and XML data formats. The \EmitsRows\ rule for CSV, \EmitsTokens\ rule for JSON, and \EmitsEvents\ rule for XML detect patterns where \Stream.emits\ is combined with \flatMap\ before passing data through format-specific operators (\rows\, \tokens\, or \events\). These rules simplify the code by replacing the verbose \flatMap\/\emits\ pattern with a direct call to the respective operator, while also ensuring the correct type parameters (such as \String\ or \Char\) are explicitly provided.

scalafix/src · high confidence

Compile-time XPath string interpolation

Users can now use the \xpath"..."\ string interpolator to write XPath expressions directly in code. This feature leverages compile-time parsing via the \org.typelevel.literally\ library to validate and convert string literals into \XPath\ objects, enabling early detection of syntax errors and potentially better performance by avoiding runtime parsing overhead.

xml/src/main/scala-2 · high confidence

Compile-time validated JSON Path literals

A new \jsonpath\ string interpolator is available in the \fs2.data.json\ package, allowing JSON paths to be defined as literals that are parsed and validated at compile time rather than at runtime. This change introduces a \JsonPathInterpolator\ backed by the \Literally\ library, which ensures that invalid JSON path syntax is caught during compilation, improving safety and reducing runtime errors for users constructing JSON queries.

json/src/main/scala-2/fs2/data/json/jsonpath, json/src/main/scala-3/fs2/data/json/jsonpath · high confidence

Compile-time validated XPath string interpolator

Users can now use the \xpath"..."\ string interpolator to construct XPath expressions at compile time. This new feature, implemented in \fs2.data.xml.xpath.literals\, leverages Scala 3 macros and the \org.typelevel.literally\ library to parse and validate XPath syntax during compilation, ensuring that invalid expressions are caught early rather than at runtime.

xml/src/main/scala-3 · high confidence

Introduce XmlQueryPipe for DFA-based XML query execution

Added XmlQueryPipe, a new internal component that executes XPath queries by processing XML events through a Deterministic Finite Automaton (DFA). This pipe handles the mapping of XML start and end tags to automaton states, enabling efficient pattern matching for XML data selection.

xml/src/main/scala/fs2/data/xml/xpath/internals · high confidence

Introduce core ESP typeclasses and exception handling

Added foundational support for the ESP (Event Stream Processing) module by introducing the \Conversion\ typeclass for creating events from tags, the \Tag2Tag\ typeclass for transforming input tags to output tags, and a dedicated \ESPException\ class for error handling. These components provide the necessary abstractions for tag manipulation and event generation within the finite-state data processing pipeline.

finite-state/shared/src/main/scala/fs2/data/esp · high confidence

Introduce core pattern matching type classes and exception handling

Added foundational type classes \IsPattern\ and \IsTag\ to the \fs2.data.pattern\ package to support pattern decomposition and tag range checking, alongside a new \PatternException\ class for error reporting. These additions provide the underlying infrastructure for pattern matching logic within the finite-state module.

finite-state/shared/src/main/scala/fs2/data/pattern · high confidence

Introduce extensible XML tree building and eventification abstractions

The \fs2.data.xml.dom\ package now provides new traits (\DocumentBuilder\, \ElementBuilder\, and \DocumentEventifier\) and a \TreeParser\ implementation that allow users to define custom XML document and element structures. This change introduces \documents\ and \elements\ pipes to parse XML events into these custom tree types, as well as an \eventify\ pipe to convert nodes back into events, enabling the library to support alternative XML DOM implementations beyond the default.

xml/src/main/scala/fs2/data/xml/dom · high confidence

Introduce high-level CBOR API for structured data processing

A new high-level API is added to the CBOR library, providing structured representations of CBOR data through the \fs2.data.cbor.high\ package. This includes new \CborValue\ types and convenience pipes (\values\, \parseValues\, \toItems\, \toBinary\) that allow users to parse byte streams into a higher-level AST and serialize them back, abstracting away the low-level item details. The package also includes a deprecated \HalfFloat\ object for binary compatibility, directing users to the main \fs2.data.cbor.HalfFloat\ implementation.

cbor/shared/src/main/scala/fs2/data/cbor/high · high confidence

Introduce high-level CBOR value parsing and serialization

Added internal \ValueParser\ and \ValueSerializer\ components to the high-level CBOR API, enabling the conversion between low-level \CborItem\ streams and a structured \CborValue\ representation. This change allows users to parse CBOR data into a tree-like structure (supporting arrays, maps, text/byte strings, and tagged values) and serialize \CborValue\ objects back to CBOR items, including proper handling of large integers via bignum tags.

cbor/shared/src/main/scala/fs2/data/cbor/high/internal · high confidence

Introduce internal types for compiled and piped jq query execution

Added the \CompiledJq\ trait and \PipedCompiledJq\ class to the \fs2.data.json.jq\ package. These components provide the internal infrastructure for executing compiled jq queries, allowing individual queries to be chained together via piping (\\|\) to process JSON token streams sequentially.

json/src/main/scala/fs2/data/json/jq · high confidence

Introduce low-level CBOR streaming API

Added a new low-level CBOR API in the \fs2.data.cbor.low\ package that provides stream-based parsing, validation, and serialization. This API allows users to parse arbitrary-length CBOR data streams into a flat sequence of \CborItem\s without building a full AST, which is more efficient for large collections. It includes pipes to validate item streams, convert items back to binary bytes (with and without validation), and handle indefinite-length constructs, following the CBOR RFC structure closely.

cbor/shared/src/main/scala/fs2/data/cbor/low · high confidence

Introduce predicate-based finite automata and tree query capabilities

This change adds new internal components to the \fs2.data.pfsa\ package to support automata that operate on predicates rather than simple symbols. It introduces the \Pred\ typeclass for defining and combining predicates, along with \Candidate\ for selection logic. The core implementation includes \PDFA\ (Deterministic Finite Automaton with Predicates) for efficient recognition and \PNFA\ (Nondeterministic Finite Automaton with Predicates) which can be determinized into a PDFA. Additionally, a \TreeQueryPipe\ abstract class is provided to enable recursive queries on tree-like structures (such as XPath or JsonPath) by leveraging the PDFA to match opening, closing, and internal tokens.

finite-state/shared/src/main/scala/fs2/data/pfsa · high confidence

Introduce streaming XPath filtering API

Adds a new \filter\ namespace in the \fs2.data.xml.xpath\ package, providing streaming XPath query capabilities for XML event streams. Users can now select matching elements using methods like \unsafeRaw\ (emitting raw event streams), \first\ (selecting the first match), \through\ (transforming matches in parallel with optional deterministic ordering), \dom\ (building element DOMs), \consume\ (side-effect processing), and \collect\ (aggregating results). The implementation compiles XPath expressions into a Deterministic Finite Automaton (PDFA) to process events in a streaming fashion, emitting matches as early as possible.

xml/src/main/scala/fs2/data/xml/xpath · high confidence

Introduce streaming pretty-printing framework for text rendering

Adds a new \fs2.data.text.render\ module providing a streaming pretty-printing infrastructure. This includes core typeclasses (\Renderable\, \Renderer\) for defining how data events are converted into document structures, utility methods for splitting text into words and handling indentation, and a \pretty\ pipe that formats these document streams into human-readable strings with configurable width and indentation.

text/shared/src/main/scala/fs2/data/text/render · high confidence

Introduces automatic and semi-automatic CSV derivation for Scala 2

The \csv/generic\ module for Scala 2 now provides automatic and semi-automatic derivation of decoders and encoders for case classes and HLists. Users can import \fs2.data.csv.generic.auto.\_\ to enable implicit derivation for \RowDecoder\, \RowEncoder\, \CsvRowDecoder\, \CsvRowEncoder\, \CellDecoder\, and \CellEncoder\ without manual boilerplate. Alternatively, the \semiauto\ object allows explicit derivation of these instances. The implementation supports both sequence-shaped and map-shaped CSV rows, handles default values, and respects \CsvName\ annotations for custom column mapping.

csv/generic/shared/src/main/scala-2 · high confidence

Introduces streaming JSONPath filtering operators

The \fs2.data.json.jsonpath\ package now provides a suite of streaming pipes for querying JSON data, including \unsafeRaw\, \first\, \through\, \values\, \deserialize\, \consume\, and \collect\. These operators allow users to filter JSON token streams based on JSONPath expressions, with configurable limits on the number of matches (\maxMatch\) and nesting depth (\maxNest\). A key behavioral feature is the \deterministic\ flag (defaulting to \true\), which controls whether results are emitted in the order they appear in the input stream or as soon as they are fully built, enabling flexible trade-offs between ordering guarantees and latency.

json/src/main/scala/fs2/data/json/jsonpath · high confidence

Introduction of JSON codec and selector DSL abstractions

The library now exposes core abstractions for JSON transformation and selection. New \Deserializer\ and \Serializer\ traits in the \codec\ package define how JSON ASTs are converted to and from Scala values, enabling new stream operators like \transform\, \transformOpt\, \transformF\, \transformOptF\, \deserialize\, and \serialize\ that allow users to modify or extract JSON data via functional pipelines. Additionally, a new \selector\ package introduces a DSL for building JSON selectors (starting with \root\), providing a structured way to target specific parts of a JSON stream for these transformations.

json/src/main/scala/fs2/data/json/codec, json/src/main/scala/fs2/data/json/selector · high confidence

New CBOR module and CSV generic derivation annotations

This release introduces a new CBOR module with low-level item parsing and validation, alongside high-level model definitions. It also adds generic derivation support for CSV, including \CsvName\ and \CsvValue\ annotations to customize field and sealed trait representations during codec derivation.

repository · high confidence

New JSON parsing, transformation, and rendering utilities

The \fs2.data.json\ package now exposes core stream-processing capabilities, including \tokens\ for parsing character streams into JSON tokens, and \unwrap.stripTopLevelArray\ to optionally remove surrounding array brackets from token streams. It also introduces \render\ for converting token streams back into compact JSON text. Additionally, legacy filtering and transformation pipes (\filter\, \transform\, \transformOpt\, \transformF\, \values\, \tokenize\) are now marked as deprecated in favor of newer \jsonpath\ and \ast\ equivalents.

json/src/main/scala/fs2/data/json · high confidence

New MFT builder DSL for defining finite-state machines

A new domain-specific language (DSL) has been added to the \fs2.data.mft\ package to simplify the construction of Multi-Forest Transducers. The \package.scala\ file introduces a \dsl\ entry point and a suite of helper functions (such as \state\, \any\, \aNode\, \leaf\, and forest manipulation helpers like \x0\, \x1\, \node\, \leaf\) that allow users to build MFT structures in a more intuitive, tree-like manner using an implicit \MFTBuilder\ context.

finite-state/shared/src/main/scala/fs2/data/mft · high confidence

New high-level MessagePack serialization and deserialization API

The \fs2.data.msgpack.high\ module now provides a new high-level API for converting between Scala types and MessagePack format. This includes \MsgpackSerializer\ and \MsgpackDeserializer\ typeclasses with built-in instances for standard types (Int, Long, BigInt, String, List, Map, etc.) and a generic \MsgpackValue\ AST. Users can now use \serialize\ and \deserialize\ pipes to convert streams of Scala values to/from MessagePack bytes, leveraging a new internal parser and serializer that handle MessagePack headers and item formats.

msgpack/shared/src/main · high confidence

New tree query compiler for MFT translation

Added a new \QueryCompiler\ in the \fs2.data.mft.query\ package that translates abstract tree queries (represented by nested for/let clauses and paths) into Macro Forest Transducers (MFT). This compiler implements the XQuery Streaming by Forest Transducers approach, allowing users to compile high-level query structures into optimized MFTs with configurable optimization passes.

finite-state/shared/src/main/scala/fs2/data/mft/query · high confidence

Removals

Removal of legacy CSV parsing API

The legacy CSV parsing implementation in \fs2.data.csv.package.scala\ has been removed. This deletes the \fromBytes\ and \fromString\ pipes that previously handled CSV parsing using line-based splitting and \ApplicativeError\ for error handling. Users relying on these specific entry points will need to migrate to the newly introduced CSV API components, which offer a different abstraction for reading and decoding CSV data.

csv · high confidence

Behavioural changes

6 commits (1 fix) modifying json/src/main/scala-2/fs2/data/json/jq

A change to existing behaviour in json/src/main/scala-2/fs2/data/json/jq — 6 commits (1 fix), 1 file.

json/src/main/scala-2/fs2/data/json/jq · medium confidence · unverified

Add internal utility for untagging TaggedJson tokens

A new internal helper function \untag\ has been added to the \fs2.data.json.tagged\ package. This utility converts \TaggedJson\ instances into standard \Token\ objects, handling specific tagged variants like \StartArrayElement\, \EndArrayElement\, and structural markers by returning \None\ where appropriate, while mapping raw tokens and object keys to their corresponding \Token\ representations.

json/src/main/scala/fs2/data/json/tagged · high confidence

Added Scala 2.12 compatibility shims for collection builders

A new package object in the Scala 2.12 source directory provides implicit extension methods (\addOne\) for \VectorBuilder\ and \ListBuffer\. This change ensures the library compiles and functions correctly on Scala 2.12 by supplying compatibility shims for collection operations that may differ or be missing in that specific version.

json/src/main/scala-2.12 · high confidence

Introduce new internal XML parsing and rendering architecture

The \fs2.data.xml.internals\ package has been replaced with a new set of components: \EventParser\ handles the low-level character stream parsing with support for XML 1.1 validation and comment preservation; \Normalizer\ merges adjacent text nodes and attributes; \ReferenceResolver\ expands character and entity references; and \Renderer\ provides configurable XML serialization with pretty-printing, indentation, and empty-tag collapsing. These changes alter the internal event flow and output formatting behavior of the XML library.

xml/src/main/scala/fs2/data/xml/internals · high confidence

Introduces an ESP-based JSON path compiler and execution engine

The library now uses a new internal execution strategy based on Event-Structure Parsing (ESP) for processing JSON path queries. This change replaces the previous implementation with \ESPJqCompiler\ to translate JSON path filters into regular expression matchers and \ESPCompiledJq\ to execute them via a stream pipeline. This new engine handles path components such as root, indices, slices, fields, and recursive descent, providing the underlying mechanism for JSON query evaluation.

json/src/main/scala/fs2/data/json/jq/internal · high confidence

Migrate documentation site to Laika

The documentation site has been migrated from the previous Ruby-based template system to Laika. This change introduces a new directory structure and navigation configuration (directory.conf) that explicitly orders the documentation pages for formats like CSV, JSON, XML, CBOR, and MsgPack, and replaces the old index content with updated introductory material explaining the library's streaming parsing architecture.

site/documentation · high confidence

Migration from Mill to sbt and introduction of Nix development shell

The project has switched its build system from Mill to sbt, removing the \build.sc\ file and adding \.sbtopts\ to configure JVM memory and read timeouts. To support this transition and provide an isolated development environment, a Nix flake (\flake.nix\ and \flake.lock\) has been added, allowing developers to use \nix develop\ for setup. Additionally, the code formatter configuration has been updated to scalafmt 3.11.5 with stricter literal casing rules and Scala 3 dialect support.

(repo-wide) · high confidence

New CharLikeChunks typeclass for efficient character stream iteration

A new \CharLikeChunks\ typeclass and its implementations have been added to the \fs2-data\ text module, enabling efficient iteration over characters in streams of \Char\, \String\, and various byte encodings (UTF-8, ASCII, ISO-8859-1, ISO-8859-15). This change introduces a specialized buffer mechanism that reduces copying overhead by sharing buffers between chunks, particularly benefiting UTF-8 byte streams which are now decoded into a reusable character buffer. The \package.scala\ file exposes implicit conversions for these encodings via \utf8\, \ascii\, \latin1\, and \latin9\ objects, allowing users to seamlessly treat byte streams as character streams. Deprecated implicit methods from previous versions are retained for binary compatibility but marked for removal in future releases.

text/shared/src/main/scala/fs2/data/text · high confidence

New streaming JSON documentation site

The JSON module documentation has been migrated to the Laika site generator and restructured into a dedicated section. The new site covers core parsing and AST building, serializers and deserializers, JSON renderers, and generating JSON streams. It also introduces dedicated pages for the experimental jq-like query language, JSONPath filtering, JSON Patch integration via diffson, and advanced transformation pipes using selectors and the selector DSL.

site/documentation/json · high confidence

Refactor CSV platform-specific encoder/decoder traits for JS/Native targets

The CSV module now includes platform-specific trait definitions for cell encoders and decoders in the js-native source directory. A new \PlatformCellDecoders\ trait defines the \javaUriDecoder\, while \PlatformCellEncoders\ (renamed from the previous \CsvException\ file) provides an empty trait structure, establishing the interface layer for JS/Native platform implementations.

csv/js-native · medium confidence

Refactor internal CSV parsing structures for Scala 3 compatibility

The internal implementation of CSV parsing has been refactored to support Scala 3. Specifically, the private \Headers\ and \ParseRowResult\ traits, which previously managed internal state for header initialization and row parsing results, have been removed. They have been replaced by new, empty \EnumDecoders\ and \EnumEncoders\ traits in the shared Scala 3 source directory, indicating a shift in how enum decoding and encoding logic is structured for the new compiler version.

csv/shared/src/main/scala-3 · medium confidence

Refactored CSV encoding and decoding into a unified Cell/Row type class hierarchy

The CSV library's internal architecture has been restructured to use a new \CellDecoder\/\CellEncoder\ type class for individual cell values and \RowDecoder\/\RowEncoder\ (along with their header-aware \CsvRowDecoder\/\CsvRowEncoder\ and generic \RowDecoderF\/\RowEncoderF\ variants) for entire rows. This change introduces \forColumns\ helper methods that allow users to decode or encode rows by specifying column headers or indices, replacing the previous ad-hoc parsing logic. The \CsvRow\ type now strictly validates that the number of values matches the number of headers, and the \EscapeMode\ trait provides explicit control over CSV escaping behavior (Auto, Always, Never).

csv/shared/src/main/scala/fs2/data/csv · high confidence

Refactored CSV internals into dedicated parser, writer, and row parser modules

The internal CSV processing logic has been reorganized into three new modules: CsvRowParser, RowParser, and RowWriter. CsvRowParser introduces an \attempt\ variant that allows consumers to handle parse errors on a per-row basis by emitting \Left\ elements instead of failing the stream, while also narrowing error types to \CsvException\. RowParser handles the low-level character stream parsing, including state management for quoted fields and line tracking. RowWriter manages column encoding, handling escaping and quoting based on the configured \EscapeMode\. This refactoring also migrates imports from \cats.implicits\ to \cats.syntax.all\ and updates collection usage to \Chunk.from\.

csv/shared/src/main/scala/fs2/data/csv/internals · high confidence

Refactored JSON parsing internals with configurable buffer capacities and new AST builder abstraction

The internal JSON parsing implementation has been restructured to improve performance and configurability. The \TokenParser\ now exposes system properties to configure buffer capacities for keys (\fs2.data.json.key-buffer-capacity\, default 64), numbers (\fs2.data.json.number-buffer-capacity\, default 16), and strings (\fs2.data.json.string-buffer-capacity\, default 128), allowing users to tune memory usage for specific workloads. Internally, the parser now uses a new \Builder\[Json\]\ trait and \ChunkAccumulator\ abstraction to construct AST values, replacing previous accumulation logic. This change includes a new \BuilderChunkAccumulator\ that builds JSON structures directly and a \LegacyTokenParser\ fallback for non-buffered character streams, ensuring compatibility while optimizing the primary parsing path.

json/src/main/scala/fs2/data/json/internal · high confidence

Refactored Scala 3 CSV derivation internals for compatibility and correctness

The internal derivation logic for Scala 3 has been rewritten to use Shapeless 3 and Scala 3 native features, introducing new internal components like CellValue, DerivedCellDecoder, DerivedCellEncoder, and OptCellDecoder. This change fixes handling of missing named columns, resolves binary compatibility issues for OptCellDecoder, and ensures that the Names given instance remains accessible to callers of deriveCsvRowDecoder/Encoder to prevent future breakage in Scala 3.7+.

csv/generic/shared/src/main/scala-3/fs2/data/csv/generic/internal · high confidence

Restored URL encoding and decoding support on the JVM

The CSV library now supports encoding and decoding \java.net.URL\ values on the JVM platform. This change introduces platform-specific encoder and decoder traits that provide \CellEncoder\[URL\]\ and \CellDecoder\[URL\]\ instances, restoring functionality that was previously unavailable or removed, while noting that URL support remains excluded from the Scala.js target.

csv/jvm · high confidence

Site migrated to Laika documentation generator

The project website has been rebuilt using the Laika static site generator, replacing the previous setup. This change introduces a new site structure with a custom navigation order defined in directory.conf, a CNAME record pointing to fs2-data.gnieh.org, and an updated home page that lists all available fs2-data modules (JSON, XML, CSV, CBOR, MessagePack) along with their respective ecosystem integrations and current adopters.

site · high confidence

Stubbed literal cell decoders and encoders for Scala 2.12

The library now provides empty trait stubs for \LiteralCellDecoders\ and \LiteralCellEncoders\ in the Scala 2.12 source directory. Since literal types are not supported in Scala 2.12, these traits are intentionally left empty to satisfy the shared API structure without enabling literal-based CSV cell conversion for this version.

csv/shared/src/main/scala-2.12- · high confidence

XML rendering and parsing enhancements

The XML module now supports pretty-printing output via new \render.prettyPrint\ and \collector.pretty\ methods, allowing users to generate indented, human-readable XML strings. Parsing capabilities have been extended with an \includeComments\ parameter on the \events\ pipe to optionally preserve comment nodes, and a new \referenceResolver\ pipe is available to handle character and entity references. Additionally, the legacy \render\ and \collector.show\ methods are deprecated in favor of the new \render.raw\ and \collector.raw\ APIs.

xml/src/main/scala/fs2/data/xml · high confidence

Test coverage

Add JMH benchmarks for CSV, JSON, MessagePack, XML, and XPath performance; Added CSV parsing test fixtures for edge cases; Added JSON test suite resources for edge-case parsing; Added Sun XML test suite resources for validation testing; Added comprehensive CBOR test suite; Added comprehensive test coverage for CBOR-JSON conversion; Added comprehensive test suite for CSV parsing, encoding, and decoding; Added comprehensive test suite for the new MessagePack serialization and parsing implementation; Added test coverage for ESP transformation capabilities; Added test coverage for generic CSV derivation; Added test for JSON selector literal parsing; Added test input for the json-parse scalafix rule; Added test utilities for MiniXML and MiniXPath; Added tests for CSV generic auto-derivation and HList encoding/decoding; Added tests for JSON Merge Patch operations; Added tests for MFT article-to-HTML transformation and pattern matching compiler; Added tests for MFT query compilation and execution; Added tests for Scala 3 tuple derivation and inaccessible name regression; Added tests for literal decoders and encoders; Added tests for pattern matching with and without guards; Added tests for regular language operations; Added tests for scala-xml integration; Added tests for the new \parse\ migration rule.

Dependencies

Upgrade to fs2 3.14 and Scala 3.3.7

The build configuration has been updated to use fs2 3.14.0, circe 0.14.16, and Scala 3.3.7 (alongside Scala 2.12.21 and 2.13.18). This upgrade brings users the latest features and performance improvements from the fs2 streaming library and the Scala compiler, while maintaining binary compatibility through MiMa filters for internal text rendering classes.

(dependencies) · high confidence

Housekeeping

Migrate CSV module documentation to Laika

The documentation for the \fs2-data-csv\ and \fs2-data-csv-generic\ modules has been migrated to the Laika documentation framework. This update introduces a new site structure with a \directory.conf\ file to define navigation order and new Markdown files (\index.md\, \generic.md\) that detail the high-level and low-level APIs, including usage examples for automatic and semi-automatic derivation of decoders and encoders.

site/documentation/csv · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

Baseline

  • First survey — no prior run to compare against. CAI 56.

Lenses

  • Code Health 79
  • Architecture 66
  • Maturity 49
  • Readiness 53
  • Security 80

Changes since last survey

  • 300 commits — 279 feature/other, 21 fixes

By area

  • (repo) — 105 commits
  • (root) — 68 commits
  • msgpack/shared — 32 commits
  • project/build.properties — 15 commits
  • project/plugins.sbt — 14 commits
  • json/src — 9 commits
  • csv/generic — 7 commits
  • site/cookbooks — 7 commits
  • .github/workflows — 5 commits
  • benchmarks/src — 5 commits
  • msgpack/src — 5 commits
  • site/documentation — 5 commits
  • xml/src — 5 commits
  • cbor-json/shared — 4 commits
  • csv/shared — 4 commits
  • finite-state/shared — 2 commits
  • msgpack/js — 2 commits
  • site/index.md — 2 commits
  • text/shared — 2 commits
  • json/circe — 1 commit

Notable commits

  • fix: Add test to check that parsing is fixed
  • fix: Bug fix for issue 745
  • fix: Fix CBOR encoding of JSON numbers with 3, 5, 6 or 7 bytes
  • fix: Fix NFA building when there are alternatives
  • fix: Fix accidental code block start in site
  • fix: Fix binary compatibility
  • fix: Fix doc for valuesFromItems
  • fix: Fix msgpack BigInt deserializer for large values
  • fix: Fix msgpack js date deserializer
  • fix: Fix parsing of alternatives
  • fix: Fix unused warnings
  • fix: Fix warning about numeric widening
  • fix: Fix wrong header in msgpack map serializer
  • fix: Merge pull request #676 from gnieh/fix-alternative-dfa-port
  • fix: Merge pull request #746 from dwalend/fix/string-chunk-escape-duplication
  • fix: Merge pull request #759 from gnieh/fix-cbor-json
  • fix: Merge pull request #783 from Dichotomia/fix/779-csv-generic-names-accessible
  • fix: Revert accidental indent
  • fix: add explicit dependency on munit to fix native tests
  • fix: fix build by updating sbt typelevel
  • …and 280 more

Architecture

  • 0 containers · 1 bounded contexts · 0 dependency edges (baseline)

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

gnieh/fs2-data was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 20 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit 0351e3b3ed05566e3629b0979d4fe59f8a3a4064 — the exact code this score is about.
  • Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-b51f968c9b10.