Skip to content
CAI
Software that uses CAICheck a score

philss/floki

71.0

Strong · 23 September 2026

8.7k

lines of production code

Elixir

with Erlang

5

measurements over time

CAI band scale
CAI trend line
CAI lens gauges

What this system is

This system is a high-performance HTML parsing and querying library for Elixir, designed to handle document and fragment parsing with configurable backends. It provides a modernized API for traversing HTML trees, extracting text, and manipulating nodes, while supporting advanced CSS selector matching including pseudo-classes and combinators. The library ensures standards compliance through a native Elixir tokenizer and optional Rust-based parsers, backed by comprehensive test suites generated from official HTML specifications.

How it got here

2014 — API modernization and project scaffolding

5 changes.

The project underwent significant structural changes, including initial repository scaffolding and the complete rewrite of the Floki library's public API with new type definitions and parsing entry points. This period also involved expanding the test suite to support configurable parser execution and updating dependencies to require Elixir 1.15.

2015–2016 — CSS selector and parser refactoring

4 changes.

This period focused on a major internal restructuring of the Floki library, introducing a new modular architecture for HTML parsing and CSS selector matching. The work replaced monolithic logic with specialized modules for attributes, combinators, and pseudo-classes, enabling support for advanced CSS features like :has and :not. Comprehensive test coverage was added to validate the new tree representation, text extraction strategies, and selector engine.

2017–2021 — HTML5-compliant parser rewrite

7 changes.

This period focused on replacing the legacy Mochiweb-based parsing with a native Elixir HTML tokenizer compliant with WHATWG/W3C specifications. The work introduced new optional backends like FastHTML and Html5ever, alongside comprehensive test generation and benchmarking to ensure accuracy and performance.

Features

Added Mix tasks to generate HTML entities and tokenizer tests

Two new Mix tasks have been added to automate the generation of internal modules from external data sources. The \generate\_entities\ task reads \priv/entities.json\ to produce \lib/floki/entities/codepoints.ex\, providing a built-in lookup for HTML entity codepoints. The \generate\_tokenizer\_tests\ task processes WHATWG HTML5lib test files from \test/html5lib-tests/tokenizer\ to generate corresponding Elixir test modules in \test/floki/html/generated/tokenizer\, ensuring the tokenizer stays aligned with the latest HTML specifications.

lib/mix · high confidence

Added benchmark suite for HTML parsing and tokenization

New benchmark scripts have been added to the \benchs\ directory to measure performance using the Benchee library. These include \parse\_document.exs\ for comparing parsing backends (mochiweb, html5ever, fast\_html), \finder.exs\ for selector query performance, \raw\_html.exs\ for serialization speed, and \tokenizers.exs\ for comparing the mochiweb and Floki tokenizers. Supporting shell scripts (\compress.sh\, \extract.sh\) are also included to manage the HTML test datasets.

benchs · high confidence

Initial project scaffolding and configuration

The repository is initialized with essential project infrastructure, including a CHANGELOG.md following the Keep a Changelog format, a Code of Conduct, and a CONTRIBUTING guide. Development tooling is configured via .credo.exs for static analysis and .formatter.exs for code formatting, while .gitignore and .gitmodules are set up to manage build artifacts and external test fixtures.

(repo-wide) · high confidence

New HTML parser backends: FastHTML and Html5ever

Floki now includes two new optional HTML parser implementations: FastHTML, which leverages the :fast\_html NIF for parsing documents and fragments, and Html5ever, which uses the Html5ever Rust library to parse documents and supports the new attributes-as-maps feature. The existing Mochiweb parser remains the default but is now explicitly documented as non-HTML5-compliant; it also gains support for parsing attributes as maps. Users can switch parsers by configuring the HTMLParser module, gaining access to potentially faster or more standards-compliant parsing depending on the chosen backend.

_lib/floki/html\parser · high confidence

Behavioural changes

Floki library rewritten with new public API and type definitions

The lib/floki.ex file has been completely rewritten to expose a modernized public API. The module now defines comprehensive types for HTML nodes, trees, and CSS selectors, and introduces new entry points \parse\_document/1\, \parse\_document!/1\, \parse\_fragment/1\, and \parse\_fragment!/1\ to replace older parsing functions. It also adds an \attributes\_as\_maps\ option to the parser, allowing users to receive HTML attributes as maps instead of lists. This change represents a significant structural update to the library's core interface.

lib · high confidence

Floki v0.20.2: Major internal refactor and new CSS/text utilities

This release introduces a significant internal restructuring of the Floki library, adding new modules for CSS escaping (Floki.CSSEscape), HTML entity handling (Floki.Entities), and text extraction strategies (Floki.DeepText, Floki.FlatText). It also adds a new filter\_out helper (Floki.FilterOut) and a configurable HTML parser dispatch layer (Floki.HTMLParser) that supports alternative parsers like html5ever. The core tree representation is now backed by a new Floki.HTMLTree module with explicit node types (HTMLNode, Text, Comment) and an ID seeder, replacing the previous tuple-based traversal logic. These changes improve performance, add support for advanced CSS selectors and pseudo-classes, and provide more flexible text extraction options including input value inclusion and custom separators.

lib/floki · high confidence

New HTML parser and CSS selector lexer

The library now includes a new HTML parser module (\floki\_mochi\_html\) and a dedicated CSS selector lexer (\floki\_selector\_lexer.xrl\). The parser supports an \attributes\_as\_maps\ option to return attributes as maps instead of lists, handles specific tags like \script\, \style\, \title\, and \textarea\ as plaintext, and fixes several parsing edge cases. The new lexer enables support for advanced CSS selectors including \:has\, \:not\, \nth-child\, general sibling combinators, and case-insensitive attribute selectors.

src · high confidence

New native Elixir HTML tokenizer and numeric character reference handling

The library now includes a native Elixir implementation of the HTML tokenizer (lib/floki/html/tokenizer.ex) compliant with WHATWG/W3C specifications, replacing previous parsing approaches. This change introduces a dedicated module for numeric character references (lib/floki/html/numeric\_charref.ex) that correctly maps specific byte values to Unicode characters and handles edge cases such as negative numbers and invalid ranges, ensuring more accurate parsing of HTML entities.

lib/floki/html · high confidence

Refactored selector engine with dedicated modules for attributes, combinators, and pseudo-classes

The selector parsing and matching logic has been restructured into specialized modules: \AttributeSelector\ handles attribute matching (including case-insensitive flags), \Combinator\ manages relationship selectors (descendant, child, adjacent/general sibling), \Functional\ parses \nth-child\ style expressions, \PseudoClass\ implements pseudo-class matching (such as \:checked\, \:disabled\, \:root\, \:has\, and \:contains\), and \Parser\ orchestrates the token stream. This refactoring replaces the previous monolithic selector handling with a modular architecture that supports more complex CSS selectors and improves maintainability.

lib/floki/selector · high confidence

Removed application configuration file

The \config/config.exs\ file has been deleted from the project. This removes the previous location where application settings and dependency configurations were defined, meaning any configuration previously managed in this file is no longer present or active.

config · high confidence

Test coverage

Added HTML tokenizer test data and generation template; Added generated HTML tokenizer tests for named entities; Added test coverage for Floki's HTML tree, text extraction, and traversal modules; Added tests for the new selector parsing and tokenizing components; Expanded test suite and configurable parser execution.

Dependencies

Floki v0.38.4 release with updated dependencies and Elixir 1.15 requirement

This release updates the project to version 0.38.4 and raises the minimum Elixir requirement to \~\> 1.15. It upgrades several development dependencies, including ex\_doc to 0.40.2, credo to 1.7.18, dialyxir to 1.4.7, earmark to 1.4.48, benchee to 1.5.0, and jason to 1.4.5. Optional parser dependencies html5ever and fast\_html are also updated to 0.18.0 and 2.5.0 respectively. The package configuration is refined to exclude mix tasks from the published library, ensuring a cleaner installation for end users.

(dependencies) · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

This is the PUBLIC form of this artifact. Findings are listed in full, but the details of SECURITY findings — which rule fired, in which file, on which line, and how to fix it — are deliberately withheld, and any secret-scanner results are excluded entirely. Where detail is absent here it was REMOVED FOR PUBLICATION; it is not missing from the analysis. The complete artifact is available from the repository owner.

Score

  • CAI 66 → 71 (+4.6)
  • Rubric changed (rubric-2026.08.19 → rubric-2026.09.15) — scores are not directly comparable.

Lenses

  • Code Health 99 → 99 (-0.2)
  • Architecture 100 → 100 (+0.0)
  • Maturity 54 → 52 (-2.0)
  • Readiness 64 → 79 (+14.3)
  • Security 83 → 96 (+12.8)

Resolved (10)

  • Coverage not included — suite not readable by the collector
  • Dependency hygiene not measured — no supported dependency manifest was read
  • High: security finding (details withheld)
  • High: security finding (details withheld)
  • High: security finding (details withheld)
  • Medium CVE: EEF-[CVE redacted] (mix.lock)
  • No exposed public API
  • Off-boarding risk: anonymized user #1
  • Test reliability not included
  • TooManyMethods: Tokenizer (lib/floki/html/tokenizer.ex)

New (21)

  • Documentation: no installation or build instructions (README.md)
  • FixmeComment (lib/floki/html/tokenizer.ex)
  • FixmeComment (lib/floki/html/tokenizer.ex)
  • FixmeComment (lib/floki/html/tokenizer.ex)
  • High: security finding (details withheld)
  • High: security finding (details withheld)
  • High: security finding (details withheld)
  • Medium CVE: EEF-[CVE redacted] (mix.lock)
  • Off-boarding risk: anonymized user #1
  • Orphaned knowledge (src/floki_mochi_html.erl)
  • Outdated: benchee
  • Outdated: credo
  • Outdated: dialyxir
  • Outdated: ex_doc
  • Retired release: earmark
  • TodoComment (lib/floki/html/tokenizer.ex)
  • TodoComment (lib/floki/html/tokenizer.ex)
  • TodoComment (lib/floki/html/tokenizer.ex)
  • TodoComment (lib/floki/html/tokenizer.ex)
  • TodoComment (lib/floki/html_tree/comment.ex)
  • …and 1 more

Changes since last survey

  • 1 commits — 1 feature/other, 0 fixes

By area

  • lib/floki — 1 commit

Notable commits

  • change: Optimize functional selector matching (#705)

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

philss/floki was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 23 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit 22fa6f4ed93e00fe67dc824ebcc70646c2a14317 — the exact code this score is about.
  • Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-955b9cee9818.