Skip to content
CAI
Software that uses CAICheck a score

infinilabs/analysis-ik

62.4

Adequate · 24 September 2026

3.9k

lines of production code

Java

primary language

5

measurements over time

CAI band scale
CAI trend line
CAI lens gauges

What this system is

This system is a Chinese text analysis engine (IK Analyzer) that provides tokenization and segmentation capabilities for search platforms. It has been refactored into a multi-module architecture supporting both Elasticsearch and OpenSearch, while replacing the legacy implementation with a rewritten core engine that handles complex Chinese word segmentation, surrogate pairs, and ambiguous path resolution. The codebase now includes comprehensive test coverage and updated configuration management for these search integrations.

Features

Add IK Analysis plugin for Elasticsearch

Introduces the IK Chinese text analysis plugin for Elasticsearch, providing 'ik\_smart' and 'ik\_max\_word' tokenizers and analyzers. The change adds the plugin's Java implementation, configuration handling, and resource files (including entitlement and security policies) to enable Chinese word segmentation within Elasticsearch.

elasticsearch · high confidence

Add OpenSearch plugin integration for IK analysis

The OpenSearch module for the IK analysis plugin has been added, introducing the 'ik\_smart' and 'ik\_max\_word' tokenizers and analyzers. This change registers the plugin with OpenSearch, allowing users to perform Chinese text segmentation and analysis using the IK engine within OpenSearch indices.

opensearch · high confidence

Complete rewrite of the IK Analyzer core engine

The core segmentation engine has been completely rewritten from scratch, introducing a new \Configuration\ class to manage analyzer settings (such as smart segmentation and lowercase handling) and a context-driven architecture where multiple \ISegmenter\ implementations (including \CJKSegmenter\, \CN\_QuantifierSegmenter\, \LetterSegmenter\, and \SurrogatePairSegmenter\) process a shared \AnalyzeContext\. This refactoring adds support for surrogate pairs (e.g., rare characters and emojis), improves buffer management, and implements a new \IKArbitrator\ to handle ambiguous word paths with a fallback mechanism for complex cross-paths.

core/src/main · high confidence

Removals

Removal of legacy IK Analyzer implementation

The legacy IK Analyzer implementation has been removed from the codebase. This includes the deletion of the \es-plugin.properties\ configuration file, the \AnalysisIkPlugin\ entry point, the \IkAnalysisBinderProcessor\ for registering the analyzer, and all underlying Java classes in the \org.wltea.analyzer\ and \org.elasticsearch.index.analysis\ packages (such as \IkAnalyzer\, \IkTokenizer\, \IKSegmentation\, \Lexeme\, and various segmenters). This change eliminates the older analyzer implementation, likely to streamline the codebase or prepare for a new architecture.

src/main · high confidence

Behavioural changes

Removed custom dictionary files

The custom dictionary files 'mydict.dic' and 'sougou.dict' have been removed from the configuration. This means any custom or third-party word lists previously loaded from these files will no longer be available for input methods or text processing.

config/ik/custom · high confidence

Reorganized IK Analyzer configuration and dictionaries

The IK Analyzer configuration file (IKAnalyzer.cfg.xml) and the legacy elasticsearch.yml and logging.yml configuration files have been removed. The IK dictionary files (main.dic, stopword.dic, etc.) have been moved from the config/ik/ subdirectory to the root config/ directory. Additionally, new configuration files (extra\_main.dic, extra\_single\_word.dic, etc.) and a new config/IKAnalyzer.cfg.xml have been added to support extended dictionary and stopword configurations for the IK analyzer.

config · high confidence

Test coverage

Added comprehensive test coverage for the analyzer core; Removed obsolete IK Analyzer test classes and dictionary files.

Dependencies

Maven build system restructured into multi-module project

The project's build configuration has been refactored from a single-module setup into a multi-module Maven structure. The root \pom.xml\ now defines a parent POM with modules for \core\, \elasticsearch\, and \opensearch\, each with their own \pom.xml\ files. This change introduces separate build artifacts for the IK analyzer's core logic, its Elasticsearch plugin, and its OpenSearch plugin, allowing for independent versioning and dependency management for each target platform.

(dependencies) · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

This is the PUBLIC form of this artifact. Findings are listed in full, but the details of SECURITY findings — which rule fired, in which file, on which line, and how to fix it — are deliberately withheld, and any secret-scanner results are excluded entirely. Where detail is absent here it was REMOVED FOR PUBLICATION; it is not missing from the analysis. The complete artifact is available from the repository owner.

Score

  • CAI 60 → 62 (+2.2)
  • Rubric changed (rubric-2026.08.19 → rubric-2026.09.15) — scores are not directly comparable.

Lenses

  • Code Health 93 → 96 (+3.1)
  • Architecture 100 → 99 (-0.5)
  • Maturity 54 → 54 (-0.8)
  • Readiness 50 → 55 (+5.2)
  • Security 75 → 75 (+0.0)

Resolved (13)

  • Coverage not included — suite not readable by the collector
  • Dependency hygiene not measured — dependency manifest found but not parsed for hygiene
  • Duplicated block (11 lines × 2) (core/src/main/java/org/wltea/analyzer/dic/Dictionary.java)
  • Duplicated block (5 lines × 2) (core/src/main/java/org/wltea/analyzer/core/CJKSegmenter.java)
  • Duplicated block (6 lines × 2) (elasticsearch/src/main/java/com/infinilabs/ik/elasticsearch/ConfigurationSub.java)
  • Duplicated block (7 lines × 2) (core/src/main/java/org/wltea/analyzer/core/CharacterUtil.java)
  • Duplicated block (8 lines × 2) (core/src/main/java/org/wltea/analyzer/core/LexemePath.java)
  • Duplicated block (9 lines × 2) (core/src/main/java/org/wltea/analyzer/dic/Dictionary.java)
  • Duplicated block (9 lines × 2) (core/src/main/java/org/wltea/analyzer/dic/Dictionary.java)
  • No exposed public API
  • Off-boarding risk: anonymized user #1
  • Scanner failed to run — not a clean result
  • Test reliability not included

New (41)

  • Dependency hygiene PARTLY measured — Maven/Gradle declarations read, no dependency graph resolved
  • Documentation: no architecture or design documentation (README.md)
  • Duplicated block (12 lines × 2) (core/src/main/java/org/wltea/analyzer/dic/Dictionary.java)
  • Duplicated block (12 lines × 2) (core/src/main/java/org/wltea/analyzer/dic/Dictionary.java)
  • Duplicated block (13 lines × 2) (core/src/main/java/org/wltea/analyzer/dic/Dictionary.java)
  • Duplicated block (37 lines × 2) (elasticsearch/src/main/java/com/infinilabs/ik/elasticsearch/AnalysisIkPlugin.java)
  • Duplicated block (6 lines × 2) (elasticsearch/src/main/java/com/infinilabs/ik/elasticsearch/ConfigurationSub.java)
  • Duplicated block (6 lines × 2) (elasticsearch/src/main/java/com/infinilabs/ik/elasticsearch/ConfigurationSub.java)
  • Duplicated block (7–8 lines × 2) (core/src/main/java/org/wltea/analyzer/core/CharacterUtil.java)
  • Duplicated block (8 lines × 2) (core/src/main/java/org/wltea/analyzer/core/LexemePath.java)
  • High: security finding (details withheld)
  • High: security finding (details withheld)
  • High: security finding (details withheld)
  • High: security finding (details withheld)
  • No dependency advisory monitoring
  • No direct assertions: testAllStopwords (core/src/test/java/org/wltea/analyzer/lucene/Issue921Test.java)
  • No direct assertions: testMultiValueWithPartialStopword (core/src/test/java/org/wltea/analyzer/lucene/Issue921Test.java)
  • No direct assertions: testNoStopwordRegression (core/src/test/java/org/wltea/analyzer/lucene/Issue921Test.java)
  • No direct assertions: testOriginalIssueScenario (core/src/test/java/org/wltea/analyzer/lucene/Issue921Test.java)
  • No direct assertions: testStopwordFiltering (core/src/test/java/org/wltea/analyzer/lucene/Issue921Test.java)
  • …and 21 more

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

infinilabs/analysis-ik was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 24 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit 6d2d70fd1a237cbf75cde254e8e4d6319b81266c — the exact code this score is about.
  • Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-ae95d6cad036.