infinilabs/analysis-ik
62.4
Adequate · 24 September 2026
3.9k
lines of production code
Java
primary language
5
measurements over time
What this system is
This system is a Chinese text analysis engine (IK Analyzer) that provides tokenization and segmentation capabilities for search platforms. It has been refactored into a multi-module architecture supporting both Elasticsearch and OpenSearch, while replacing the legacy implementation with a rewritten core engine that handles complex Chinese word segmentation, surrogate pairs, and ambiguous path resolution. The codebase now includes comprehensive test coverage and updated configuration management for these search integrations.
Features
Add IK Analysis plugin for Elasticsearch
Introduces the IK Chinese text analysis plugin for Elasticsearch, providing 'ik\_smart' and 'ik\_max\_word' tokenizers and analyzers. The change adds the plugin's Java implementation, configuration handling, and resource files (including entitlement and security policies) to enable Chinese word segmentation within Elasticsearch.
elasticsearch · high confidence
Add OpenSearch plugin integration for IK analysis
The OpenSearch module for the IK analysis plugin has been added, introducing the 'ik\_smart' and 'ik\_max\_word' tokenizers and analyzers. This change registers the plugin with OpenSearch, allowing users to perform Chinese text segmentation and analysis using the IK engine within OpenSearch indices.
opensearch · high confidence
Complete rewrite of the IK Analyzer core engine
The core segmentation engine has been completely rewritten from scratch, introducing a new \Configuration\ class to manage analyzer settings (such as smart segmentation and lowercase handling) and a context-driven architecture where multiple \ISegmenter\ implementations (including \CJKSegmenter\, \CN\_QuantifierSegmenter\, \LetterSegmenter\, and \SurrogatePairSegmenter\) process a shared \AnalyzeContext\. This refactoring adds support for surrogate pairs (e.g., rare characters and emojis), improves buffer management, and implements a new \IKArbitrator\ to handle ambiguous word paths with a fallback mechanism for complex cross-paths.
core/src/main · high confidence
Removals
Removal of legacy IK Analyzer implementation
The legacy IK Analyzer implementation has been removed from the codebase. This includes the deletion of the \es-plugin.properties\ configuration file, the \AnalysisIkPlugin\ entry point, the \IkAnalysisBinderProcessor\ for registering the analyzer, and all underlying Java classes in the \org.wltea.analyzer\ and \org.elasticsearch.index.analysis\ packages (such as \IkAnalyzer\, \IkTokenizer\, \IKSegmentation\, \Lexeme\, and various segmenters). This change eliminates the older analyzer implementation, likely to streamline the codebase or prepare for a new architecture.
src/main · high confidence
Behavioural changes
Removed custom dictionary files
The custom dictionary files 'mydict.dic' and 'sougou.dict' have been removed from the configuration. This means any custom or third-party word lists previously loaded from these files will no longer be available for input methods or text processing.
config/ik/custom · high confidence
Reorganized IK Analyzer configuration and dictionaries
The IK Analyzer configuration file (IKAnalyzer.cfg.xml) and the legacy elasticsearch.yml and logging.yml configuration files have been removed. The IK dictionary files (main.dic, stopword.dic, etc.) have been moved from the config/ik/ subdirectory to the root config/ directory. Additionally, new configuration files (extra\_main.dic, extra\_single\_word.dic, etc.) and a new config/IKAnalyzer.cfg.xml have been added to support extended dictionary and stopword configurations for the IK analyzer.
config · high confidence
Test coverage
Added comprehensive test coverage for the analyzer core; Removed obsolete IK Analyzer test classes and dictionary files.
Dependencies
Maven build system restructured into multi-module project
The project's build configuration has been refactored from a single-module setup into a multi-module Maven structure. The root \pom.xml\ now defines a parent POM with modules for \core\, \elasticsearch\, and \opensearch\, each with their own \pom.xml\ files. This change introduces separate build artifacts for the IK analyzer's core logic, its Elasticsearch plugin, and its OpenSearch plugin, allowing for independent versioning and dependency management for each target platform.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
This is the PUBLIC form of this artifact. Findings are listed in full, but the details of SECURITY findings — which rule fired, in which file, on which line, and how to fix it — are deliberately withheld, and any secret-scanner results are excluded entirely. Where detail is absent here it was REMOVED FOR PUBLICATION; it is not missing from the analysis. The complete artifact is available from the repository owner.
Score
- CAI 60 → 62 (+2.2)
- Rubric changed (rubric-2026.08.19 → rubric-2026.09.15) — scores are not directly comparable.
Lenses
- Code Health 93 → 96 (+3.1)
- Architecture 100 → 99 (-0.5)
- Maturity 54 → 54 (-0.8)
- Readiness 50 → 55 (+5.2)
- Security 75 → 75 (+0.0)
Resolved (13)
- Coverage not included — suite not readable by the collector
- Dependency hygiene not measured — dependency manifest found but not parsed for hygiene
- Duplicated block (11 lines × 2) (core/src/main/java/org/wltea/analyzer/dic/Dictionary.java)
- Duplicated block (5 lines × 2) (core/src/main/java/org/wltea/analyzer/core/CJKSegmenter.java)
- Duplicated block (6 lines × 2) (elasticsearch/src/main/java/com/infinilabs/ik/elasticsearch/ConfigurationSub.java)
- Duplicated block (7 lines × 2) (core/src/main/java/org/wltea/analyzer/core/CharacterUtil.java)
- Duplicated block (8 lines × 2) (core/src/main/java/org/wltea/analyzer/core/LexemePath.java)
- Duplicated block (9 lines × 2) (core/src/main/java/org/wltea/analyzer/dic/Dictionary.java)
- Duplicated block (9 lines × 2) (core/src/main/java/org/wltea/analyzer/dic/Dictionary.java)
- No exposed public API
- Off-boarding risk: anonymized user #1
- Scanner failed to run — not a clean result
- Test reliability not included
New (41)
- Dependency hygiene PARTLY measured — Maven/Gradle declarations read, no dependency graph resolved
- Documentation: no architecture or design documentation (README.md)
- Duplicated block (12 lines × 2) (core/src/main/java/org/wltea/analyzer/dic/Dictionary.java)
- Duplicated block (12 lines × 2) (core/src/main/java/org/wltea/analyzer/dic/Dictionary.java)
- Duplicated block (13 lines × 2) (core/src/main/java/org/wltea/analyzer/dic/Dictionary.java)
- Duplicated block (37 lines × 2) (elasticsearch/src/main/java/com/infinilabs/ik/elasticsearch/AnalysisIkPlugin.java)
- Duplicated block (6 lines × 2) (elasticsearch/src/main/java/com/infinilabs/ik/elasticsearch/ConfigurationSub.java)
- Duplicated block (6 lines × 2) (elasticsearch/src/main/java/com/infinilabs/ik/elasticsearch/ConfigurationSub.java)
- Duplicated block (7–8 lines × 2) (core/src/main/java/org/wltea/analyzer/core/CharacterUtil.java)
- Duplicated block (8 lines × 2) (core/src/main/java/org/wltea/analyzer/core/LexemePath.java)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- No dependency advisory monitoring
- No direct assertions: testAllStopwords (core/src/test/java/org/wltea/analyzer/lucene/Issue921Test.java)
- No direct assertions: testMultiValueWithPartialStopword (core/src/test/java/org/wltea/analyzer/lucene/Issue921Test.java)
- No direct assertions: testNoStopwordRegression (core/src/test/java/org/wltea/analyzer/lucene/Issue921Test.java)
- No direct assertions: testOriginalIssueScenario (core/src/test/java/org/wltea/analyzer/lucene/Issue921Test.java)
- No direct assertions: testStopwordFiltering (core/src/test/java/org/wltea/analyzer/lucene/Issue921Test.java)
- …and 21 more
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
infinilabs/analysis-ik was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 24 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 6d2d70fd1a237cbf75cde254e8e4d6319b81266c — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-ae95d6cad036.