XiaoMi/MiNLP
38.8
Weak · 28 September 2026
37.5k
lines of production code
JavaScript
with Scala
2
measurements over time
What this system is
This system is a specialized natural language processing toolkit designed for Chinese text, comprising a Scala-based entity extraction engine and a Python-based word segmentation module. The core engine parses unstructured text to identify and normalize specific entities such as time, place, currency, and zodiac signs using rule-based and machine learning techniques. Complementing this, the tokenizer provides high-performance, multi-process word segmentation using deep learning models. Together, these components form a comprehensive infrastructure for Chinese NLP tasks, offering both API access and local processing capabilities.
Features
Added development, benchmarking, and release utility scripts
This change introduces a suite of shell scripts in the \duckling-fork-chinese/bin\ directory to support the Chinese fork's development workflow. Users can now use \bayes\, \numeral\, \time\, and \diff\ to run specific debugging or testing tasks via Bloop, while \time\_test.sh\, \time\_single\_case.sh\, and \time\_all\_case\_cost.sh\ provide targeted benchmarking and test execution. Operational convenience is improved with \stop\ to kill running server processes and \package\ to build and stage the server release. Finally, \release.sh\ automates the versioning, tagging, and git push process for new releases.
duckling-fork-chinese/bin · high confidence
Chinese NLP learning module adds sample data for multiple dimensions
The \duckling-fork-chinese/learning\ module now includes a comprehensive set of training examples for various data dimensions, including time, numerals, currency, temperature, distance, and more. This addition supports the Chinese language learning capabilities by providing structured input for model training and testing.
duckling-fork-chinese/learning · high confidence
Initial Chinese localization data and configuration
This commit introduces the core resource files for the Chinese Duckling fork, enabling Chinese language support for entity extraction. It adds \reference.conf\ to define the analyzer and available dimensions (such as Time, Place, and Constellation), \constellation.json\ to map Chinese zodiac signs to their aliases, \places4.json\ containing a comprehensive list of countries with Chinese names and codes, and \solar\_terms.csv\ providing data for Chinese solar terms. These resources allow the engine to recognize and parse Chinese-specific entities like locations, zodiac signs, and traditional calendar dates.
duckling-fork-chinese/core/src/main/resources · high confidence
Initial release of MiNLP-Tokenizer with multi-process support
The MiNLP-Tokenizer module is introduced as a new component, providing a Chinese word segmentation tool based on deep learning sequence labeling. This release includes the core tokenizer implementation, setup configuration, and documentation. A key feature is the addition of multi-process cut functionality, allowing users to accelerate tokenization speed by processing large volumes of text in parallel (controlled via the \n\_jobs\ parameter). The tool supports both coarse and fine-grained segmentation, custom user dictionaries, and is compatible with Python 3.5 through 3.8.
minlp-tokenizer · high confidence
Initial release of duckling-fork-chinese
This entry introduces the \duckling-fork-chinese\ project, a Scala-based reimplementation of Facebook's Duckling library specifically tailored for Chinese text parsing. The repository includes the core library code, a web server for API access, and documentation. It is currently set to version 1.4-SNAPSHOT and supports Scala 2.11, 2.12, and 2.13. The project provides APIs for analyzing dimensions like time and place, with specific support for Chinese contexts and holidays.
duckling-fork-chinese · high confidence
Initial server resource configuration and UI assets
The server now includes its foundational configuration and static assets. Application properties define the view resolution for JSP templates and set the listening port to 11559. Logging is configured via Logback to output to the console and rotate log files. Additionally, the static CSS for the Element UI component library is included to support the frontend interface.
duckling-fork-chinese/server/src/main/resources · high confidence
Introduce MiNLP Chinese tokenizer with multi-process support
The \minlp-tokenizer\ module is added, providing a Chinese word segmentation tool that uses a CNN-CRF model for fine and coarse granularity tokenization. It includes built-in lexicons (default and idioms) to influence segmentation, supports full-width to half-width character conversion, and allows users to define custom dictionaries. The tokenizer supports multi-process execution via the \n\_jobs\ parameter in the \cut\ method to improve performance on large text batches.
minlp-tokenizer/minlptokenizer · high confidence
Introduction of Chinese-specific NLP dimensions and rule engine
The library now includes a comprehensive set of new parsing dimensions tailored for Chinese text, including Act, Age, BloodType, Constellation, Currency, Duplicate, Episode, Gender, Level, and others. These additions enable the extraction of specific entity types such as anime episodes, blood types, zodiac signs, and Chinese currency formats, alongside foundational rule infrastructure to support these new linguistic patterns.
duckling-fork-chinese/core/src/main/scala/com/xiaomi/duckling/dimension · high confidence
Introduction of core Duckling parsing components for Chinese support
This change introduces the foundational Scala classes for the Duckling text parsing engine within the \com.xiaomi.duckling\ package. It adds \Api.scala\ as the primary entry point for parsing entities and analyzing text, \DuckParser.scala\ to handle the core parsing logic including rule application, ranking, and overlap filtering, and \Document.scala\ to manage input text, tokenization, and character classification (including Chinese character detection). Additionally, it includes \JsonSerde.scala\ for serializing parsed results, \Resources.scala\ for loading language data with explicit classloader support for Android compatibility, \Rules.scala\ for rule resolution, and \Types.scala\ defining core data structures like Context, Options, and Token. These files collectively enable the system to ingest text, apply linguistic rules, and return structured entity extractions.
duckling-fork-chinese/core/src/main/scala/com/xiaomi/duckling · high confidence
Introduction of core parsing infrastructure and Chinese-specific extraction utilities
This change introduces foundational components for the Chinese language processing module within the core library. It adds a constraint system (Constraint.scala and TokenSpan.scala) to filter out invalid parsing results that cross token boundaries, ensuring more accurate semantic ranges. A new PlaceExtractor utility is provided to simplify location extraction by returning a single, highest-confidence result while handling foreign place logic. Additionally, core type definitions are established, including LanguageInfo for managing sentence tokenization and dependency structures, Node for representing parse tree elements with validation capabilities, and LunarOutOfRangeException to handle specific date-related errors. These files collectively form the structural basis for Chinese-specific dimension processing.
(repo-wide) · high confidence
New Java-based Spring Boot server implementation
The server module now provides a Java-based REST API built on Spring Boot, replacing the previous implementation. This includes a new controller with endpoints for text extraction (/duckling/extract), general API queries (/api), and place extraction (/api/place), along with helper classes for JSON serialization and application startup.
duckling-fork-chinese/server/src/main/scala/com/xiaomi/duckling/server · high confidence
New place resolution and token visualization utilities
Added PlaceQuery to handle Chinese place entity resolution by filtering candidates (prioritizing non-China countries and selecting the best match based on range and path length) and TokenVisualization to generate HTML tables for debugging NLP tokens, including highlighting and grain details for time and duration values.
duckling-fork-chinese/server/src/main/scala/com/xiaomi/duckling · high confidence
Behavioural changes
New Chinese NLP engine implementation with Aho-Corasick lexicon lookup
This change introduces the core parsing engine for the Chinese language fork, replacing previous implementations with a new architecture. The engine now uses an Aho-Corasick automaton (via HanLP) for efficient lexicon matching, supports multi-character symbol detection (such as emojis), and implements phrase and regex lookups. It also includes variable-length character expansion logic and a node-limiting mechanism to control performance during parsing.
duckling-fork-chinese/core/src/main/scala/com/xiaomi/duckling/engine · high confidence
New Naive Bayes ranking system with Kryo serialization
The ranking module now uses a Naive Bayes classifier to score and rank parsing results, replacing the previous approach. This includes a new Bayes implementation for training and classification, a Kryo-based serialization layer (KryoSerde) for loading pre-trained models, and an OverlapResolver to handle conflicts between overlapping time and numeral interpretations. The system is exposed via a Ranker enum and integrated into the ranking pipeline through Rank.scala.
duckling-fork-chinese/core/src/main/scala/com/xiaomi/duckling/ranking · high confidence
Switches Chinese text segmentation to BERT-based tokenizer
The Chinese analyzer now uses a BERT-based tokenizer (via \com.robrua.nlp.bert.BasicTokenizer\) for word segmentation instead of the previous HanLP-based approach. This change is implemented in the new \SplitAnalyzer\ class within the analyzer package, which replaces the prior default segmentation logic to improve tokenization accuracy for Chinese text.
duckling-fork-chinese/core/src/main/scala/com/xiaomi/duckling/analyzer · high confidence
Test coverage
Added JMH benchmarks for Chinese date/time and number parsing; Added base unit test specification class; Added comprehensive test suite for Chinese NLP dimensions; Added test for Chinese place query extraction; Added unit tests for MiNLP tokenizer and vocabulary components.
Dependencies
Initial SBT build configuration and dependency definitions
The project now uses SBT (version 1.11.5) for building, with a centralized dependency management file defining core libraries such as Spring Boot 2.4.5, Kryo 5.3.0, and Apache Commons Text 1.10.0, alongside build plugins for native packaging, release management, and code coverage.
duckling-fork-chinese/project · high confidence
Initial dependency manifests for duckling-fork-chinese and minlp-tokenizer
Added build configuration and dependency manifests for two new modules. The \duckling-fork-chinese\ module (Scala) is configured with sbt, targeting Scala versions 2.11.12, 2.12.12, and 2.13.10, and includes sub-projects for core, learning, test, server, and benchmark functionality. The \minlp-tokenizer\ module (Python) specifies dependencies on tensorflow (\>=1.14), pyahocorasick, and regex.
(dependencies) · high confidence
Housekeeping
Initial documentation for the Chinese Duckling fork
This change introduces the initial documentation set for the duckling-fork-chinese project, including a CONTRIBUTING guide for developers, an FAQ addressing usage constraints and dependency optimization, a ROADMAP outlining future enhancements like NER integration and performance improvements, a class diagram illustrating the core architecture, and a detailed reference of supported dimensions (such as Time, Numeral, Act, Age, Currency, and others) with parsing examples and configuration options.
duckling-fork-chinese/doc · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
This is the PUBLIC form of this artifact. Findings are listed in full, but the details of SECURITY findings — which rule fired, in which file, on which line, and how to fix it — are deliberately withheld, and any secret-scanner results are excluded entirely. Where detail is absent here it was REMOVED FOR PUBLICATION; it is not missing from the analysis. The complete artifact is available from the repository owner.
Score
- CAI 49 → 39 (-10.3)
- Rubric changed (rubric-2026.09.8 → rubric-2026.09.16) — scores are not directly comparable.
Lenses
- Code Health 92 → 44 (-48.0)
- Architecture 98 → 87 (-10.9)
- Maturity 50 → 50 (-0.0)
- Readiness 46 → 35 (-10.6)
- Security 69 → 78 (+9.0)
- Accessibility 41 → 30 (-11.2)
- Performance 100 (new)
Resolved (3)
- High: security finding (details withheld)
- High: security finding (details withheld)
- Off-boarding risk: anonymized user #1
New (69)
- Dependency hygiene PARTLY measured — Python dependencies read, no exact pin to grade for currency
- FileTooLong: js/vue.js (duckling-fork-chinese/server/src/main/resources/static/js/vue.js)
- No ADRs found
- Off-boarding risk: anonymized user #1
- Split duckling-fork-chinese
- bfwone.loadjs (cognitive 37) (duckling-fork-chinese/server/src/main/resources/static/js/bfwone.js)
- bfwone.loadjs (cyclomatic 25) (duckling-fork-chinese/server/src/main/resources/static/js/bfwone.js)
- vue.Vue.$mount (cognitive 31) (duckling-fork-chinese/server/src/main/resources/static/js/vue.js)
- vue.Vue.$mount (cyclomatic 16) (duckling-fork-chinese/server/src/main/resources/static/js/vue.js)
- vue._createElement (cognitive 29) (duckling-fork-chinese/server/src/main/resources/static/js/vue.js)
- vue._createElement (cyclomatic 25) (duckling-fork-chinese/server/src/main/resources/static/js/vue.js)
- vue._update (cognitive 27) (duckling-fork-chinese/server/src/main/resources/static/js/vue.js)
- vue.actuallySetSelected (cognitive 17) (duckling-fork-chinese/server/src/main/resources/static/js/vue.js)
- vue.addHandler (cognitive 24) (duckling-fork-chinese/server/src/main/resources/static/js/vue.js)
- vue.addHandler (cyclomatic 22) (duckling-fork-chinese/server/src/main/resources/static/js/vue.js)
- vue.assertProp (cognitive 16) (duckling-fork-chinese/server/src/main/resources/static/js/vue.js)
- vue.bindObjectProps (cognitive 28) (duckling-fork-chinese/server/src/main/resources/static/js/vue.js)
- vue.bindObjectProps (cyclomatic 17) (duckling-fork-chinese/server/src/main/resources/static/js/vue.js)
- vue.checkNode (cognitive 23) (duckling-fork-chinese/server/src/main/resources/static/js/vue.js)
- vue.createCompileToFunctionFn (cognitive 19) (duckling-fork-chinese/server/src/main/resources/static/js/vue.js)
- …and 49 more
Architecture
- Unchanged — 0 containers · 1 contexts · 0 edges
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
XiaoMi/MiNLP was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 28 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 3745b4a8f9a46c85d9ac1f732779b0de332cb7ee — the exact code this score is about.
- Scored under rubric-2026.09.16 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-2d9048c36d26.