Skip to content
CAI
Software that uses CAICheck a score

hankcs/HanLP

53.7

Adequate · 2 August 2026

50.3k

lines of production code

Python

primary language

3

measurements over time

CAI band scale
CAI trend line
CAI lens gauges

What this system is

HanLP is a comprehensive natural language processing library that provides a modular, pipeline-based architecture for building and training NLP models. It supports a wide range of tasks including tokenization, part-of-speech tagging, named entity recognition, dependency and constituency parsing, semantic role labeling, and text classification. The system offers both PyTorch and TensorFlow backends, enabling users to construct complex workflows by chaining components, and includes utilities for model compression, multi-task learning, and evaluation metrics.

How it got here

2014–2020 — Modular architecture and backend expansion

28 changes.

The project underwent a significant architectural shift towards a modular, component-based design, introducing a unified pipeline system and abstract base classes for NLP components. This period also focused on expanding framework support by adding comprehensive TensorFlow implementations for existing PyTorch models, alongside building out a rich ecosystem of layers, metrics, and dataset loaders.

2021–2025 — Multi-task learning and component expansion

27 changes.

This period was defined by the introduction of a comprehensive Multi-Task Learning (MTL) framework, enabling joint training across diverse NLP tasks such as parsing, tagging, and NER. The codebase expanded significantly with new components for knowledge distillation, semantic textual similarity, and masked language modeling, alongside extensive demo scripts for multiple languages.

Features

Add Biaffine-based NER component with transformer support

A new Named Entity Recognition component based on the Biaffine architecture has been added to HanLP. This implementation treats every possible span as a candidate entity and predicts its label, supporting both flat and nested NER tasks. The component integrates with the existing TorchComponent framework, allowing for optional transformer-based backends and standard PyTorch training loops with F1 metric tracking.

_hanlp/components/ner/biaffine\ner · high confidence

Add CRF layer and file-reading utility

The \hanlp/layers/crf\ module now includes a PyTorch-based Conditional Random Field (CRF) layer (\crf.py\) and a TensorFlow-based counterpart (\crf\_tf.py\), enabling sequence labeling tasks in both frameworks. Additionally, the \hanlp/utils/file\_read\_backwards\ package has been added, providing a \FileReadBackwards\ utility for reading files from the end to the beginning, which supports multiple encodings and chunked reading.

_hanlp/layers/crf, hanlp/utils/file\_read\backwards · high confidence

Add Chinese NLP demo and training scripts for TensorFlow

Added a new directory for Chinese language processing demos and training scripts using TensorFlow. This includes demo scripts for tasks such as word segmentation, part-of-speech tagging, named entity recognition, dependency parsing, and sentiment analysis, as well as training scripts for various models including BERT, ALBERT, and Electra on datasets like MSRA, CTB, and PTB.

_plugins/hanlp\_demo/hanlp\demo/zh/tf · high confidence

Add Chinese NLP training scripts and Windows compatibility check

Added new scripts for fine-tuning Named Entity Recognition (NER) and training multi-task learning models for Chinese NLP tasks (tokenization, POS, NER, SRL, dependency parsing, etc.) in the HanLP demo plugin. The \finetune\_ner.py\ script demonstrates how to load a pretrained model and fine-tune it on a custom dataset. The \open\_base.py\ and \open\_small.py\ scripts provide complete training pipelines for multiple Chinese NLP tasks using the Electra base and small models. Additionally, a \block\_windows.py\ module was added to prevent execution on Windows, as the training scripts are not supported on that platform.

_plugins/hanlp\_demo/hanlp\demo/zh/train · high confidence

Add Japanese multi-task learning demo

A new demo script for Japanese NLP has been added, showcasing a multi-task learning model trained on NPCMJ, UD, and Kyoto corpora. The script loads the pre-trained Japanese model (tok, pos, con, BERT base char) and demonstrates tokenization, part-of-speech tagging, and dependency parsing on sample sentences.

_plugins/hanlp\_demo/hanlp\demo/ja · high confidence

Add Masked Language Model component for text completion

A new Masked Language Model (MLM) component has been introduced to support filling blanks in text. The implementation includes a dataset class and a TorchComponent that leverages Hugging Face's AutoModelForMaskedLM. The model's prediction logic explicitly blocks special tokens from the output, ensuring that only valid vocabulary tokens are returned during inference.

hanlp/components/lm · high confidence

Add STS component for semantic textual similarity

A new STS (Semantic Textual Similarity) component has been added to the library. This introduces a transformer-based baseline model that fine-tunes a pre-trained transformer with a regression head to predict similarity scores between sentence pairs. The implementation includes a data loader for the SemanticTextualSimilarityDataset, a training loop with Spearman correlation metrics, and a predict method that supports both flat and nested list inputs for sentence pairs.

hanlp/components/sts · high confidence

Add Semantic Role Labeling tasks to the multi-task learning framework

Introduced two new Semantic Role Labeling (SRL) tasks for the multi-task learning (MTL) framework: a span-based BIO tagger and a span-ranking model. These additions enable joint training of SRL alongside other NLP tasks within HanLP's unified pipeline.

hanlp/components/mtl/tasks/srl · high confidence

Add TensorFlow-based Biaffine Dependency Parser

The \hanlp/components/parsers\ directory now includes a complete TensorFlow implementation of the Biaffine Dependency Parser (\BiaffineDependencyParserTF\), mirroring the existing PyTorch version. This adds support for training and inference using TensorFlow/Keras, including algorithmic components like K-Means clustering, Eisner's algorithm, and Chu-Liu/Edmonds decoding adapted for TF, enabling users to leverage the parser with TensorFlow backends.

hanlp/components/parsers · high confidence

Add TensorFlow-based NLP components for parsing, tagging, and CRF

Added new TensorFlow implementations for sequence labeling and parsing tasks. This includes a Conditional Random Field (CRF) layer and loss function for sequence classification, a CNN-based tagger for word-level tagging, an n-gram convolutional tagger for character/word-level tagging, and a Transformer-based tagger for sequence tagging. Additionally, a TensorFlow implementation of the Biaffine dependency parser was added. These components provide native TensorFlow support for these NLP tasks, mirroring existing PyTorch/MXNet counterparts.

python · high confidence

Add chunking and sequence-labeling metrics (SpanF1, ChunkingF1, BMES/IOBES F1)

The hanlp/metrics/chunking package now includes new metric classes for evaluating chunking and sequence-labeling tasks. This includes SpanF1 and ChunkingF1 for standard Python/tensor-based evaluation, as well as TensorFlow-specific implementations (BMES\_F1\_TF, IOBES\_F1\_TF) that integrate with Keras training loops. Users can now compute F1 scores for named entity recognition, part-of-speech tagging, or any sequence-labeling task using these built-in metrics.

hanlp/metrics/chunking · high confidence

Add parsing evaluation metrics and evaluation scripts

The \hanlp/metrics/parsing\ module now includes new metric classes for evaluating parsing performance: \AttachmentScore\ (computing UAS/LAS), \LabeledF1\ (computing UF/LF F1 scores), \LabeledF1TF\ (TensorFlow-based Labeled F1), \LabeledScore\ (TensorFlow-based UAS/LAS), and \SpanMetric\ (computing UCM/LCM and related precision/recall/F1 for span-based parsing). Additionally, evaluation scripts for CoNLL-X and semantic dependencies have been added to support standard benchmark evaluation.

hanlp/metrics/parsing · high confidence

Added AdamW optimizer with weight decay support

Introduced a new \AdamWeightDecay\ optimizer and a \WarmUp\ learning rate schedule in \hanlp/optimizers/adamw/\. This adds support for L2 weight decay that is decoupled from the Adam moment estimates, along with a polynomial warmup schedule for training stability.

hanlp/optimizers · high confidence

Added CRF-based constituency parser with MBR decoding

Users can now use a new constituency parser based on a Conditional Random Field (CRF) model, which supports Minimum Bayes-Risk (MBR) decoding for improved parsing accuracy. This component, located in the constituency parser module, introduces a two-stage neural CRF architecture that computes span and label scores, enabling more robust syntactic analysis.

hanlp/components/parsers/constituency · high confidence

Added custom loss functions for sparse categorical cross-entropy

The hanlp/losses module now includes three new loss functions: SparseCategoricalCrossentropyOverNonzeroWeights, SparseCategoricalCrossentropyOverBatchFirstDim, and MaskedSparseCategoricalCrossentropyOverBatchFirstDim. These provide specialized handling for sample weighting and masking in classification tasks, addressing previous compatibility issues with TensorFlow 2.3 and resolving errors during model export.

hanlp/losses · high confidence

Added demo scripts for Ancient Chinese (LZH) models

New demo scripts (demo\_mtl.py, demo\_tok.py) were added to the lzh directory, providing examples for loading and running the KYOTO\_EVAHAN\_TOK\_LZH and KYOTO\_EVAHAN\_TOK\_LEM\_POS\_UDEP\_LZH models on Ancient Chinese text.

_plugins/hanlp\_demo/hanlp\demo/lzh · high confidence

Added language-specific utility modules for English, Japanese, and Chinese

New utility modules have been added to support specific languages. For English, a regex-based word tokenizer is now available to handle contractions, possessives, and line breaks. For Japanese, a custom BERT Japanese tokenizer wrapper is provided to handle tokenization and offset mapping. For Chinese, a character normalization table and localization mappings for tasks, POS tags, NER, and dependency parsing are now available. These utilities provide foundational tools for processing text in these languages.

hanlp/utils/lang · high confidence

Added new components for knowledge distillation, sentence boundary detection, and semantic role labeling

The HanLP library now includes new modules for model compression and NLP tasks. The \hanlp.components.distillation\ package introduces a \DistillableComponent\ and associated loss functions (e.g., \kd\_ce\_loss\, \hid\_mse\_loss\) and temperature schedulers (\flsw\, \cwsm\) to support knowledge distillation training. Additionally, new components have been added for sentence boundary detection (\hanlp.components.eos\), semantic role labeling (\hanlp.components.srl\), and RNN-based tagging (\hanlp.components.taggers.rnn\), along with corresponding metrics for evaluation.

(repo-wide) · high confidence

English NLP demo scripts added

Added demo scripts for English language processing, including tokenization, part-of-speech tagging, dependency parsing, semantic dependency parsing, named entity recognition, language modeling, sentiment analysis, and an AMR (Abstract Meaning Representation) parser. A training script for the Stanford Sentiment Treebank (SST2) with ALBERT is also included.

_plugins/hanlp\_demo/hanlp\demo/en · high confidence

HanLP v2.1.3 release with AMR metrics and load/pipeline APIs

This update releases HanLP version 2.1.3, introducing a new \load()\ function in the main \hanlp\ namespace to simplify loading pretrained components and a \pipeline()\ helper for bundling components. It also adds support for Abstract Meaning Representation (AMR) evaluation via a new \smatch\_eval\ metric, and includes various bug fixes and compatibility improvements for TensorFlow, PyTorch, and macOS M1 chips.

hanlp · high confidence

Introduce HanLP RESTful API client with new NLP capabilities

Adds the \hanlp\_restful\ package, providing a Python client (\HanLPClient\) to interact with HanLP's RESTful APIs. This enables users to perform various NLP tasks including text style transfer, abstract meaning representation, keyphrase extraction, extractive summarization, and coreference resolution. The client supports specifying the language (e.g., 'ja' for Japanese, 'mul' for multilingual) and allows for coarse tokenization. It also includes a \verify\ parameter to control SSL certificate verification, supporting environments behind private TLS gateways. Tests are added to verify client functionality.

_plugins/hanlp\restful · high confidence

Introduce HanLP Trie plugin with tokenization and batch processing

Added the \plugins/hanlp\_trie\ module, providing a \Trie\ data structure and \TrieDict\/\TupleTrieDict\ classes for fast, rule-based tokenization and longest-prefix matching. The implementation supports batch processing via \split\_batch\ and \merge\_batch\ methods, enabling efficient integration of custom dictionaries into NLP pipelines. Tests verify core trie operations and dictionary tokenization behavior.

_plugins/hanlp\trie · high confidence

Introduce HanLP common utilities and visualization

Added a new \hanlp\_common\ plugin providing shared utilities, data structures, and visualization tools. This includes a \Document\ class for handling parsed annotations, \CoNLL\ and \CoNLLU\ word/sentence classes for parsing formats, and enhanced visualization for constituency trees, dependency graphs, and NER/SRL/CON outputs. The update also adds support for HTML visualization in Jupyter notebooks, improves the robustness of various visualizations, and allows users to disable IPython-specific features.

_plugins/hanlp\_common/hanlp\common · medium confidence

Introduce Java RESTful client with new NLP task methods

Added a new Java RESTful client (\HanLPClient\) and associated data models (\BaseInput\, \DocumentInput\, \SentenceInput\, \TokenInput\, \Span\, \CoreferenceResolutionOutput\, and MRP classes) that expose methods for tokenization, parsing, coreference resolution, text style transfer, semantic textual similarity, extractive/abstractive summarization, grammatical error correction, language identification, sentiment analysis, and abstract meaning representation. The client supports specifying tasks to run or skip, handles authentication via an \auth\ parameter or \HANLP\_AUTH\ environment variable, and includes tests for these new capabilities.

_plugins/hanlp\_restful\java · high confidence

Introduce Multi-Task Learning (MTL) framework

Users can now train multiple NLP tasks jointly using a shared encoder and task-specific decoders. This new MTL component supports dynamic task sampling during training, allowing for more efficient multi-task optimization.

hanlp/components/mtl · high confidence

Introduce TensorFlow-based data transformation and tokenization components

Added a new \hanlp/transform\ module containing a suite of data transformation classes for the TensorFlow backend. This includes \TransformerSequenceTokenizer\ and \TransformerTextTokenizer\ for handling Hugging Face tokenizers with support for sliding windows, prefix matching, and subtoken span tracking. The module also introduces \TableTransform\ and \TsvTaggingTransform\ for processing tabular and TSV datasets, as well as \CoNLLTransform\ for parsing CoNLL-style NLP tasks. Additionally, a \PrependSpace\ transform is provided to handle whitespace normalization for tokenizers that require it.

hanlp/transform · high confidence

Introduce Transformer-based tagger with PyTorch and TensorFlow backends

Users can now use a new TransformerTagger component that supports both PyTorch and TensorFlow backends. The PyTorch implementation (transformer\_tagger.py) provides a shallow tagging model using a linear layer with an optional CRF layer, featuring support for extra embeddings, teacher-student distillation, and dynamic vocabulary resizing. The TensorFlow implementation (transformer\_transform\_tf.py) includes a custom accuracy metric that handles masking and supports variable sequence lengths. This change enables training and inference for sequence labeling tasks like NER and POS tagging using modern transformer architectures.

hanlp/components/taggers/transformers · high confidence

Introduce a new Transformer encoder layer with sliding window and mirror support

The \hanlp/layers/transformers\ package now includes a new \TransformerEncoder\ layer that supports sliding window processing for long sequences, word dropout, and scalar mix aggregation. The implementation adds a \resource.py\ module that maps model identifiers to local mirror paths, enabling faster access to pre-trained models. Additionally, the \pt\_imports.py\ and \tf\_imports.py\ modules provide wrapper classes for Hugging Face and TensorFlow transformers, while \relative\_transformer.py\ introduces relative positional embeddings for sequence modeling.

hanlp/layers/transformers · high confidence

Introduce common component and dataset base classes

The \hanlp/common\ package now provides foundational classes for building NLP components. \Component\ serves as the abstract base for all components, while \TorchComponent\ and \KerasComponent\ offer PyTorch and TensorFlow-specific implementations with shared workflows for building models, handling vocabs, and managing configuration. Additionally, \TransformableDataset\ and \Transform\ classes enable flexible data transformation pipelines for both PyTorch and TensorFlow backends.

hanlp/common · high confidence

Introduce span-based Semantic Role Labeling (SRL) component

Added a new \span\_rank\ component for joint predicate and argument prediction in Semantic Role Labeling. This implementation, adapted from an external research repository, introduces a span-ranking architecture that generates candidate spans and ranks them using a bidirectional LSTM with highway connections, biaffine scoring, and greedy/DP-based decoding utilities. The addition includes the core model (\SpanRankingSRLModel\), the network layers (\highway\_variational\_lstm.py\, \layer.py\), and inference/evaluation utilities (\inference\_utils.py\, \srl\_eval\_utils.py\), enabling users to perform SRL tasks using this specific ranking-based approach.

_hanlp/components/srl/span\_bio, hanlp/components/srl/span\rank · high confidence

Introduces Universal Dependencies parser with lemma rule-based lemmatization

The \hanlp/components/parsers/ud\ directory now contains a complete implementation for Universal Dependencies parsing, including the \UniversalDependenciesParser\ and its associated model, decoder, and utility modules. This adds support for POS tagging, dependency parsing, feature prediction, and lemmatization. A key feature is the use of edit-distance-based lemma rules (\lemma\_edit.py\) to handle lemmatization, which allows the model to generate lemmas by applying transformation rules rather than relying on external dependencies like \alnlp\. The implementation supports both adaptive and standard loss functions for tag prediction and integrates with the existing \ContextualWordEmbedding\ encoder.

hanlp/components/parsers/ud · high confidence

Introduction of HanLP Common library

A new 'hanlp\_common' package has been added to the repository, providing shared utilities and structures for the HanLP NLP library. This includes the package's setup configuration, which defines the 'hanlp\_common' distribution (version 0.0.23) and its dependencies, alongside a README file that outlines the library's purpose and installation instructions.

_plugins/hanlp\common · high confidence

New BART-based AMR parser and generator components

The AMR component now includes a new BART-based architecture for both AMR parsing (text to AMR graph) and AMR generation (AMR graph to text). This introduces \BART\_AMR\_Parser\ and \BART\_AMR\_Generation\ classes, along with supporting modules for tokenization, dataset handling, and post-processing, enabling state-of-the-art performance on English AMR tasks.

hanlp/components/amr · high confidence

New Biaffine parser components for dependency and semantic parsing

The \hanlp/components/parsers/biaffine\ directory now contains the full implementation of the Biaffine parser, including the core \Biaffine\ layer, \MLP\ modules, and parser classes for standard dependency parsing (\BiaffineDependencyParser\), semantic dependency parsing (\BiaffineSemanticDependencyParser\), and secondary/structural attention parsing. This adds support for joint and separate decoding of arc and relation scores, with options for single-root constraints and cycle avoidance in the parsing logic.

hanlp/components/parsers/biaffine · high confidence

New Chinese NLP demos and utilities for HanLP v2.1

Added a suite of new demo scripts and utilities for Chinese NLP tasks in the \plugins/hanlp\_demo/hanlp\_demo/zh\ directory. These include demonstrations for AMR parsing, custom dictionaries for tokenization and POS tagging, task removal from multi-task learning models, document handling, masked language modeling, constituency parsing pipelines, semantic textual similarity, word2vec similarity, and training scripts for state-of-the-art Chinese word segmentation models.

_plugins/hanlp\_demo/hanlp\demo/zh · high confidence

New NER components for RNN and Transformer backends

Added new named entity recognition (NER) components for both RNN and Transformer architectures. The RNN implementation (\rnn\_ner.py\) supports word2vec/fasttext embeddings and handles IOBES/BIO/BIOUL tagging schemes. The Transformer implementation (\transformer\_ner.py\) integrates with pre-trained models and supports dictionary-based longest-prefix matching for entity recognition, along with blacklist filtering and entity type merging. A new TensorFlow-based NER component (\ner\_tf.py\) provides equivalent functionality using Keras/TF backends.

hanlp/components/ner · high confidence

New NER task implementations with dictionary support

Added two new Named Entity Recognition (NER) task implementations: a Biaffine-based model and a Tagging-based model (with optional CRF). The Tagging model introduces support for conditional-matching custom dictionaries (whitelist/blacklist) to override or filter entity predictions, and allows merging of consecutive entities of the same type.

hanlp/components/mtl/tasks/ner · high confidence

New character-level and contextual embedding layers for PyTorch and TensorFlow

The \hanlp/layers/embeddings\ module now includes new embedding implementations: \CharCNN\ and \CharRNN\ for character-level embeddings, \ContextualStringEmbedding\ for loading pre-trained language models, \ContextualWordEmbedding\ for transformer-based contextual embeddings, and \FastText\/\Word2Vec\ support. These are available for both PyTorch and TensorFlow backends, enabling users to integrate character-level and contextualized word embeddings into their models.

hanlp/layers/embeddings · high confidence

New classifier components for Hugging Face models and regression tasks

Added new components to support loading Hugging Face models directly for both classification and regression tasks. The \TransformerClassifierHF\ class enables inference using \AutoModelForSequenceClassification\ from the \transformers\ library, while \TransformerRegressionHF\ provides a similar interface for regression tasks. Additionally, a \FastTextClassifier\ component was introduced to support classification using the FastText library. These additions expand the range of pre-trained models and algorithms available for text classification and regression within the library.

hanlp/components/classifiers · high confidence

New dataset loaders and dataset classes for multiple NLP tasks

The \hanlp/datasets\ module has been reorganized and expanded with new dataset loaders and classes for various NLP tasks. This includes a \LanguageModelDataset\ for language modeling, a \SentenceBoundaryDetectionDataset\ for end-of-sentence detection, and loaders for coreference resolution, named entity recognition (NER), and parsing tasks (including CoNLL formats and Chinese Treebank variants). These changes provide standardized interfaces for loading and processing diverse linguistic datasets within the library.

hanlp/datasets · high confidence

New evaluation metrics for classification, regression, and ranking tasks

The \hanlp.metrics\ module now provides a suite of new metric classes for model evaluation. This includes \CategoricalAccuracy\ and \BooleanAccuracy\ for classification tasks, \F1\ and \F1\_\ for precision/recall/F1 scoring, \SpearmanCorrelation\ for ranking/regression correlation, and \MetricDict\ for managing multiple metrics simultaneously. These additions enable more comprehensive monitoring of model performance across different NLP tasks.

hanlp/metrics · high confidence

New multi-task learning (MTL) task components for parsing and tagging

The \hanlp/components/mtl/tasks\ directory now contains a complete set of new task implementations for multi-task learning, including \Task\ (the abstract base class), \amr\ (Abstract Meaning Representation parsing), \constituency\ (constituency parsing), \dep\ (dependency parsing), \dep\_2nd\ (secondary dependency parsing), \lem\ (lemmatization), \pos\ (part-of-speech tagging), and \sdp\ (semantic dependency parsing). These components provide the core logic for data loading, loss computation, and model building for these specific NLP tasks within the MTL framework.

hanlp/components/mtl/tasks · high confidence

New multilingual NLP demos and training scripts

Added new demo scripts for language identification (supporting 176 languages) and multi-task learning (tokenization, POS, lemmatization, NER, SRL, dependency parsing, semantic dependency parsing, and constituency parsing) using the XLMR base model, along with a training script for the UD-MTL pipeline.

_plugins/hanlp\_demo/hanlp\demo/mul · medium confidence

New neural network layer implementations for HanLP

The \hanlp.layers\ package now includes several new PyTorch and TensorFlow layer modules to support advanced NLP architectures. These additions include \CnnEncoder\ for convolutional sequence encoding, \FeedForward\ for multi-layer feed-forward networks, \ScalarMixWithDropout\ for weighted layer averaging with dropout, \TimeDistributed\ for applying modules across time steps, and \WeightNormalization\ for accelerating convergence. The package also introduces specialized dropout mechanisms (\WordDropout\, \SharedDropout\, \IndependentDropout\, \LockedDropout\) and a \ConfigTracker\-enabled wrapper for feed-forward layers, expanding the toolkit available for building and configuring deep learning models within HanLP.

hanlp/layers · high confidence

New pipeline and component architecture for NLP tasks

The \hanlp.components\ module now introduces a modular pipeline system (\Pipeline\ and \Pipe\ classes) that allows chaining multiple processing steps, each with configurable input and output keys. This enables users to construct complex NLP workflows by combining components like \LambdaComponent\ (for custom functions) and specific task models such as \TransformerLemmatizer\ and \TransformerTaggingLemmatizer\. The pipeline supports serialization, copying, and flexible data flow between stages, facilitating the creation of multi-stage NLP pipelines.

hanlp/components · high confidence

New pretrained model registry for NLP tasks

The library now provides a comprehensive registry of pretrained models for a wide range of natural language processing tasks, including tokenization, part-of-speech tagging, named entity recognition, dependency and semantic dependency parsing, constituency parsing, semantic role labeling, text classification, language identification, and word embeddings. This update introduces support for modern architectures such as ModernBERT, MiniLMv2, and ERNIE-GRAM, alongside traditional models like BERT, ALBERT, and Electra. The registry covers 130+ languages for multilingual tasks and includes specialized models for Ancient Chinese, Japanese, and English. Users can now easily access and download these models for various NLP pipelines.

hanlp/pretrained · high confidence

New regression and tagging tokenization tasks for multi-task learning

Added new \RegressionTokenization\ and \TaggingTokenization\ task implementations in the \hanlp/components/mtl/tasks/tok\ directory. These introduce support for regression-based tokenization and sequence tagging-based tokenization within the multi-task learning framework, enabling users to perform tokenization tasks using different modeling approaches (regression vs. CRF/tagging) with configurable parameters like \crf\, \tagging\_scheme\, and dictionary integration.

hanlp/components/mtl/tasks/tok · high confidence

New tokenizer components for RNN, Transformer, and multi-criteria segmentation

The \hanlp/components/tokenizers\ package now includes new implementations for tokenization: \RNNTokenizer\ and \TransformerTokenizer\ (PyTorch), \BMESTokenizerTF\, \NgramConvTokenizerTF\, \TransformerTokenizerTF\, and \RNNTokenizerTF\ (TensorFlow), as well as \MultiCriteriaTransformerTaggingTokenizer\ for multi-criteria word segmentation. These components introduce support for span-based tokenization with dictionary integration (\dict\_force\ and \dict\_combine\) and multi-criteria evaluation, enabling users to perform word segmentation using either RNN or Transformer architectures across both PyTorch and TensorFlow backends.

hanlp/components/tokenizers · high confidence

Project initialization and repository structure

The repository was initialized with essential project files: a comprehensive .gitignore for Python, Java, and IDE artifacts; a CITATION.cff file for academic citation; an Apache 2.0 LICENSE; a detailed README.md documenting the library's features and usage; and a setup.py for Python package distribution. Additionally, legacy build files (HanLP.iml, config.yaml) were removed.

(repo-wide) · high confidence

Removals

Removal of legacy data structures and algorithms

Removed several legacy classes and files, including the BinTrie and DoubleArrayTrie implementations, as well as utility classes for array comparison, binary search, edit distance, vector distance, and Viterbi algorithms. These components, which were previously used for dictionary management and text processing, have been deleted from the codebase.

main/java · high confidence

Behavioural changes

Major refactoring of the hanlp.utils package and introduction of new utility modules

The hanlp.utils package has been significantly reorganized and expanded with new modules for handling components, I/O, logging, string processing, and framework-specific utilities. Key changes include the addition of component loading logic in component\_util.py that supports loading from meta files and upgrading legacy TensorFlow components to the new architecture. New utility functions have been introduced for sentence splitting, span manipulation (BMES/IOBES to BILOU conversion), and time tracking. The logging system has been enhanced with colored output and better error reporting. Additionally, TensorFlow and PyTorch-specific utility functions have been added to support model loading, GPU management, and data processing.

hanlp/utils · high confidence

Restructured tagger components and added support for conditional-matching custom dictionaries

The tagger module has been reorganized into a new \hanlp.components.taggers\ package, separating PyTorch and TensorFlow implementations into distinct files (e.g., \rnn\_tagger.py\, \rnn\_tagger\_tf.py\). This refactor introduces support for conditional-matching custom dictionaries (\dict\_tags\) in the \pos\ and \ner\ taggers, allowing users to override model predictions with predefined tag mappings. The change also includes a fix for training CRF models in \TaggingNamedEntityRecognition\ and adds support for \eval\_trn\ to speed up training loops.

hanlp/components/taggers · medium confidence

Test coverage

Add comprehensive test coverage for multi-task learning and utility functions; Removed obsolete test files.

Dependencies

Migrate to modular Maven build for the Java RESTful client

The project's build structure has been refactored by moving the Java RESTful client into its own \plugins/hanlp\_restful\_java\ directory with a dedicated \pom.xml\. This new module explicitly declares dependencies on \jackson-databind\ (v2.14.1) and \junit-jupiter\ (RELEASE), while the previous root \pom.xml\ containing older dependencies like \slf4j\, \logback\, and \jackson\ v1.9.13 has been removed.

(dependencies) · medium confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

Score

  • CAI 52 → 54 (+1.8)

Lenses

  • Code Health 65 → 65 (-0.2)
  • Architecture 99 → 99 (+0.0)
  • Maturity 50 → 57 (+6.6)
  • Readiness 46 → 46 (+0.0)
  • Security 48 → 51 (+3.3)
  • Domain Modelling 100 → 100 (+0.0)

Resolved (86)

  • Duplicated block (10 lines × 2) (hanlp/components/amr/amrbart/common/postprocessing.py)
  • Duplicated block (10 lines × 2) (hanlp/components/amr/amrbart/model_interface/modeling_bart.py)
  • Duplicated block (10 lines × 2) (hanlp/components/ner/biaffine_ner/biaffine_ner.py)
  • Duplicated block (10 lines × 2) (hanlp/components/parsers/constituency/crf_constituency_parser.py)
  • Duplicated block (10 lines × 2) (hanlp/components/parsers/ud/ud_parser.py)
  • Duplicated block (10 lines × 2) (hanlp/metrics/parsing/semdep_eval.py)
  • Duplicated block (10 lines × 2) (hanlp/transform/conll_tf.py)
  • Duplicated block (11 lines × 2) (hanlp/components/amr/amrbart/common/postprocessing.py)
  • Duplicated block (11 lines × 2) (hanlp/components/amr/seq2seq/dataset/postprocessing.py)
  • Duplicated block (11 lines × 2) (hanlp/components/amr/seq2seq/dataset/tokenization_bart.py)
  • Duplicated block (11 lines × 2) (hanlp/components/amr/seq2seq/dataset/tokenization_bart.py)
  • Duplicated block (11 lines × 2) (hanlp/components/amr/seq2seq/dataset/tokenization_bart.py)
  • Duplicated block (11 lines × 2) (hanlp/components/ner/biaffine_ner/biaffine_ner.py)
  • Duplicated block (11 lines × 2) (hanlp/components/srl/span_rank/highway_variational_lstm.py)
  • Duplicated block (11 lines × 3) (hanlp/components/amr/amrbart/preprocess/read_and_process.py)
  • Duplicated block (12 lines × 2) (hanlp/components/amr/amrbart/bart_amr_generation.py)
  • Duplicated block (12 lines × 2) (hanlp/components/amr/amrbart/common/postprocessing.py)
  • Duplicated block (12 lines × 2) (hanlp/transform/conll_tf.py)
  • Duplicated block (12 lines × 3) (hanlp/components/amr/amrbart/model_interface/tokenization_bart.py)
  • Duplicated block (13 lines × 2) (hanlp/components/amr/amrbart/common/postprocessing.py)
  • …and 66 more

New (94)

  • Duplicated block (10 lines × 2) (hanlp/components/amr/amrbart/common/postprocessing.py)
  • Duplicated block (10 lines × 2) (hanlp/components/amr/seq2seq/dataset/postprocessing.py)
  • Duplicated block (10 lines × 2) (hanlp/components/amr/seq2seq/dataset/tokenization_bart.py)
  • Duplicated block (10 lines × 2) (hanlp/components/amr/seq2seq/dataset/tokenization_bart.py)
  • Duplicated block (10 lines × 2) (hanlp/components/amr/seq2seq/dataset/tokenization_bart.py)
  • Duplicated block (10 lines × 2) (hanlp/components/srl/span_rank/highway_variational_lstm.py)
  • Duplicated block (10 lines × 3) (hanlp/components/amr/amrbart/model_interface/tokenization_bart.py)
  • Duplicated block (10 lines × 3) (hanlp/components/amr/amrbart/preprocess/read_and_process.py)
  • Duplicated block (11 lines × 2) (hanlp/components/amr/amrbart/bart_amr_generation.py)
  • Duplicated block (11 lines × 2) (hanlp/components/amr/amrbart/common/postprocessing.py)
  • Duplicated block (11 lines × 2) (hanlp/components/amr/seq2seq/dataset/tokenization_bart.py)
  • Duplicated block (11 lines × 2) (hanlp/components/ner/biaffine_ner/biaffine_ner.py)
  • Duplicated block (11 lines × 2) (hanlp/transform/conll_tf.py)
  • Duplicated block (12 lines × 2) (hanlp/components/amr/amrbart/common/postprocessing.py)
  • Duplicated block (12 lines × 2) (hanlp/components/amr/amrbart/common/postprocessing.py)
  • Duplicated block (12 lines × 2) (hanlp/components/amr/seq2seq/dataset/tokenization_bart.py)
  • Duplicated block (12 lines × 2) (hanlp/components/amr/seq2seq/dataset/tokenization_bart.py)
  • Duplicated block (12 lines × 2) (hanlp/components/parsers/constituency/crf_constituency_parser.py)
  • Duplicated block (12 lines × 2) (hanlp/components/parsers/ud/ud_parser.py)
  • Duplicated block (12 lines × 3) (hanlp/components/amr/amrbart/model_interface/tokenization_bart.py)
  • …and 74 more

Architecture

  • Unchanged — 0 containers · 1 contexts · 0 edges

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

hankcs/HanLP was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 2 August 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit 942ce9306c953cc9f445a448c8e25584bb561453 — the exact code this score is about.
  • Scored under rubric-2026.08.18 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer latest.