hankcs/HanLP
53.7
Adequate · 2 August 2026
50.3k
lines of production code
Python
primary language
3
measurements over time
What this system is
HanLP is a comprehensive natural language processing library that provides a modular, pipeline-based architecture for building and training NLP models. It supports a wide range of tasks including tokenization, part-of-speech tagging, named entity recognition, dependency and constituency parsing, semantic role labeling, and text classification. The system offers both PyTorch and TensorFlow backends, enabling users to construct complex workflows by chaining components, and includes utilities for model compression, multi-task learning, and evaluation metrics.
How it got here
2014–2020 — Modular architecture and backend expansion
28 changes.
The project underwent a significant architectural shift towards a modular, component-based design, introducing a unified pipeline system and abstract base classes for NLP components. This period also focused on expanding framework support by adding comprehensive TensorFlow implementations for existing PyTorch models, alongside building out a rich ecosystem of layers, metrics, and dataset loaders.
2021–2025 — Multi-task learning and component expansion
27 changes.
This period was defined by the introduction of a comprehensive Multi-Task Learning (MTL) framework, enabling joint training across diverse NLP tasks such as parsing, tagging, and NER. The codebase expanded significantly with new components for knowledge distillation, semantic textual similarity, and masked language modeling, alongside extensive demo scripts for multiple languages.
Features
Add Biaffine-based NER component with transformer support
A new Named Entity Recognition component based on the Biaffine architecture has been added to HanLP. This implementation treats every possible span as a candidate entity and predicts its label, supporting both flat and nested NER tasks. The component integrates with the existing TorchComponent framework, allowing for optional transformer-based backends and standard PyTorch training loops with F1 metric tracking.
_hanlp/components/ner/biaffine\ner · high confidence
Add CRF layer and file-reading utility
The \hanlp/layers/crf\ module now includes a PyTorch-based Conditional Random Field (CRF) layer (\crf.py\) and a TensorFlow-based counterpart (\crf\_tf.py\), enabling sequence labeling tasks in both frameworks. Additionally, the \hanlp/utils/file\_read\_backwards\ package has been added, providing a \FileReadBackwards\ utility for reading files from the end to the beginning, which supports multiple encodings and chunked reading.
_hanlp/layers/crf, hanlp/utils/file\_read\backwards · high confidence
Add Chinese NLP demo and training scripts for TensorFlow
Added a new directory for Chinese language processing demos and training scripts using TensorFlow. This includes demo scripts for tasks such as word segmentation, part-of-speech tagging, named entity recognition, dependency parsing, and sentiment analysis, as well as training scripts for various models including BERT, ALBERT, and Electra on datasets like MSRA, CTB, and PTB.
_plugins/hanlp\_demo/hanlp\demo/zh/tf · high confidence
Add Chinese NLP training scripts and Windows compatibility check
Added new scripts for fine-tuning Named Entity Recognition (NER) and training multi-task learning models for Chinese NLP tasks (tokenization, POS, NER, SRL, dependency parsing, etc.) in the HanLP demo plugin. The \finetune\_ner.py\ script demonstrates how to load a pretrained model and fine-tune it on a custom dataset. The \open\_base.py\ and \open\_small.py\ scripts provide complete training pipelines for multiple Chinese NLP tasks using the Electra base and small models. Additionally, a \block\_windows.py\ module was added to prevent execution on Windows, as the training scripts are not supported on that platform.
_plugins/hanlp\_demo/hanlp\demo/zh/train · high confidence
Add Japanese multi-task learning demo
A new demo script for Japanese NLP has been added, showcasing a multi-task learning model trained on NPCMJ, UD, and Kyoto corpora. The script loads the pre-trained Japanese model (tok, pos, con, BERT base char) and demonstrates tokenization, part-of-speech tagging, and dependency parsing on sample sentences.
_plugins/hanlp\_demo/hanlp\demo/ja · high confidence
Add Masked Language Model component for text completion
A new Masked Language Model (MLM) component has been introduced to support filling blanks in text. The implementation includes a dataset class and a TorchComponent that leverages Hugging Face's AutoModelForMaskedLM. The model's prediction logic explicitly blocks special tokens from the output, ensuring that only valid vocabulary tokens are returned during inference.
hanlp/components/lm · high confidence
Add STS component for semantic textual similarity
A new STS (Semantic Textual Similarity) component has been added to the library. This introduces a transformer-based baseline model that fine-tunes a pre-trained transformer with a regression head to predict similarity scores between sentence pairs. The implementation includes a data loader for the SemanticTextualSimilarityDataset, a training loop with Spearman correlation metrics, and a predict method that supports both flat and nested list inputs for sentence pairs.
hanlp/components/sts · high confidence
Add Semantic Role Labeling tasks to the multi-task learning framework
Introduced two new Semantic Role Labeling (SRL) tasks for the multi-task learning (MTL) framework: a span-based BIO tagger and a span-ranking model. These additions enable joint training of SRL alongside other NLP tasks within HanLP's unified pipeline.
hanlp/components/mtl/tasks/srl · high confidence
Add TensorFlow-based Biaffine Dependency Parser
The \hanlp/components/parsers\ directory now includes a complete TensorFlow implementation of the Biaffine Dependency Parser (\BiaffineDependencyParserTF\), mirroring the existing PyTorch version. This adds support for training and inference using TensorFlow/Keras, including algorithmic components like K-Means clustering, Eisner's algorithm, and Chu-Liu/Edmonds decoding adapted for TF, enabling users to leverage the parser with TensorFlow backends.
hanlp/components/parsers · high confidence
Add TensorFlow-based NLP components for parsing, tagging, and CRF
Added new TensorFlow implementations for sequence labeling and parsing tasks. This includes a Conditional Random Field (CRF) layer and loss function for sequence classification, a CNN-based tagger for word-level tagging, an n-gram convolutional tagger for character/word-level tagging, and a Transformer-based tagger for sequence tagging. Additionally, a TensorFlow implementation of the Biaffine dependency parser was added. These components provide native TensorFlow support for these NLP tasks, mirroring existing PyTorch/MXNet counterparts.
python · high confidence
Add chunking and sequence-labeling metrics (SpanF1, ChunkingF1, BMES/IOBES F1)
The hanlp/metrics/chunking package now includes new metric classes for evaluating chunking and sequence-labeling tasks. This includes SpanF1 and ChunkingF1 for standard Python/tensor-based evaluation, as well as TensorFlow-specific implementations (BMES\_F1\_TF, IOBES\_F1\_TF) that integrate with Keras training loops. Users can now compute F1 scores for named entity recognition, part-of-speech tagging, or any sequence-labeling task using these built-in metrics.
hanlp/metrics/chunking · high confidence
Add parsing evaluation metrics and evaluation scripts
The \hanlp/metrics/parsing\ module now includes new metric classes for evaluating parsing performance: \AttachmentScore\ (computing UAS/LAS), \LabeledF1\ (computing UF/LF F1 scores), \LabeledF1TF\ (TensorFlow-based Labeled F1), \LabeledScore\ (TensorFlow-based UAS/LAS), and \SpanMetric\ (computing UCM/LCM and related precision/recall/F1 for span-based parsing). Additionally, evaluation scripts for CoNLL-X and semantic dependencies have been added to support standard benchmark evaluation.
hanlp/metrics/parsing · high confidence
Added AdamW optimizer with weight decay support
Introduced a new \AdamWeightDecay\ optimizer and a \WarmUp\ learning rate schedule in \hanlp/optimizers/adamw/\. This adds support for L2 weight decay that is decoupled from the Adam moment estimates, along with a polynomial warmup schedule for training stability.
hanlp/optimizers · high confidence
Added CRF-based constituency parser with MBR decoding
Users can now use a new constituency parser based on a Conditional Random Field (CRF) model, which supports Minimum Bayes-Risk (MBR) decoding for improved parsing accuracy. This component, located in the constituency parser module, introduces a two-stage neural CRF architecture that computes span and label scores, enabling more robust syntactic analysis.
hanlp/components/parsers/constituency · high confidence
Added custom loss functions for sparse categorical cross-entropy
The hanlp/losses module now includes three new loss functions: SparseCategoricalCrossentropyOverNonzeroWeights, SparseCategoricalCrossentropyOverBatchFirstDim, and MaskedSparseCategoricalCrossentropyOverBatchFirstDim. These provide specialized handling for sample weighting and masking in classification tasks, addressing previous compatibility issues with TensorFlow 2.3 and resolving errors during model export.
hanlp/losses · high confidence
Added demo scripts for Ancient Chinese (LZH) models
New demo scripts (demo\_mtl.py, demo\_tok.py) were added to the lzh directory, providing examples for loading and running the KYOTO\_EVAHAN\_TOK\_LZH and KYOTO\_EVAHAN\_TOK\_LEM\_POS\_UDEP\_LZH models on Ancient Chinese text.
_plugins/hanlp\_demo/hanlp\demo/lzh · high confidence
Added language-specific utility modules for English, Japanese, and Chinese
New utility modules have been added to support specific languages. For English, a regex-based word tokenizer is now available to handle contractions, possessives, and line breaks. For Japanese, a custom BERT Japanese tokenizer wrapper is provided to handle tokenization and offset mapping. For Chinese, a character normalization table and localization mappings for tasks, POS tags, NER, and dependency parsing are now available. These utilities provide foundational tools for processing text in these languages.
hanlp/utils/lang · high confidence
Added new components for knowledge distillation, sentence boundary detection, and semantic role labeling
The HanLP library now includes new modules for model compression and NLP tasks. The \hanlp.components.distillation\ package introduces a \DistillableComponent\ and associated loss functions (e.g., \kd\_ce\_loss\, \hid\_mse\_loss\) and temperature schedulers (\flsw\, \cwsm\) to support knowledge distillation training. Additionally, new components have been added for sentence boundary detection (\hanlp.components.eos\), semantic role labeling (\hanlp.components.srl\), and RNN-based tagging (\hanlp.components.taggers.rnn\), along with corresponding metrics for evaluation.
(repo-wide) · high confidence
English NLP demo scripts added
Added demo scripts for English language processing, including tokenization, part-of-speech tagging, dependency parsing, semantic dependency parsing, named entity recognition, language modeling, sentiment analysis, and an AMR (Abstract Meaning Representation) parser. A training script for the Stanford Sentiment Treebank (SST2) with ALBERT is also included.
_plugins/hanlp\_demo/hanlp\demo/en · high confidence
HanLP v2.1.3 release with AMR metrics and load/pipeline APIs
This update releases HanLP version 2.1.3, introducing a new \load()\ function in the main \hanlp\ namespace to simplify loading pretrained components and a \pipeline()\ helper for bundling components. It also adds support for Abstract Meaning Representation (AMR) evaluation via a new \smatch\_eval\ metric, and includes various bug fixes and compatibility improvements for TensorFlow, PyTorch, and macOS M1 chips.
hanlp · high confidence
Introduce HanLP RESTful API client with new NLP capabilities
Adds the \hanlp\_restful\ package, providing a Python client (\HanLPClient\) to interact with HanLP's RESTful APIs. This enables users to perform various NLP tasks including text style transfer, abstract meaning representation, keyphrase extraction, extractive summarization, and coreference resolution. The client supports specifying the language (e.g., 'ja' for Japanese, 'mul' for multilingual) and allows for coarse tokenization. It also includes a \verify\ parameter to control SSL certificate verification, supporting environments behind private TLS gateways. Tests are added to verify client functionality.
_plugins/hanlp\restful · high confidence
Introduce HanLP Trie plugin with tokenization and batch processing
Added the \plugins/hanlp\_trie\ module, providing a \Trie\ data structure and \TrieDict\/\TupleTrieDict\ classes for fast, rule-based tokenization and longest-prefix matching. The implementation supports batch processing via \split\_batch\ and \merge\_batch\ methods, enabling efficient integration of custom dictionaries into NLP pipelines. Tests verify core trie operations and dictionary tokenization behavior.
_plugins/hanlp\trie · high confidence
Introduce HanLP common utilities and visualization
Added a new \hanlp\_common\ plugin providing shared utilities, data structures, and visualization tools. This includes a \Document\ class for handling parsed annotations, \CoNLL\ and \CoNLLU\ word/sentence classes for parsing formats, and enhanced visualization for constituency trees, dependency graphs, and NER/SRL/CON outputs. The update also adds support for HTML visualization in Jupyter notebooks, improves the robustness of various visualizations, and allows users to disable IPython-specific features.
_plugins/hanlp\_common/hanlp\common · medium confidence
Introduce Java RESTful client with new NLP task methods
Added a new Java RESTful client (\HanLPClient\) and associated data models (\BaseInput\, \DocumentInput\, \SentenceInput\, \TokenInput\, \Span\, \CoreferenceResolutionOutput\, and MRP classes) that expose methods for tokenization, parsing, coreference resolution, text style transfer, semantic textual similarity, extractive/abstractive summarization, grammatical error correction, language identification, sentiment analysis, and abstract meaning representation. The client supports specifying tasks to run or skip, handles authentication via an \auth\ parameter or \HANLP\_AUTH\ environment variable, and includes tests for these new capabilities.
_plugins/hanlp\_restful\java · high confidence
Introduce Multi-Task Learning (MTL) framework
Users can now train multiple NLP tasks jointly using a shared encoder and task-specific decoders. This new MTL component supports dynamic task sampling during training, allowing for more efficient multi-task optimization.
hanlp/components/mtl · high confidence
Introduce TensorFlow-based data transformation and tokenization components
Added a new \hanlp/transform\ module containing a suite of data transformation classes for the TensorFlow backend. This includes \TransformerSequenceTokenizer\ and \TransformerTextTokenizer\ for handling Hugging Face tokenizers with support for sliding windows, prefix matching, and subtoken span tracking. The module also introduces \TableTransform\ and \TsvTaggingTransform\ for processing tabular and TSV datasets, as well as \CoNLLTransform\ for parsing CoNLL-style NLP tasks. Additionally, a \PrependSpace\ transform is provided to handle whitespace normalization for tokenizers that require it.
hanlp/transform · high confidence
Introduce Transformer-based tagger with PyTorch and TensorFlow backends
Users can now use a new TransformerTagger component that supports both PyTorch and TensorFlow backends. The PyTorch implementation (transformer\_tagger.py) provides a shallow tagging model using a linear layer with an optional CRF layer, featuring support for extra embeddings, teacher-student distillation, and dynamic vocabulary resizing. The TensorFlow implementation (transformer\_transform\_tf.py) includes a custom accuracy metric that handles masking and supports variable sequence lengths. This change enables training and inference for sequence labeling tasks like NER and POS tagging using modern transformer architectures.
hanlp/components/taggers/transformers · high confidence
Introduce a new Transformer encoder layer with sliding window and mirror support
The \hanlp/layers/transformers\ package now includes a new \TransformerEncoder\ layer that supports sliding window processing for long sequences, word dropout, and scalar mix aggregation. The implementation adds a \resource.py\ module that maps model identifiers to local mirror paths, enabling faster access to pre-trained models. Additionally, the \pt\_imports.py\ and \tf\_imports.py\ modules provide wrapper classes for Hugging Face and TensorFlow transformers, while \relative\_transformer.py\ introduces relative positional embeddings for sequence modeling.
hanlp/layers/transformers · high confidence
Introduce common component and dataset base classes
The \hanlp/common\ package now provides foundational classes for building NLP components. \Component\ serves as the abstract base for all components, while \TorchComponent\ and \KerasComponent\ offer PyTorch and TensorFlow-specific implementations with shared workflows for building models, handling vocabs, and managing configuration. Additionally, \TransformableDataset\ and \Transform\ classes enable flexible data transformation pipelines for both PyTorch and TensorFlow backends.
hanlp/common · high confidence
Introduce span-based Semantic Role Labeling (SRL) component
Added a new \span\_rank\ component for joint predicate and argument prediction in Semantic Role Labeling. This implementation, adapted from an external research repository, introduces a span-ranking architecture that generates candidate spans and ranks them using a bidirectional LSTM with highway connections, biaffine scoring, and greedy/DP-based decoding utilities. The addition includes the core model (\SpanRankingSRLModel\), the network layers (\highway\_variational\_lstm.py\, \layer.py\), and inference/evaluation utilities (\inference\_utils.py\, \srl\_eval\_utils.py\), enabling users to perform SRL tasks using this specific ranking-based approach.
_hanlp/components/srl/span\_bio, hanlp/components/srl/span\rank · high confidence
Introduces Universal Dependencies parser with lemma rule-based lemmatization
The \hanlp/components/parsers/ud\ directory now contains a complete implementation for Universal Dependencies parsing, including the \UniversalDependenciesParser\ and its associated model, decoder, and utility modules. This adds support for POS tagging, dependency parsing, feature prediction, and lemmatization. A key feature is the use of edit-distance-based lemma rules (\lemma\_edit.py\) to handle lemmatization, which allows the model to generate lemmas by applying transformation rules rather than relying on external dependencies like \alnlp\. The implementation supports both adaptive and standard loss functions for tag prediction and integrates with the existing \ContextualWordEmbedding\ encoder.
hanlp/components/parsers/ud · high confidence
Introduction of HanLP Common library
A new 'hanlp\_common' package has been added to the repository, providing shared utilities and structures for the HanLP NLP library. This includes the package's setup configuration, which defines the 'hanlp\_common' distribution (version 0.0.23) and its dependencies, alongside a README file that outlines the library's purpose and installation instructions.
_plugins/hanlp\common · high confidence
New BART-based AMR parser and generator components
The AMR component now includes a new BART-based architecture for both AMR parsing (text to AMR graph) and AMR generation (AMR graph to text). This introduces \BART\_AMR\_Parser\ and \BART\_AMR\_Generation\ classes, along with supporting modules for tokenization, dataset handling, and post-processing, enabling state-of-the-art performance on English AMR tasks.
hanlp/components/amr · high confidence
New Biaffine parser components for dependency and semantic parsing
The \hanlp/components/parsers/biaffine\ directory now contains the full implementation of the Biaffine parser, including the core \Biaffine\ layer, \MLP\ modules, and parser classes for standard dependency parsing (\BiaffineDependencyParser\), semantic dependency parsing (\BiaffineSemanticDependencyParser\), and secondary/structural attention parsing. This adds support for joint and separate decoding of arc and relation scores, with options for single-root constraints and cycle avoidance in the parsing logic.
hanlp/components/parsers/biaffine · high confidence
New Chinese NLP demos and utilities for HanLP v2.1
Added a suite of new demo scripts and utilities for Chinese NLP tasks in the \plugins/hanlp\_demo/hanlp\_demo/zh\ directory. These include demonstrations for AMR parsing, custom dictionaries for tokenization and POS tagging, task removal from multi-task learning models, document handling, masked language modeling, constituency parsing pipelines, semantic textual similarity, word2vec similarity, and training scripts for state-of-the-art Chinese word segmentation models.
_plugins/hanlp\_demo/hanlp\demo/zh · high confidence
New NER components for RNN and Transformer backends
Added new named entity recognition (NER) components for both RNN and Transformer architectures. The RNN implementation (\rnn\_ner.py\) supports word2vec/fasttext embeddings and handles IOBES/BIO/BIOUL tagging schemes. The Transformer implementation (\transformer\_ner.py\) integrates with pre-trained models and supports dictionary-based longest-prefix matching for entity recognition, along with blacklist filtering and entity type merging. A new TensorFlow-based NER component (\ner\_tf.py\) provides equivalent functionality using Keras/TF backends.
hanlp/components/ner · high confidence
New NER task implementations with dictionary support
Added two new Named Entity Recognition (NER) task implementations: a Biaffine-based model and a Tagging-based model (with optional CRF). The Tagging model introduces support for conditional-matching custom dictionaries (whitelist/blacklist) to override or filter entity predictions, and allows merging of consecutive entities of the same type.
hanlp/components/mtl/tasks/ner · high confidence
New character-level and contextual embedding layers for PyTorch and TensorFlow
The \hanlp/layers/embeddings\ module now includes new embedding implementations: \CharCNN\ and \CharRNN\ for character-level embeddings, \ContextualStringEmbedding\ for loading pre-trained language models, \ContextualWordEmbedding\ for transformer-based contextual embeddings, and \FastText\/\Word2Vec\ support. These are available for both PyTorch and TensorFlow backends, enabling users to integrate character-level and contextualized word embeddings into their models.
hanlp/layers/embeddings · high confidence
New classifier components for Hugging Face models and regression tasks
Added new components to support loading Hugging Face models directly for both classification and regression tasks. The \TransformerClassifierHF\ class enables inference using \AutoModelForSequenceClassification\ from the \transformers\ library, while \TransformerRegressionHF\ provides a similar interface for regression tasks. Additionally, a \FastTextClassifier\ component was introduced to support classification using the FastText library. These additions expand the range of pre-trained models and algorithms available for text classification and regression within the library.
hanlp/components/classifiers · high confidence
New dataset loaders and dataset classes for multiple NLP tasks
The \hanlp/datasets\ module has been reorganized and expanded with new dataset loaders and classes for various NLP tasks. This includes a \LanguageModelDataset\ for language modeling, a \SentenceBoundaryDetectionDataset\ for end-of-sentence detection, and loaders for coreference resolution, named entity recognition (NER), and parsing tasks (including CoNLL formats and Chinese Treebank variants). These changes provide standardized interfaces for loading and processing diverse linguistic datasets within the library.
hanlp/datasets · high confidence
New evaluation metrics for classification, regression, and ranking tasks
The \hanlp.metrics\ module now provides a suite of new metric classes for model evaluation. This includes \CategoricalAccuracy\ and \BooleanAccuracy\ for classification tasks, \F1\ and \F1\_\ for precision/recall/F1 scoring, \SpearmanCorrelation\ for ranking/regression correlation, and \MetricDict\ for managing multiple metrics simultaneously. These additions enable more comprehensive monitoring of model performance across different NLP tasks.
hanlp/metrics · high confidence
New multi-task learning (MTL) task components for parsing and tagging
The \hanlp/components/mtl/tasks\ directory now contains a complete set of new task implementations for multi-task learning, including \Task\ (the abstract base class), \amr\ (Abstract Meaning Representation parsing), \constituency\ (constituency parsing), \dep\ (dependency parsing), \dep\_2nd\ (secondary dependency parsing), \lem\ (lemmatization), \pos\ (part-of-speech tagging), and \sdp\ (semantic dependency parsing). These components provide the core logic for data loading, loss computation, and model building for these specific NLP tasks within the MTL framework.
hanlp/components/mtl/tasks · high confidence
New multilingual NLP demos and training scripts
Added new demo scripts for language identification (supporting 176 languages) and multi-task learning (tokenization, POS, lemmatization, NER, SRL, dependency parsing, semantic dependency parsing, and constituency parsing) using the XLMR base model, along with a training script for the UD-MTL pipeline.
_plugins/hanlp\_demo/hanlp\demo/mul · medium confidence
New neural network layer implementations for HanLP
The \hanlp.layers\ package now includes several new PyTorch and TensorFlow layer modules to support advanced NLP architectures. These additions include \CnnEncoder\ for convolutional sequence encoding, \FeedForward\ for multi-layer feed-forward networks, \ScalarMixWithDropout\ for weighted layer averaging with dropout, \TimeDistributed\ for applying modules across time steps, and \WeightNormalization\ for accelerating convergence. The package also introduces specialized dropout mechanisms (\WordDropout\, \SharedDropout\, \IndependentDropout\, \LockedDropout\) and a \ConfigTracker\-enabled wrapper for feed-forward layers, expanding the toolkit available for building and configuring deep learning models within HanLP.
hanlp/layers · high confidence
New pipeline and component architecture for NLP tasks
The \hanlp.components\ module now introduces a modular pipeline system (\Pipeline\ and \Pipe\ classes) that allows chaining multiple processing steps, each with configurable input and output keys. This enables users to construct complex NLP workflows by combining components like \LambdaComponent\ (for custom functions) and specific task models such as \TransformerLemmatizer\ and \TransformerTaggingLemmatizer\. The pipeline supports serialization, copying, and flexible data flow between stages, facilitating the creation of multi-stage NLP pipelines.
hanlp/components · high confidence
New pretrained model registry for NLP tasks
The library now provides a comprehensive registry of pretrained models for a wide range of natural language processing tasks, including tokenization, part-of-speech tagging, named entity recognition, dependency and semantic dependency parsing, constituency parsing, semantic role labeling, text classification, language identification, and word embeddings. This update introduces support for modern architectures such as ModernBERT, MiniLMv2, and ERNIE-GRAM, alongside traditional models like BERT, ALBERT, and Electra. The registry covers 130+ languages for multilingual tasks and includes specialized models for Ancient Chinese, Japanese, and English. Users can now easily access and download these models for various NLP pipelines.
hanlp/pretrained · high confidence
New regression and tagging tokenization tasks for multi-task learning
Added new \RegressionTokenization\ and \TaggingTokenization\ task implementations in the \hanlp/components/mtl/tasks/tok\ directory. These introduce support for regression-based tokenization and sequence tagging-based tokenization within the multi-task learning framework, enabling users to perform tokenization tasks using different modeling approaches (regression vs. CRF/tagging) with configurable parameters like \crf\, \tagging\_scheme\, and dictionary integration.
hanlp/components/mtl/tasks/tok · high confidence
New tokenizer components for RNN, Transformer, and multi-criteria segmentation
The \hanlp/components/tokenizers\ package now includes new implementations for tokenization: \RNNTokenizer\ and \TransformerTokenizer\ (PyTorch), \BMESTokenizerTF\, \NgramConvTokenizerTF\, \TransformerTokenizerTF\, and \RNNTokenizerTF\ (TensorFlow), as well as \MultiCriteriaTransformerTaggingTokenizer\ for multi-criteria word segmentation. These components introduce support for span-based tokenization with dictionary integration (\dict\_force\ and \dict\_combine\) and multi-criteria evaluation, enabling users to perform word segmentation using either RNN or Transformer architectures across both PyTorch and TensorFlow backends.
hanlp/components/tokenizers · high confidence
Project initialization and repository structure
The repository was initialized with essential project files: a comprehensive .gitignore for Python, Java, and IDE artifacts; a CITATION.cff file for academic citation; an Apache 2.0 LICENSE; a detailed README.md documenting the library's features and usage; and a setup.py for Python package distribution. Additionally, legacy build files (HanLP.iml, config.yaml) were removed.
(repo-wide) · high confidence
Removals
Removal of legacy data structures and algorithms
Removed several legacy classes and files, including the BinTrie and DoubleArrayTrie implementations, as well as utility classes for array comparison, binary search, edit distance, vector distance, and Viterbi algorithms. These components, which were previously used for dictionary management and text processing, have been deleted from the codebase.
main/java · high confidence
Behavioural changes
Major refactoring of the hanlp.utils package and introduction of new utility modules
The hanlp.utils package has been significantly reorganized and expanded with new modules for handling components, I/O, logging, string processing, and framework-specific utilities. Key changes include the addition of component loading logic in component\_util.py that supports loading from meta files and upgrading legacy TensorFlow components to the new architecture. New utility functions have been introduced for sentence splitting, span manipulation (BMES/IOBES to BILOU conversion), and time tracking. The logging system has been enhanced with colored output and better error reporting. Additionally, TensorFlow and PyTorch-specific utility functions have been added to support model loading, GPU management, and data processing.
hanlp/utils · high confidence
Restructured tagger components and added support for conditional-matching custom dictionaries
The tagger module has been reorganized into a new \hanlp.components.taggers\ package, separating PyTorch and TensorFlow implementations into distinct files (e.g., \rnn\_tagger.py\, \rnn\_tagger\_tf.py\). This refactor introduces support for conditional-matching custom dictionaries (\dict\_tags\) in the \pos\ and \ner\ taggers, allowing users to override model predictions with predefined tag mappings. The change also includes a fix for training CRF models in \TaggingNamedEntityRecognition\ and adds support for \eval\_trn\ to speed up training loops.
hanlp/components/taggers · medium confidence
Test coverage
Add comprehensive test coverage for multi-task learning and utility functions; Removed obsolete test files.
Dependencies
Migrate to modular Maven build for the Java RESTful client
The project's build structure has been refactored by moving the Java RESTful client into its own \plugins/hanlp\_restful\_java\ directory with a dedicated \pom.xml\. This new module explicitly declares dependencies on \jackson-databind\ (v2.14.1) and \junit-jupiter\ (RELEASE), while the previous root \pom.xml\ containing older dependencies like \slf4j\, \logback\, and \jackson\ v1.9.13 has been removed.
(dependencies) · medium confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Score
- CAI 52 → 54 (+1.8)
Lenses
- Code Health 65 → 65 (-0.2)
- Architecture 99 → 99 (+0.0)
- Maturity 50 → 57 (+6.6)
- Readiness 46 → 46 (+0.0)
- Security 48 → 51 (+3.3)
- Domain Modelling 100 → 100 (+0.0)
Resolved (86)
- Duplicated block (10 lines × 2) (hanlp/components/amr/amrbart/common/postprocessing.py)
- Duplicated block (10 lines × 2) (hanlp/components/amr/amrbart/model_interface/modeling_bart.py)
- Duplicated block (10 lines × 2) (hanlp/components/ner/biaffine_ner/biaffine_ner.py)
- Duplicated block (10 lines × 2) (hanlp/components/parsers/constituency/crf_constituency_parser.py)
- Duplicated block (10 lines × 2) (hanlp/components/parsers/ud/ud_parser.py)
- Duplicated block (10 lines × 2) (hanlp/metrics/parsing/semdep_eval.py)
- Duplicated block (10 lines × 2) (hanlp/transform/conll_tf.py)
- Duplicated block (11 lines × 2) (hanlp/components/amr/amrbart/common/postprocessing.py)
- Duplicated block (11 lines × 2) (hanlp/components/amr/seq2seq/dataset/postprocessing.py)
- Duplicated block (11 lines × 2) (hanlp/components/amr/seq2seq/dataset/tokenization_bart.py)
- Duplicated block (11 lines × 2) (hanlp/components/amr/seq2seq/dataset/tokenization_bart.py)
- Duplicated block (11 lines × 2) (hanlp/components/amr/seq2seq/dataset/tokenization_bart.py)
- Duplicated block (11 lines × 2) (hanlp/components/ner/biaffine_ner/biaffine_ner.py)
- Duplicated block (11 lines × 2) (hanlp/components/srl/span_rank/highway_variational_lstm.py)
- Duplicated block (11 lines × 3) (hanlp/components/amr/amrbart/preprocess/read_and_process.py)
- Duplicated block (12 lines × 2) (hanlp/components/amr/amrbart/bart_amr_generation.py)
- Duplicated block (12 lines × 2) (hanlp/components/amr/amrbart/common/postprocessing.py)
- Duplicated block (12 lines × 2) (hanlp/transform/conll_tf.py)
- Duplicated block (12 lines × 3) (hanlp/components/amr/amrbart/model_interface/tokenization_bart.py)
- Duplicated block (13 lines × 2) (hanlp/components/amr/amrbart/common/postprocessing.py)
- …and 66 more
New (94)
- Duplicated block (10 lines × 2) (hanlp/components/amr/amrbart/common/postprocessing.py)
- Duplicated block (10 lines × 2) (hanlp/components/amr/seq2seq/dataset/postprocessing.py)
- Duplicated block (10 lines × 2) (hanlp/components/amr/seq2seq/dataset/tokenization_bart.py)
- Duplicated block (10 lines × 2) (hanlp/components/amr/seq2seq/dataset/tokenization_bart.py)
- Duplicated block (10 lines × 2) (hanlp/components/amr/seq2seq/dataset/tokenization_bart.py)
- Duplicated block (10 lines × 2) (hanlp/components/srl/span_rank/highway_variational_lstm.py)
- Duplicated block (10 lines × 3) (hanlp/components/amr/amrbart/model_interface/tokenization_bart.py)
- Duplicated block (10 lines × 3) (hanlp/components/amr/amrbart/preprocess/read_and_process.py)
- Duplicated block (11 lines × 2) (hanlp/components/amr/amrbart/bart_amr_generation.py)
- Duplicated block (11 lines × 2) (hanlp/components/amr/amrbart/common/postprocessing.py)
- Duplicated block (11 lines × 2) (hanlp/components/amr/seq2seq/dataset/tokenization_bart.py)
- Duplicated block (11 lines × 2) (hanlp/components/ner/biaffine_ner/biaffine_ner.py)
- Duplicated block (11 lines × 2) (hanlp/transform/conll_tf.py)
- Duplicated block (12 lines × 2) (hanlp/components/amr/amrbart/common/postprocessing.py)
- Duplicated block (12 lines × 2) (hanlp/components/amr/amrbart/common/postprocessing.py)
- Duplicated block (12 lines × 2) (hanlp/components/amr/seq2seq/dataset/tokenization_bart.py)
- Duplicated block (12 lines × 2) (hanlp/components/amr/seq2seq/dataset/tokenization_bart.py)
- Duplicated block (12 lines × 2) (hanlp/components/parsers/constituency/crf_constituency_parser.py)
- Duplicated block (12 lines × 2) (hanlp/components/parsers/ud/ud_parser.py)
- Duplicated block (12 lines × 3) (hanlp/components/amr/amrbart/model_interface/tokenization_bart.py)
- …and 74 more
Architecture
- Unchanged — 0 containers · 1 contexts · 0 edges
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
hankcs/HanLP was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 2 August 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 942ce9306c953cc9f445a448c8e25584bb561453 — the exact code this score is about.
- Scored under rubric-2026.08.18 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer latest.