NaturalNode/natural
55.5
Weak · 2 October 2026
18.4k
lines of production code
JavaScript
primary language
2
measurements over time
What this system is
This system is a comprehensive natural language processing library for JavaScript and TypeScript, providing a unified API for text analysis tasks. It supports a wide range of capabilities including tokenization, stemming, phonetic matching, spell checking, and sentiment analysis across numerous languages. The library also includes machine learning components for classification and part-of-speech tagging, alongside utilities for string distance metrics, n-grams, and graph algorithms. Modern infrastructure features include TypeScript support, pluggable storage backends, and performance benchmarks.
How it got here
2011 — TypeScript migration and infrastructure overhaul
22 changes.
The project underwent a comprehensive modernization, migrating to TypeScript and establishing a new ES Module build pipeline with Rollup and npm scripts. This period involved restructuring core NLP modules like classifiers, stemmers, and tokenizers to support multi-language capabilities and parallel training, while removing legacy text processing components. The effort also expanded the library's feature set with new phonetic algorithms, graph utilities, and extensive example code to improve usability and type safety.
2012–2018 — multilingual NLP and structural modernization
13 changes.
This period focused on expanding the library's multilingual capabilities by adding comprehensive support for Japanese, French, Dutch, and Indonesian, alongside new string distance metrics and sentiment analysis features. Concurrently, the codebase underwent significant structural modernization, including the introduction of a Trie data structure, the adoption of ES6 classes, and the integration of TypeScript definitions across core modules like spellcheck and classifiers.
2020–2024 — TypeScript migration and storage abstraction
6 changes.
This period focused on modernizing the codebase by adding TypeScript declarations for key modules like the Brill POS Tagger and expanding test coverage for core NLP components. It also introduced a pluggable storage backend system to support diverse persistence options and added new features such as a French Carry stemmer and tokenizer examples.
Features
Added Dutch and English Brill POS tagger data files
Added transformation rules and lexicon data files for the Dutch and English Brill part-of-speech taggers. The Dutch data includes context rules and a lexicon sourced from ml.nl.net, while the English data includes rules derived from Eric Brill's original paper and the pos-js library, along with a lexicon from pos-js. These files enable the Brill tagger to process Dutch and English text.
_lib/natural/brill\_pos\tagger/data · high confidence
Added French count and noun inflectors
New French language support has been added to the Natural library, introducing two new modules: a count inflector for generating French ordinal numbers (e.g., '1er', '2e') and a noun inflector for handling singular/plural transformations. The noun inflector includes rules for irregular plurals (such as 'ail' to 'aulx'), invariant nouns ending in -s, -x, and -z, and standard suffix changes (e.g., -al to -aux, -eau to -eaux).
lib/natural/inflectors/fr · high confidence
Added Japanese Kana transliterator
Users can now transliterate Japanese Katakana and Hiragana into roman characters using a modified Hepburn system. This new capability, exposed via the \TransliterateJa\ function in the \lib/natural/transliterators\ module, handles complex combinations such as small vowels, long vowel signs, and voiced consonants, correcting several bugs found in the standard CLDR transform rules.
lib/natural/transliterators · high confidence
Added Sentence Analyzer for syntactic and sentence-type analysis
The \lib/natural/analyzers\ module now includes a \SentenceAnalyzer\ class that processes part-of-speech tagged sentences to determine their type (declarative, interrogative, exclamatory, or command) and extract grammatical components like the subject and predicate. This addition is accompanied by TypeScript type definitions and a \SenType\ enum to support static typing for the analyzer's output.
lib/natural/analyzers · high confidence
Added TF-IDF usage examples
The examples/tfidf directory now includes three JavaScript files (array\_example.js, multiple\_terms.js, tfidf\_example.js) demonstrating how to use the natural library's TfIdf class. These examples show adding documents and calculating TF-IDF scores for single terms, multiple terms, and arrays of terms.
examples/tfidf · high confidence
Added TypeScript declarations for Brill POS Tagger
The \lib/natural/brill\_pos\_tagger\ module now includes an \index.d.ts\ file providing TypeScript type definitions for the Brill Part-of-Speech Tagger components, including \BrillPOSTagger\, \BrillPOSTrainer\, \BrillPOSTester\, \Lexicon\, \RuleSet\, \Corpus\, and \Sentence\. This enables TypeScript users to consume the module with proper type checking and IntelliSense support, while the existing \index.js\ continues to export the underlying JavaScript implementations.
_lib/natural/brill\_pos\tagger · high confidence
Added classification examples and tests
New example scripts and a test file have been added to the classification directory to demonstrate the Natural library's classification capabilities. The examples include basic usage of the BayesClassifier, methods for saving and loading trained classifiers to JSON, and an example using the LogisticRegressionClassifier with the Spanish Porter stemmer that listens to training events. Additionally, a new test file demonstrates how to apply the Maximum Entropy Classifier to Part-of-Speech tagging using a Brown Corpus excerpt.
examples/classification · high confidence
Added example scripts for CountInflector and NounInflector
New example files have been added to the inflection examples directory to demonstrate the library's inflection capabilities. The count.js script illustrates how to use CountInflector to generate ordinal numbers (e.g., 1st, 2nd, 3rd), while noun.js demonstrates NounInflector for pluralizing and singularizing nouns (e.g., radius/radii, beer/beers).
examples/inflection · high confidence
Added performance benchmarks for phonetic algorithms
New benchmark scripts have been added to measure the performance of the Metaphone, SoundEx, and Indonesian Stemmer algorithms. These benchmarks evaluate processing speed on single words, small text bodies, and medium-sized documents, providing a baseline for performance tracking.
benchmarks · high confidence
Added phonetic comparison and tokenization examples
New example scripts have been added to the phonetics module to demonstrate practical usage of the Metaphone and DoubleMetaphone algorithms. The \compare.js\ example shows how to check if two words sound alike using the Metaphone algorithm, while \tokenize\_and\_phoneticize.js\ demonstrates tokenizing a sentence and finding matches for user input using DoubleMetaphone.
examples/phonetics · high confidence
Added sentence tokenizer example script
A new example script (testSentenceTokenizer.js) was added to the examples/tokenizer directory. This script demonstrates how to use the SentenceTokenizer from the natural library by loading predefined abbreviations and demarkers, then tokenizing a sample text about renewable energy and logging the result.
examples/tokenizer · high confidence
Added stemming examples for Porter and Lancaster stemmers
New example scripts have been added to the \examples/stemming\ directory to demonstrate the usage of the \natural\ library's stemming capabilities. \stem\_corpus.js\ showcases the Porter stemmer with tokenization, while \stem\_word.js\ demonstrates the Lancaster stemmer for individual word stemming. These files provide users with concrete code samples for integrating stemming into their applications.
examples/stemming · high confidence
Expanded language support and improved sentence tokenization in the tokenizer library
The tokenizer module now includes aggressive tokenizers for additional languages, specifically adding support for Hindi (aggressive\_tokenizer\_hi.js), Indonesian (aggressive\_tokenizer\_id.js), and Ukrainian (aggressive\_tokenizer\_uk.js), while also updating existing tokenizers for languages like German, Spanish, French, and Swedish to better handle language-specific characters and diacritics. Additionally, the SentenceTokenizer has been rewritten to use a placeholder-based approach that more accurately handles abbreviations, URIs, numbers, and complex punctuation, and now includes a configurable \trimSentences\ option to allow users to preserve original whitespace if needed.
lib/natural/tokenizers · high confidence
Introduction of Daitch-Mokotoff Soundex and Double Metaphone phonetic algorithms
The \lib/natural/phonetics\ module now includes implementations for the Daitch-Mokotoff Soundex (\dm\_soundex.js\) and Double Metaphone (\double\_metaphone.js\) algorithms, in addition to the existing Metaphone and SoundEx. These new processors are exported via the main \index.js\ entry point and have corresponding TypeScript definitions in \index.d.ts\, allowing users to perform more nuanced phonetic matching, particularly for Slavic and Yiddish names with the Daitch-Mokotoff variant.
lib/natural/phonetics · high confidence
Introduction of the Trie data structure with TypeScript support
A new Trie data structure has been added to the library, providing methods for string storage and retrieval including case-sensitive or case-insensitive modes. The implementation exposes methods such as addString, contains, keysWithPrefix, findMatchesOnPath, and findPrefix, along with a getSize utility. TypeScript declaration files have been included to provide type definitions for the Trie class, enabling better development experience for TypeScript users.
lib/natural/trie · high confidence
Multilingual sentiment analysis with multiple vocabulary sources
The sentiment analyzer now supports multiple languages (English, Spanish, Portuguese, Dutch, Italian, French, German, Galician, Catalan, and Basque) by integrating three distinct vocabulary sources: Afinn, Senticon, and Pattern. Users can select the language and the specific vocabulary type (e.g., 'afinn', 'senticon', 'pattern') when instantiating the analyzer to tailor the sentiment analysis to their needs.
lib/natural/sentiment · high confidence
N-gram module adds left/right padding and statistics support
The n-gram generation functions (ngrams, bigrams, trigrams, multrigrams) now accept optional startSymbol and endSymbol arguments to pad sequences with left and right boundary markers, which is useful for language modeling. Additionally, passing stats=true returns an object containing ngram frequencies and counts alongside the generated ngrams, while the default behavior remains unchanged for backward compatibility. A separate Chinese-language variant (NGramsZH) is also provided, which splits input strings into characters instead of using a tokenizer.
lib/natural/ngrams · high confidence
New French Carry stemmer added
A new French stemmer based on the Carry algorithm has been added to the library. This implementation, located in lib/natural/stemmers/Carry, integrates with the existing natural stemmer infrastructure and includes configuration files for transformation steps and utility functions for processing.
lib/natural/stemmers/Carry · high confidence
New Japanese noun inflector added
A new Japanese noun inflector has been added to the library, enabling singularization and pluralization of Japanese nouns. The implementation handles specific suffixes such as -たち, -達, -等, -ども, and -がた, including exception lists for invariant forms and irregular nouns like 神 (kami) and 人 (hito).
lib/natural/inflectors/ja · high confidence
New MaxEnt-based POS tagging components
Added new classes (ME\_Corpus, ME\_Sentence, POS\_Element) to the MaxEnt POS classifier module, enabling corpus splitting for training/testing and defining specific feature generation logic for part-of-speech tagging.
lib/natural/classifiers/maxent/POS · high confidence
New normalizers for English, Japanese, Norwegian, and Swedish
The \lib/natural/normalizers\ module now exposes dedicated normalizers for multiple languages. The English normalizer expands contractions (e.g., 'can't' to 'can not') and handles tokenization. A Japanese normalizer is added with full support for converting between full-width and half-width characters (alphabet, numbers, punctuation, katakana) and between hiragana and katakana. Norwegian and Swedish normalizers are introduced to remove specific diacritic marks (e.g., 'š' to 's' for Norwegian, while preserving 'å' for Swedish). A generic \removeDiacritics\ function is also available for broader character normalization.
lib/natural/normalizers · high confidence
New pluggable storage backend system
A new storage abstraction layer has been added to the library, allowing users to persist data using one of five backends: PostgreSQL, MongoDB, Redis, Memcached, or local file storage. The \StorageBackend\ class in \lib/natural/util/storage\ acts as a unified interface, delegating \store\ and \retrieve\ operations to the specific plugin (e.g., \PostgresPlugin\, \RedisPlugin\) selected at runtime. This change introduces the infrastructure for configurable data persistence, with a \docker-compose.yml\ provided to easily spin up the required database services for development or testing.
lib/natural/util/storage · high confidence
New string distance metrics: Dice Coefficient and Hamming Distance
The \lib/natural/distance\ module now exposes two additional string similarity metrics: Dice Coefficient and Hamming Distance. The Dice Coefficient implementation handles single-character strings by padding them with a space and normalizes input by lowercasing and collapsing whitespace. The Hamming Distance function compares two strings of equal length, optionally ignoring case, and returns -1 if the inputs are not strings or differ in length. These new functions are integrated into the module's public API via \index.js\ and accompanied by TypeScript declarations in \index.d.ts\.
lib/natural/distance · high confidence
New utility modules for graph algorithms and language-specific data
The lib/natural/util area now includes new classes for graph processing—ShortestPathTree, LongestPathTree, EdgeWeightedDigraph, DirectedEdge, Bag, and Topological—along with updated index exports. It also adds language-specific stopword lists (French, Spanish, Farsi, Indonesian, Italian, Japanese, Dutch, Norwegian, Polish, Portuguese) and abbreviation lists for English and Spanish, enabling more accurate text analysis for these languages.
lib/natural/util · high confidence
Stemmer module restructured with TypeScript definitions and multi-language support
The stemmer module has been reorganized to expose a unified API for multiple languages, including English, French, German, Spanish, Italian, Dutch, Norwegian, Portuguese, Swedish, Ukrainian, Russian, Farsi, Japanese, and Indonesian. A new TypeScript declaration file (index.d.ts) defines the shared Stemmer interface and Token class, while index.js centralizes exports for all language-specific implementations. This change provides consistent stemming capabilities across supported languages and improves type safety for TypeScript users.
lib/natural/stemmers · high confidence
Removals
Removal of legacy text processing modules
The library has removed the legacy text processing components: the Bayes classifier, Porter stemmer, stopwords list, and string tokenizers. This eliminates the ability to perform naive Bayesian classification and Porter stemming on text within this module, as the underlying implementation files (\bayes\_classifier.js\, \porter\_stemmer.js\, \stopwords.js\, \tokenizers.js\) have been deleted.
lib · high confidence
Behavioural changes
Brill POS Tagger modernized to ES6 classes with new training and testing capabilities
The Brill POS Tagger implementation in lib/natural/brill\_pos\_tagger/lib has been refactored from legacy code into a modern ES6 class-based architecture. This change introduces dedicated classes for core components: BrillPOSTagger for applying rules, BrillPOSTrainer for deriving transformation rules from corpora, and BrillPOSTester for evaluating accuracy. Supporting classes include Sentence, Corpus, Lexicon, RuleSet, TransformationRule, Predicate, and RuleTemplate, along with a new PEG.js-based parser (TF\_Parser) for reading transformation rule files. The tagger now supports both English and Dutch lexicons and rule sets, and allows for extension of the lexicon. Logging has been simplified to use native console.log statements controlled by a DEBUG flag.
_lib/natural/brill\_pos\tagger/lib · high confidence
Classifiers now delegate to the 'apparatus' library and support parallel training
The Bayes and Logistic Regression classifiers in lib/natural/classifiers have been rewritten to wrap the external 'apparatus' library, replacing the previous internal implementations. This change introduces parallel multi-core training capabilities (trainParallel, trainParallelBatches) for faster model building, while also modernizing the codebase by removing deprecated prototype usage and adding TypeScript declarations.
lib/natural/classifiers · high confidence
Indonesian stemmer reimplemented with strict mode support and JSON dictionary
The Indonesian stemmer in lib/natural/stemmers/indonesian has been replaced with a new implementation that runs in strict mode, fixing previous bugs where prototype-inherited methods caused iteration errors. The dictionary data (kata-dasar.txt) has been converted to a JSON file (data/kata-dasar.json) and loaded via a Set for efficient lookup. The new stemmer supports full stemming algorithms including prefix and suffix rule disambiguation, plural word handling, and stopword management, providing a more robust and performant stemming experience for Indonesian text.
lib/natural/stemmers/indonesian · high confidence
New MaxEnt classifier implementation with deterministic serialization
The MaxEnt classifier module has been modernized with a complete rewrite of its core components (Classifier, Context, Distribution, Element, Feature, FeatureSet, GISScaler, Sample), introducing support for saving and loading trained models to and from JSON files. To ensure reproducible model serialization, the Context class now uses the safe-stable-stringify library for deterministic key generation, replacing the previous json-stable-stringify dependency.
lib/natural/classifiers/maxent · high confidence
New unified entry-point exports for Natural.js modules
The library now provides a consolidated entry point via \lib/natural/index.js\ and its TypeScript declaration \lib/natural/index.d.ts\. This change allows consumers to import the entire Natural.js API (including analyzers, classifiers, distance metrics, inflectors, ngrams, normalizers, phonetics, sentiment analysis, spellcheck, stemmers, TF-IDF, tokenizers, transliterators, tries, utilities, and WordNet) from a single location, rather than importing individual sub-modules separately.
lib/natural · high confidence
Refactored inflector architecture with new base classes and TypeScript support
The inflector module has been restructured to improve code organization and type safety. A new \SingularPluralInflector\ base class (internally named \TenseInflector\ in the source) now handles core pluralization and singularization logic, including case preservation and ambiguous word handling. \NounInflector\ and \PresentVerbInflector\ extend this base class to provide specific linguistic rules. Additionally, a new \CountInflector\ has been added to handle ordinal number suffixes (e.g., 'st', 'nd', 'rd', 'th'). TypeScript declaration files (\index.d.ts\) have been introduced to provide type definitions for these classes, and the module entry point (\index.js\) has been updated to export the new structure.
lib/natural/inflectors · high confidence
Repository modernization and infrastructure overhaul
This change introduces significant structural and operational updates to the project. It adds TypeScript support via a new \tsconfig.json\ and ESLint configuration (\standard-with-typescript\), and sets up an ES Module build pipeline using Rollup. The build system migrates from a Makefile to npm scripts, with a new \.env\ file introduced to configure storage backends (Postgres, Redis, Memcached, MongoDB). Documentation and community standards are enhanced with the addition of \CODE\_OF\_CONDUCT.md\, \CONTRIBUTING.md\, and \SECURITY.md\ (detailing supported versions and vulnerability reporting). The license is updated to MIT, and the README is expanded with badges and detailed licensing information for dependencies like WordNet.
(repo-wide) · high confidence
Spellcheck module restructured with TypeScript definitions and Trie-based implementation
The spellcheck module has been reorganized to support TypeScript and improve performance. A new TypeScript declaration file (index.d.ts) exposes the Spellcheck class interface, while the core implementation (spellcheck.js) now utilizes a Trie data structure for faster word lookups instead of previous methods. The module entry point (index.js) has been updated to export the Spellcheck class from the new implementation file.
lib/natural/spellcheck · high confidence
TfIdf module restructured with TypeScript definitions and caching improvements
The TfIdf module has been reorganized into a dedicated folder with a new entry point (index.js) and TypeScript declarations (index.d.ts). The core implementation now includes an IDF cache to significantly speed up searches on large datasets, and the idf calculation has been corrected to prevent NaN values. Additionally, the module now supports loading documents from files via addFileSync with optional encoding, allows setting custom stopword lists and tokenizers, and includes a new removeDocument method.
lib/natural/tfidf · high confidence
WordNet module refactored with TypeScript definitions and improved data parsing
The WordNet library has been restructured to include TypeScript declarations (index.d.ts) and a cleaner file-based architecture (DataFile, IndexFile, WordNetFile). A key behavioral change is that glosses are now split into separate definition and example fields in the returned data records, and synonym lookups can now be performed directly using synset objects. The module also fixes buffer overruns for large results and ensures proper file handle closure.
lib/natural/wordnet · high confidence
Test coverage
Added IO specifications for classifiers, TfIdf, and WordNet; Added TypeScript unit tests for core NLP components; Added test counting utility and Jasmine configuration; Added test fixtures for NLP processing and dictionary conversion.
Dependencies
Upgrade to version 8.1.1 with modernized dependencies and build tooling
The project has been upgraded to version 8.1.1, introducing a comprehensive update to its dependency tree and build infrastructure. Key runtime dependencies now include Mongoose 9, PostgreSQL driver 8, Redis client 5, and Underscore 1.13, alongside new additions like dotenv, memjs, and safe-stable-stringify. The development environment has been modernized with TypeScript 5, Rollup 4 for ESM builds, and updated testing tools including Jasmine 6 and ESLint 8. This change also establishes a formal build process using npm scripts for compilation and linting, replacing previous manual or ad-hoc methods.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
This is the PUBLIC form of this artifact. Findings are listed in full, but the details of SECURITY findings — which rule fired, in which file, on which line, and how to fix it — are deliberately withheld, and any secret-scanner results are excluded entirely. Where detail is absent here it was REMOVED FOR PUBLICATION; it is not missing from the analysis. The complete artifact is available from the repository owner.
Score
- CAI 54 → 55 (+1.2)
- Rubric changed (rubric-2026.09.12 → rubric-2026.09.18) — scores are not directly comparable.
Lenses
- Code Health 57 → 57 (-0.0)
- Architecture 100 → 88 (-12.2)
- Maturity 46 → 46 (+0.0)
- Readiness 65 → 64 (-1.2)
- Security 56 → 61 (+4.6)
- Performance 100 (new)
Resolved (11)
- Dependency hygiene PARTLY measured — npm pinning read, dependency currency not (no pnpm-resolved versions to grade)
- Documentation: no contributor guidance (README.md)
- Documentation: no installation or build instructions (README.md)
- Documentation: no usage examples (README.md)
- High CVE: [GHSA redacted] (package-lock.json)
- High CVE: [GHSA redacted] (package-lock.json)
- High CVE: [GHSA redacted] (package-lock.json)
- stem (cognitive 19) (lib/natural/stemmers/porter_stemmer_ru.js)
- stem (cognitive 19) (lib/natural/stemmers/porter_stemmer_uk.js)
- stem (cognitive 40) (lib/natural/stemmers/porter_stemmer_it.js)
- stem (cyclomatic 34) (lib/natural/stemmers/porter_stemmer_it.js)
New (16)
- High CVE: [GHSA redacted] (package-lock.json)
- High CVE: [GHSA redacted] (package-lock.json)
- High CVE: [GHSA redacted] (package-lock.json)
- Medium: security finding (details withheld)
- Medium: security finding (details withheld)
- Outdated (npm): dotenv
- Outdated (npm): mongoose
- Outdated (npm): pg
- Outdated (npm): redis
- Outdated (npm): underscore
- Outdated (npm): uuid
- Projects may be oversized for their cohesion
- porter_stemmer_it.PorterStemmer.stem (cognitive 40) (lib/natural/stemmers/porter_stemmer_it.js)
- porter_stemmer_it.PorterStemmer.stem (cyclomatic 34) (lib/natural/stemmers/porter_stemmer_it.js)
- porter_stemmer_ru.PorterStemmer.stem (cognitive 19) (lib/natural/stemmers/porter_stemmer_ru.js)
- porter_stemmer_uk.PorterStemmer.stem (cognitive 19) (lib/natural/stemmers/porter_stemmer_uk.js)
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
NaturalNode/natural was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 2 October 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 69c59f91963045abea5b718430af07f24f01dfb0 — the exact code this score is about.
- Scored under rubric-2026.09.18 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-e569280dd5e2.