spencermountain/compromise
60.0
Adequate · 2 October 2026
60.9k
lines of production code
JavaScript
primary language
2
measurements over time
What this system is
This system is a modular, client-side natural language processing library that tokenizes text and applies hierarchical tagging to identify parts of speech, named entities, and grammatical structures. It provides extensive morphological capabilities, including verb conjugation and noun inflection, alongside advanced features like coreference resolution and sense disambiguation. The architecture supports high-performance processing through parallel workers and lazy parsing, while offering a plugin system for extending functionality with specialized tools for sentiment analysis, date parsing, and data redaction.
How it got here
2013–2021 — Modular architecture and plugin expansion
152 changes.
The project restructured its core NLP engine into a modular, multi-layer plugin system with lazy parsing, supported by a massive expansion of the centralized lexicon. This architectural shift enabled the development of new plugins for paragraph processing, dates, speech pronunciation, and number formatting, alongside significant refactoring and API enhancements.
2022 — advanced NLP features and performance plugins
67 changes.
This period focused on expanding the library's linguistic capabilities by introducing complex features such as coreference resolution, morphological transformations, and structured fact extraction. It also delivered a suite of new plugins for statistical analysis, Wikipedia entity recognition, and high-performance parallel processing, alongside significant improvements to the matching engine and text normalization utilities.
2023–2026 — plugin ecosystem and NLP feature expansion
16 changes.
This period focused on expanding the library's capabilities through a diverse set of new plugins and features, including coreference resolution, sentiment analysis, and morphological transformations. Significant work was also done to enhance data handling with payload attachment, term freezing, and money parsing, alongside the introduction of experimental command-prompt support. The development cycle concluded with the addition of performance benchmarking tools to monitor library stability.
Features
Add Wikipedia plugin for named-entity recognition
Introduces a new \compromise-wikipedia\ plugin that scans text for approximately 38,000 popular Wikipedia article titles. The plugin provides a \wikipedia()\ method on NLP documents to identify these entities, using a compressed lexicon of about 300kb (minified) to enable efficient client-side processing.
plugins/wikipedia · high confidence
Add heuristic verb tense detection based on suffix matching
The pre-tagging system now includes a new heuristic module to detect verb tenses by analyzing word endings. This change introduces a lookup table mapping specific suffixes (such as 'ing', 'ed', 's', and various 2- and 3-character endings) to tense categories like Gerund, PastTense, PresentTense, and Participle. The \getTense\ function checks the last one, two, or three characters of a verb string against this table to return the most likely tense, enabling downstream components to understand conjugation context.
src/2-two/preTagger/methods/transform/verbs/getTense · high confidence
Add number formatting capabilities for text, ordinal, and cardinal representations
This change introduces a new module in \src/3-three/numbers/numbers/format\ that converts numeric values into human-readable text formats. It supports generating text cardinals (e.g., 'five dollars', 'twenty percent'), text ordinals (e.g., 'first', 'twentieth'), and numeric ordinals (e.g., '5th', '20th'). The implementation handles currency symbols (like $, €, £) and percentage signs by appending their word equivalents, manages singular/plural suffixes, and supports decimal numbers (e.g., 'point eight nine') and negative numbers.
src/3-three/numbers/numbers/format · high confidence
Add redact plugin to replace sensitive text with configurable blocks
The redact plugin is now available in the three library, adding a View.prototype.redact method that replaces detected entities (such as people, places, emails, and URLs) with a block string (defaulting to '██████████'). The method supports configuration options to enable or disable specific entity types, allows a custom block string, and includes a 'keep' parameter to control whether the original tags are preserved or removed during the replacement process.
src/3-three/redact · high confidence
Add set operations for JSON pointers
The API now exposes set-based operations on pointers, allowing users to compute the union, intersection, and difference of pointer sets. These methods (exposed as \union\/\and\, \intersection\, and \not\/\difference\) handle overlapping ranges by merging or subtracting segments, and also include a \complement\ method to find parts of the document not covered by the current pointers, and a \settle\ method to remove internal overlaps within a single pointer set.
src/1-one/pointers/api · high confidence
Add speech plugin compute logic for sounds-like and syllable processing
The speech plugin now includes compute modules for phonetic matching and syllable segmentation. The \soundsLike\ module applies a Metaphone algorithm to generate phonetic keys for terms, enabling fuzzy matching based on pronunciation. The \syllables\ module splits text into pronounced syllables using a recursive consonant-vowel analysis with post-processing rules to handle edge cases like silent endings and specific suffixes. These capabilities are exposed via the \plugins/speech/src/compute\ entry point.
plugins/speech/src/compute · high confidence
Add support for detecting and stripping various quotation marks
Users can now identify and remove a wide range of quotation characters, including straight, curly, angle, and prime quotes, via the new \quotations\ method on views. This feature allows for normalizing text by stripping these specific punctuation marks from document content.
src/3-three/misc/quotations · high confidence
Add support for parsing duration values in the dates plugin
The dates plugin now includes a new \durations\ API that allows users to extract and parse duration information from text. This feature supports phrases like '2 months' or '2mins' by matching patterns such as '\#Value+ \#Duration' and handling both multi-word expressions and compact formats like '2mins'. The implementation includes a parser that normalizes various unit abbreviations (e.g., 'min', 'hr', 'wk') and handles plurals, returning structured duration data. This enables downstream components to access duration information directly from date-related views.
plugins/dates/src/api/durations · high confidence
Add swap plugin for morphological transformation of words
The swap plugin now allows users to replace words in text with their morphological variants (e.g., conjugating verbs, pluralizing nouns, converting adjectives to comparative/superlative, or deriving adverbs) based on grammatical tags. This is achieved by registering a new \swap\ method on the View prototype, which internally uses specific handlers for verbs, nouns, adverbs, and adjectives to perform the transformations.
src/2-two/swap · high confidence
Add verb-to-infinitive transformation logic
Introduces a new module for converting conjugated verbs back to their infinitive form. The implementation handles phrasal verbs by separating and reattaching particles and prefixes, supports irregular conjugations via a copula map for forms of 'be', and applies suffix transformations based on the detected tense (past, present, participle, or gerund) using the 'suffix-thumb' library.
src/2-two/preTagger/methods/transform/verbs/toInfinitive · high confidence
Added NLP performance benchmarking script
A new benchmarking script (scripts/bench/index.js) has been added to measure the performance of the NLP library. It runs a suite of tests including tokenization, parsing, syntax matching, entity selection, text transforms, and output generation against a fixed corpus. The script uses performance metrics to calculate scores and stability, helping users and developers track performance regressions or improvements in the library's processing speed.
scripts/bench · high confidence
Added TF-IDF model generation script
A new build script (generate.js) and its helper (pack.js) were added to the stats plugin to generate and package the TF-IDF model data. The script processes a full NLP corpus, computes term frequencies using a minimum word length of 4, and outputs a compressed JSON model file (\_model.js) that the plugin uses for statistical analysis.
plugins/stats/scripts · high confidence
Added TypeScript type definitions for the library
This release introduces a complete set of TypeScript declaration files (\.d.ts\ and \.d.cts\) for the library, covering the core \nlp\ entry points, the \View\ API, and the \Plugin\ interface. Users can now import and use the library in TypeScript projects with full type safety, including support for generic plugin types in the 'three' entry point and detailed definitions for match options, document structures, and view methods.
types · high confidence
Added adverb-to-adjective and adjective-to-adverb conjugation logic
The preTagger now includes specific modules for converting between adverbs and adjectives. The new \fromAdverb.js\ handles stripping suffixes (like '-ly', '-ically') to derive adjectives, including a list of exceptions and words that do not have an adjectival form. The \toAdverb.js\ module adds '-ly' or specific suffixes (like '-ily', '-ically') to adjectives to form adverbs, also respecting exceptions like 'good' to 'well'. A shared \lib.js\ provides the core suffix-matching logic used by both converters.
src/2-two/preTagger/methods/transform/adjectives/conjugate · high confidence
Added browser demo for compromise-dates plugin
A new demo page has been added for the compromise-dates plugin, allowing users to test date extraction functionality directly in the browser. The demo loads the minified plugin build and demonstrates parsing a relative date string ('lets meet in 32 days') to output the resulting start date.
plugins/dates/demo · high confidence
Added compiled builds for the compromise-paragraphs plugin
The \plugins/paragraphs/builds\ directory now contains the compiled distribution files for the \compromise-paragraphs\ plugin, including CommonJS (\compromise-paragraphs.cjs\), ES Module (\compromise-paragraphs.mjs\), and minified (\compromise-paragraphs.min.js\) versions. These builds expose the \paragraphs\ API on the View prototype, allowing users to split text into paragraphs based on double newlines and perform operations like filtering, mapping, and matching on the resulting paragraph views.
plugins/paragraphs/builds · high confidence
Added compromise-stats demo page
A new demo page has been added for the compromise-stats plugin, allowing users to view an example of the plugin in action. The page loads the compiled plugin script and demonstrates how to register the plugin with the nlp instance and use the ngrams() method to process text.
plugins/stats/demo · high confidence
Added compromise-stats plugin for text n-gram analysis
The \plugins/stats/builds\ directory now includes the \compromise-stats\ library (both source and minified builds), enabling users to perform statistical text analysis such as counting n-grams (unigrams, bigrams, etc.) and calculating term frequency-inverse document frequency (TF-IDF) scores. This adds a new capability for analyzing word patterns and importance within text documents.
plugins/stats/builds · high confidence
Added contraction expansion and number suffix models
The model layer for the 'one' contraction module now includes a comprehensive set of contraction mappings and number suffix definitions. The \contractions.js\ file provides a lookup table for expanding informal text (e.g., 'gonna' to 'going to', 'cant' to 'can not') and handling suffix-based contractions (e.g., 'll' to 'will'). Additionally, \number-suffix.js\ defines valid ordinal and temporal suffixes (such as 'st', 'nd', 'rd', 'th', 'am', 'pm') to support number formatting. These models are exported via \index.js\ to be consumed by the rest of the application.
src/1-one/contraction-one/model · high confidence
Added coreference resolution script
A new script at scripts/coreference/index.js has been added to process text for coreference resolution. It loads a corpus, slices a batch of 1,000 documents starting at index 80,000, and uses the NLP library to identify pronouns and map them to their referenced entities, outputting the results to an object.
scripts/coreference · high confidence
Added date and timezone phrase matching rules
The post-tagging model now includes new pattern definitions for recognizing date phrases, weekdays, months, and timezones. These additions enable the system to identify specific date formats (such as '5th of March', 'march 5th', or 'aug 20-21'), relative weekday references (like 'next sun' or 'this sat'), and various timezone indicators (including 'eastern time', 'central european time', and abbreviations like 'est' or 'pst').
src/2-two/postTagger/model/dates · high confidence
Added date-related lexicon files for durations, weekdays, months, and relative dates
New lexicon files have been added to the \data/lexicon/dates\ directory to support date parsing and recognition. \durations.js\ provides terms for time spans (e.g., 'day', 'months', 'qtr'), \weekdays.js\ includes full and abbreviated day names, \months.js\ lists specific month names, and \dates.js\ covers relative terms like 'today' and 'yesterday'.
data/lexicon/dates · high confidence
Added demo pages for speech, speed, and wikipedia plugins
New interactive demo pages have been added for the compromise-speech, compromise-speed, and compromise-wikipedia plugins. These HTML files allow users to test the plugins' capabilities directly in the browser: the speech demo demonstrates syllable counting, the speed demo showcases efficient keypress-based parsing to avoid full re-parsing, and the wikipedia demo illustrates entity recognition for locations.
plugins/speech/demo, plugins/speed/demo, plugins/wikipedia/demo · high confidence
Added dictionary lookup API with trie-based scanning
The lookup API now exposes a \View.prototype.lookup\ method that accepts input strings or objects and returns all matches within the document. This functionality is implemented using a trie structure built via \buildTrie\ and scanned via \scan.js\, which iterates through document terms to find matches based on the trie's transition and failure links, supporting optional form normalization and caching for performance.
src/1-one/lookup/api · high confidence
Added experimental command-prompt parsing plugins for search bangs and slash commands
This change introduces a new experimental plugin (\compromise-cmd-k\) that extends the NLP engine to recognize command-style syntax in text. Users can now detect 'search bangs' (e.g., \!g\, \!gh\) which are tagged as \SearchBang\ and extracted via the new \searchBangs()\ method, as well as 'slash commands' (e.g., \/me\) which are tagged as \SlashCmd\ and extracted via \slashCmds()\. The plugin is designed to be installed via npm and integrated into the main library using \nlp.extend()\.
_plugins/\experiments/cmd-k · high confidence
Added lexicon expansion utility for normalizing and indexing word entries
A new \expand\ method has been introduced to process lexicon key-value pairs, normalizing words by converting them to lowercase, trimming whitespace, and stripping possessive apostrophes (e.g., removing 's). It also identifies and caches multi-word terms to prioritize longer matches. The resulting normalized lexicon and multi-word cache are exported via the \lexicon/methods\ module, enabling downstream components to access pre-processed word data.
src/1-one/lexicon/methods · high confidence
Added morphological transformation data for adjectives and verbs
The \data/pairs\ module now includes new data files that provide mappings for converting adjectives to nouns (\AdjToNoun\), forming comparative and superlative degrees (\Comparative\, \Superlative\), and generating verb gerunds, past tenses, present tenses, and participles (\Gerund\, \PastTense\, \PresentTense\, \Participle\). These additions expand the library's ability to handle various grammatical transformations for both adjectives and verbs, with all new datasets exported via the central \index.js\.
data/pairs · high confidence
Added number lexicon data files
Added new data files for cardinal numbers, ordinal numbers, number multiples, and physical units to support number-related text processing.
data/lexicon/numbers · high confidence
Added organization and place name detection lexicons
The preTagger model now includes dedicated lexicons for identifying organizations and geographical locations. New files \orgWords.js\ and \placeWords.js\ provide lists of keywords (such as 'company', 'university', 'river', 'city') that enable the system to recognize and tag these entity types during text processing.
src/2-two/preTagger/model · high confidence
Added parentheses detection and stripping capabilities
The parentheses module now provides functionality to locate and remove parentheses within documents. The \find\ function scans document terms to identify matching open and close parentheses, returning the indices of matched pairs, while the \strip\ function removes the outermost parentheses from the first and last terms of a match. These utilities are exposed via a \parentheses(n)\ method on the View prototype, allowing users to retrieve specific parenthetical matches, and a \strip()\ method on the resulting Parentheses view instance to remove the detected parentheses from the document.
src/3-three/misc/parentheses · high confidence
Added payload plugin for attaching and retrieving data on text matches
The \plugins/payload/builds\ directory now includes the compiled distribution files (CommonJS, ES Module, and minified) for the \compromise-payload\ plugin. This addition enables users to attach arbitrary data payloads to specific text matches within a compromise document and retrieve or clear them using the new \addPayload\, \getPayloads\, and \clearPayloads\ methods on the View prototype.
plugins/payload/builds · high confidence
Added pre-tagging utilities for lemmatization and Penn Treebank tag conversion
The \src/2-two/preTagger/compute\ module now exposes functions to compute word roots and map internal tags to the Penn Treebank standard. Users can now access \root\ to normalize words to their base form (e.g., converting 'drinks' to 'drink' or 'walked' to 'walk') and \penn\ to retrieve standard POS tags (e.g., mapping 'Adverb' to 'RB') for terms in the document view.
src/2-two/preTagger/compute · high confidence
Added punctuation and case detection methods for match terms
The match system now exposes a set of term methods that allow users to query specific punctuation and casing characteristics of matched terms. These include checks for quotation marks, commas, periods, exclamation marks, question marks, ellipses, semicolons, colons, slashes, hyphens, and dashes, as well as detection of contractions, acronyms, title case, and uppercase text. These capabilities are implemented via the new \termMethods.js\ module and integrated into the \one\ match configuration.
src/1-one/match/methods · high confidence
Added sentence-level caching for keypress parsing
The speed plugin now includes a new keypress module that caches parsed NLP results per sentence to avoid redundant processing. When processing text, the module splits input into sentences and checks a local cache; if a sentence has already been parsed, it reuses the stored JSON data, otherwise it parses the sentence and stores the result. The cache is automatically cleaned up by removing entries that were not used during the current operation, improving performance for repeated or overlapping text inputs.
plugins/speed/src/keypress · high confidence
Added spec-based output parsing and testing utilities
The library now includes \fromSpec\ and \testSpec\ methods to handle spec-formatted output. \fromSpec\ parses text lines containing optional tag blocks (e.g., \{tag1, tag2}\) and comments, stripping them to return clean text. \testSpec\ validates input against expected tags, supporting tag aliases and multi-tag conditions, and reports pass/fail status with optional verbose logging or error throwing. These utilities are exposed via the plugin API and lib interface.
src/1-one/output · high confidence
Added static lexicon data for non-comparable and comparable adjectives
The application now includes two new data files, \adjectives.js\ and \comparables.js\, within the \data/lexicon/adjectives\ directory. These files provide static lists of adjectives categorized by their grammatical behavior: one list for adjectives that do not conjugate into superlative, adverb, or verb forms (e.g., 'ultra', 'specific'), and another for adjectives that do convert to comparative and superlative forms (e.g., 'bad', 'big'). This data supports the lexicon system's ability to handle adjective morphology correctly.
data/lexicon/adjectives · high confidence
Added structured abbreviation dictionaries for tokenization
The tokenization model now includes dedicated abbreviation dictionaries for honorifics, miscellaneous terms, months, nouns, organizations, places, and units. These new files provide specific lists of common abbreviations (such as 'dr', 'st', 'dept', 'kg') to improve the accuracy of text tokenization by recognizing these patterns as distinct units rather than splitting them arbitrarily.
src/1-one/tokenize/model/abbreviations · high confidence
Added structured place lexicon data
The application now includes a new set of lexical data files for places, located in \data/lexicon/places\. This adds structured lists of cities, countries, regions (including US states, Canadian provinces, and other international subdivisions), and general places (such as airports, bodies of water, and notable neighborhoods). This data supports improved place recognition and tagging within the system.
data/lexicon/places · high confidence
Added trie compression and Aho-Corasick builder for lookup API
The lookup API now includes a new \compress.js\ module that prunes redundant trailing values from the Aho-Corasick trie structure (removing undefined, zero, and null tails) to optimize memory usage. Additionally, \index.js\ introduces the core \buildTrie\ function, which constructs the Aho-Corasick automaton from input phrases using tokenization, enabling efficient multi-pattern string matching within the lookup service.
src/1-one/lookup/api/buildTrie · high confidence
Added typeahead prefix auto-fill capability
The typeahead module now includes logic to automatically complete terms based on discovered prefixes. When a user types a term that matches a known prefix, the system fills in the full word, marks it as a machine-generated typeahead entry, and triggers the preTagger to apply lexicon tags. This is implemented via new api, compute, and plugin files that manage the prefix lookup and auto-fill behavior.
src/1-one/typeahead · high confidence
Adds helper methods for selecting specific text patterns
The selections module now provides convenience methods on the View prototype to easily extract specific types of content from text. Users can call methods like \phoneNumbers\, \addresses\, \emails\, \hashTags\, \urls\, and others to retrieve matches for those patterns. The module also supports aliases (e.g., \emojis\ for \emoji\) and allows retrieving the nth match by passing an index argument.
src/3-three/misc/selections · high confidence
Coreference resolution for pronouns in text
The system now resolves pronouns (he, she, they, etc.) to their antecedents within the text. This new coreference engine links pronouns to specific nouns or previous pronouns by searching the current sentence, the immediately preceding sentence, and the one before that. It handles gendered pronouns by filtering for male or female references and supports singular 'they' by matching plural nouns or indefinite pronouns like 'everyone'.
src/3-three/coreference/compute · high confidence
Expanded noun lexicon with categorized data files
The noun lexicon has been reorganized and significantly expanded by introducing dedicated data files for specific noun categories, including actors, demonyms, organizations, possessives, pronouns, proper nouns, singulars, sports teams, and uncountables. This structural change provides a more comprehensive and categorized vocabulary base, improving the system's ability to recognize and process a wider variety of noun types in user input.
data/lexicon/nouns · high confidence
Expanded person lexicon with gendered names and honorifics
The person lexicon has been significantly expanded to improve name recognition and gender detection. New data files introduce distinct lists for male and female first names, a list of ambiguous first names, and a list of common last names. Additionally, a new file adds specific military and professional honorifics (such as 'lieutenant general' and 'field marshal') to the lexicon, while the existing list of famous people and tricky names has been updated.
data/lexicon/people · high confidence
Experimental markdown plugin for parsing and plaintext extraction
Adds a new experimental plugin in \plugins/\_experiments/markdown\ that integrates with the unified/remark ecosystem to parse Markdown into a unified AST. The plugin exposes a \fromMarkdown\ API that converts Markdown strings into a list of plaintext text segments, supporting GFM tables and utilizing utilities for tree traversal and ID generation.
_plugins/\experiments/markdown · high confidence
Experimental sentiment analysis plugin
Adds an experimental, rule-based sentiment analysis plugin for English text. The plugin registers a \.sentiment()\ method on the NLP view, calculating polarity, subjectivity, and intensity scores by analyzing opinion words, emoticons, and emojis. It handles linguistic nuances such as negations and intensifiers, and provides optional mood and summary outputs.
_plugins/\experiments/sentiment · high confidence
Expose noun and verb transformation utilities
The preTagger now exposes dedicated modules for transforming nouns and verbs. The noun module provides functions to convert strings to their plural and singular forms, as well as a combined 'all' method that returns both variants if they differ from the input. The verb module exposes functions to convert verbs to their infinitive form and conjugate them, with a combined 'all' method that returns all conjugated forms except the future tense. These utilities are now available for import and use within the application's tagging logic.
src/2-two/preTagger/methods/transform/nouns, src/2-two/preTagger/methods/transform/verbs · high confidence
Fractions module: new API for parsing and converting fraction representations
This change introduces a new \Fractions\ plugin and supporting utilities in \src/3-three/numbers/fractions\ that enable the system to identify, parse, and convert various fraction formats. The \parse.js\ module handles input normalization for slash notation (e.g., '1/2'), 'out of' phrasing (e.g., '1 out of 4'), ordinal text (e.g., 'three fifths'), and named fractions (e.g., 'half', 'quarter'). The \api.js\ plugin exposes methods to transform these parsed fractions into other representations: \toDecimal\ converts to decimal numbers, \toPercentage\ calculates percentage values, \toCardinal\ produces 'numerator out of denominator' text, and \toOrdinal\/\toText\ generate natural language ordinal forms (e.g., 'one half', 'two thirds').
src/3-three/numbers/fractions · high confidence
Initial release of the Speed plugin
The Speed plugin is introduced as a new capability, aggregating four sub-modules: streamFile, keyPress, workerPool, and lazyParse. This release establishes the plugin's structure and exports its combined library and version (0.1.2) for use.
plugins/speed/src · high confidence
Initial sentence parsing logic for subject, verb, and predicate extraction
This change introduces the core sentence parsing module (\src/3-three/sentences/parse\), which identifies the main clause and extracts the subject, verb, and predicate components. The parser handles complex structures by skipping verbs in relative clauses (e.g., 'the boy who you saw') and correctly identifies tense from the main verb. It also distinguishes main clauses from secondary clauses by recognizing subordinating conjunctions and gerunds only when they appear at the beginning of a clause, ensuring accurate grammatical breakdown for downstream features.
src/3-three/sentences/parse · high confidence
Initial speech plugin with syllable and sounds-like API
The speech plugin is introduced, providing new View methods \syllables\ and \soundsLike\ that aggregate computed linguistic data from documents. This initial implementation exposes the core API surface for speech-related features, relying on a compute module to process the underlying data.
plugins/speech/src · high confidence
Introduce centralized date parsing entry point
A new entry point at \plugins/dates/src/api/parse/one/index.js\ has been added to standardize how date strings are processed. This module orchestrates the parsing pipeline by sequentially invoking tokenization, parsing, and transformation steps, while also handling timezone adjustments and debug logging based on environment configuration.
plugins/dates/src/api/parse/one · high confidence
Introduce chunker plugin entry point
A new plugin entry point has been added to the chunker module, exposing the compute logic and API while registering the 'chunks' hook. This establishes the integration point for the chunking functionality within the system.
src/3-three/chunker · high confidence
Introduce compromise-payload plugin for attaching metadata to matches
Adds the compromise-payload plugin, which allows users to attach arbitrary key-value metadata to specific text matches within a document. Users can use addPayload to store data (such as instrument or height) on matched terms, retrieve it via getPayloads, and clear it with clearPayloads. The plugin stores payloads by sentence number and term position, meaning that removing sentences can shift payload indices, and payloads may persist across different documents sharing the same nlp instance unless explicitly cleared.
plugins/payload · high confidence
Introduce dates and times plugin API
The dates and times plugin is now available, providing methods to parse, format, and filter natural language date and time expressions. Users can access parsed date objects via the \.dates()\ method, which supports filtering by range (\.isBefore()\, \.isAfter()\, \.isSame()\), formatting output, and handling recurring intervals (e.g., 'every tuesday'). The \.times()\ method allows parsing and formatting of time expressions, including conversion to 24-hour format. Both APIs return structured JSON data with start/end times, timezones, and durations, enabling integration with scheduling and calendar features.
plugins/dates/src/api · high confidence
Introduce freeze/unfreeze capability to protect term tags
Users can now freeze terms to prevent destructive tag modifications and unfreeze them to restore mutability. This change adds a new \freeze\ module containing a \compute\ function that marks terms as frozen based on a lexicon, a \debug\ function for visualizing frozen state, and a \plugin\ that registers \freeze\, \unfreeze\, and \isFrozen\ methods on the View API, along with an \@isFrozen\ tag matcher.
src/1-one/freeze · high confidence
Introduce lazy parsing capability in the Speed plugin
The Speed plugin now exposes a new \lazy\ parsing method via its library interface. This feature optimizes processing by tokenizing input first and then filtering sentences to tag only those containing relevant keywords, rather than processing the entire document upfront. The implementation includes a \maybeMatch\ helper that leverages document caching to efficiently identify sentences requiring further analysis.
plugins/speed/src/lazyParse · high confidence
Introduce lexicon plugin with frozen and multi-cache support
The lexicon module now exposes a plugin that initializes a model containing a primary lexicon, a multi-cache, and a separate 'frozenLex' store. The new lib.js implementation allows words to be added to the lexicon, with a specific 'isFrozen' flag that routes entries to the frozenLex store instead of the main lexicon, while still updating the multi-cache. This structure supports both standard lexicon expansion and a distinct frozen state for specific word sets.
src/1-one/lexicon · high confidence
Introduce match module with fuzzy matching and plugin structure
Added a new match module under src/1-one/match that provides fuzzy matching capabilities via a lib.js entry point. The lib.js file exports a parseMatch function that optionally normalizes input by removing unicode characters before delegating to the core parsing logic. A corresponding plugin.js file registers this module alongside API and method handlers, making the fuzzy match feature available as part of the application's plugin system.
src/1-one/match · high confidence
Introduce noun singularization logic with regex rules
The preTagger now includes a new \toSingular\ method for converting plural nouns to their singular forms. This feature uses a dedicated set of regex-based transformation rules (covering cases like 'ies' to 'y', 'ves' to 'f', and irregular plurals) to handle common English inflection patterns, allowing the system to correctly process singular variants of nouns during tagging.
src/2-two/preTagger/methods/transform/nouns/toSingular · high confidence
Introduce pointer repair and document retrieval methods
Added a new \getDoc\ method that retrieves document subsets based on pointers, featuring automatic repair logic to handle mismatched start IDs via blind sweeping and to correct ending pointers when the stored end ID is missing. These capabilities are exposed through the \src/1-one/pointers/methods\ module alongside the existing \termList\ and pointer utility functions.
src/1-one/pointers/methods · high confidence
Introduce postTagger compute module with comma-split and sweep-based tagging
A new compute module for the postTagger has been added, implementing a tagging pipeline that splits documents by commas, builds a match network from model definitions, and applies a sweep operation to update the view. This module also introduces freeze and unfreeze steps to the broader tagger workflow, allowing for controlled processing of lexicon and pre-tagging stages before the final post-tagging sweep.
src/2-two/postTagger/compute · high confidence
Introduce postTagger plugin with confidence API and tagger method
The postTagger module is now available as a plugin, exposing a new confidence API method that calculates the average tagger score across all documents and terms, and a tagger method to re-run the POS-tagger. This provides users with a way to assess tagging quality and trigger re-tagging operations directly through the View prototype.
src/2-two/postTagger · high confidence
Introduce pre-tagging model expansion logic
Added the \\_expand\ module to handle the expansion of the pre-tagging lexicon and model. This includes processing switch terms to resolve ambiguous tags (e.g., Actor\|Verb, Adj\|Gerund) into specific forms, expanding verbs into their conjugated forms (past, present, gerund), inflecting adjectives into comparative and superlative degrees, and generating plurals for nouns. It also integrates support for irregular plurals and uncountable nouns, ensuring the model's lexicon is fully populated with derived forms before tagging.
_src/2-two/preTagger/model/\expand · high confidence
Introduce structured pre-tagging taxonomy for nouns, verbs, values, dates, and misc entities
The pre-tagging system now uses a comprehensive, hierarchical tag set defined in \src/2-two/preTagger/tagSet\. This change adds specific tags for temporal expressions (dates, times, durations), detailed noun categories (persons, places, organizations, pronouns, proper nouns), verb aspects and tenses (present, past, future, imperative, passive, modal), numeric values (cardinals, ordinals, fractions, money), and miscellaneous entities (URLs, emails, hashtags, abbreviations). Users will see more granular and consistent tagging for these entity types, enabling downstream components to distinguish between, for example, a 'Person' and a 'Place', or a 'FutureTense' verb and a 'PastTense' verb, with clear inheritance rules (e.g., 'Month' is a 'Date') and exclusion constraints (e.g., 'Date' is not a 'Verb').
src/2-two/preTagger/tagSet · high confidence
Introduce tag ranking and plugin structure for the 'one' tag system
The 'one' tag module now includes a new tag ranking computation that sorts tags by their number of children in the tag set, prioritizing less common tags. This is exposed via a new plugin structure that wires up the tag set model, compute step, methods, API, and library functions, enabling more sophisticated tag handling within the application.
src/1-one/tag · high confidence
Introduces parallel NLP processing via a worker pool plugin
The speed plugin now includes a new \workerPool\ capability that processes text using Node.js worker threads. This implementation splits input text into chunks using a fast split-and-repair strategy (in \rip.js\) and distributes the work across a pool of workers (created in \pool/create.js\ and executed in \pool/worker.js\) to perform NLP matching in parallel, returning a combined result via a Promise.
plugins/speed/src/workerPool · high confidence
Introduction of preTagger plugin module
A new plugin entry point has been added to the preTagger component, aggregating the core model, methods, compute logic, and tag set definitions. This module registers itself with the 'preTagger' hook, enabling the system to utilize these specific tagging and computation capabilities during the pre-tagging phase.
src/2-two/preTagger · high confidence
Lexicon expansion for morphological variants and multi-word terms
The pre-tagger now automatically expands the lexicon with derived forms during initialization. For nouns, it adds plural forms (including specific handling for 'Actor' and 'Demonym' tags). For adjectives, it generates comparative and superlative variants. Verbs are fully conjugated, with special support for phrasal verbs that conjugates the root verb and preserves the multi-word structure. Number words are expanded to include ordinal, fraction, and cardinal variants. Additionally, the system caches multi-word terms to prioritize longer matches.
src/2-two/preTagger/methods/expand · high confidence
Lexicon-based tagging for single and multi-word terms
The system now applies tags to text terms by looking them up in a defined lexicon. This includes exact matches for single words, support for word aliases, and prefix handling (e.g., 'un-', 're-') for verbs and adjectives. It also supports multi-word term matching (e.g., 'jack rabbit') and includes special logic to tag the second word of phrasal verbs as a particle.
src/1-one/lexicon/compute · high confidence
New 'sweep' capability for bulk matching and tagging
This change introduces a new 'sweep' plugin that adds a \View.prototype.sweep\ method, enabling users to perform bulk matching of a sequence of matches against documents using a compiled 'net'. The method supports optional tagging of results via a \tagger\ option and automatically updates the view with the matched pointers. It also includes a \buildNet\ utility in the library to compile a list of matches into a net for efficient processing.
src/1-one/sweep · high confidence
New 3rd-pass tagging rules for acronyms, organizations, places, and context-aware disambiguation
The pre-tagger now applies a new 3rd-pass analysis to refine entity detection and word sense. It introduces specific logic to identify acronyms (including those without periods and plural forms), detect organizations and places by analyzing title-casing and neighbor words, and handle ambiguous terms via a 'switch' mechanism that considers surrounding context. Additionally, it adds a fallback to tag unclassified words as nouns and infers grammatical details like singular/plural and verb tense for bare tags.
src/2-two/preTagger/compute/tagger/3rd-pass · high confidence
New API methods for computation, iteration, and utility operations
The \src/API/methods\ module now exposes a new set of methods for interacting with the library's view objects. Users can apply metadata or processing steps via the new \compute\ method, which accepts a single method name, an array of names, or a custom function. Iteration over matched results is now supported through \forEach\, \map\, \filter\, \find\, \some\, and \random\, allowing callbacks to process individual term views. Additionally, utility methods have been added for navigation and inspection, including \terms\, \groups\, \eq\, \first\, \last\, \slice\, \all\, \fullSentences\, \none\, \isDoc\, \wordCount\, \isFull\, and \getNth\, providing more granular control over document traversal and data extraction.
src/API/methods · high confidence
New Acronyms view with period formatting and nth selection
A new Acronyms view is introduced, allowing users to retrieve the nth acronym from a document via the \acronyms(n)\ method. This view provides utilities to strip periods from acronym text or add periods between characters, enabling flexible formatting of acronym representations within the document structure.
src/3-three/misc/acronyms · high confidence
New Adjectives plugin for morphological transformations
A new plugin file has been added to the adjectives module, introducing an Adjectives class that extends the View base. This component enables users to retrieve and transform adjectives into their comparative, superlative, adverb, and noun forms via methods like toComparative, toSuperlative, toAdverb, and toNoun. It also provides a json method to export these transformations and exposes new View prototype methods (adjectives, superlatives, comparatives) to select and manipulate these specific grammatical categories.
src/3-three/adjectives · high confidence
New Nouns API for parsing and inflection
The \src/3-three/nouns/api\ module now exposes a dedicated Nouns API that allows users to parse noun phrases and perform grammatical transformations. This includes methods to check if a noun is plural or singular, extract adjectives, and convert nouns between singular and plural forms. The pluralization logic intelligently handles uncountable nouns, proper nouns, and possessives, while also updating associated determiners (e.g., 'a' to 'the') and copulas (e.g., 'is' to 'are') to maintain grammatical correctness.
src/3-three/nouns/api · high confidence
New People topic extraction with gender inference
Added a new 'People' topic that extracts person entities from documents, parsing first names, last names, and honorifics. The feature includes a gender inference module that predicts gender based on names, honorifics (e.g., Mr, Mrs), and pronouns to support co-reference resolution, while also providing methods to filter for presumed male or female individuals.
src/3-three/topics/people · high confidence
New Pronouns API for coreference resolution
A new \Pronouns\ class has been added to the coreference API, exposing methods to resolve pronoun references within the document. Users can now call \hasReference()\ to check if a pronoun has a resolved antecedent, or \refersTo()\ to retrieve the specific noun phrase it points to. The API also provides a \View.prototype.pronouns\ helper to easily instantiate these objects from the document view.
src/3-three/coreference/api · high confidence
New Slashes view component with split functionality
A new Slashes view component has been added to the misc/slashes module. This component extends the base View class and provides a split method that divides text content by forward slashes, replacing the original text with space-separated parts and appending a regex pattern indicating the split structure. Additionally, a slashes helper method is exposed on the View prototype to match and retrieve specific slashed terms by index.
src/3-three/misc/slashes · high confidence
New Verbs API for conjugation and grammatical analysis
This change introduces a new \Verbs\ API class in \src/3-three/verbs/api\ that exposes methods for conjugating verbs (to infinitive, present, past, future, gerund, and past participle forms) and analyzing their grammatical structure (identifying subjects, adverbs, singular/plural status, and negative/positive polarity). The API also provides a \json\ method to retrieve detailed verb metadata, including root form, auxiliaries, and grammatical info, enabling users to programmatically transform and inspect verb phrases within the document.
src/3-three/verbs/api · high confidence
New cache plugin implementation in src/1-one/cache
A new cache plugin has been added to the src/1-one/cache location, providing functionality to cache and uncachecache document data. The plugin exposes API methods (cache, uncache) that modify the view's internal cache state, a compute module for caching logic, and a methods index. This change introduces a new capability for caching within the one-one context.
src/1-one/cache · high confidence
New chunker API for clause and chunk extraction
The chunker module now exposes a new API entry point that allows users to extract syntactic clauses and noun/verb chunks from text. The \clauses\ function splits text based on punctuation (commas, semicolons, dashes), conjunctions, and specific grammatical patterns (e.g., 'said John', 'if...then'), while also handling edge cases like parentheticals and quotations. The \chunks\ function further groups terms into noun phrases, verb phrases, etc., based on clause boundaries. These capabilities are exposed via \View.prototype.chunks\ and \View.prototype.clauses\, returning \Chunks\ objects that support filtering by part of speech (verb, noun, adjective, pivot) and debugging.
src/3-three/chunker/api · high confidence
New compromise plugin for paragraph-level text processing
This change introduces the \compromise-paragraphs\ plugin, enabling users to split and manipulate text at the paragraph level using the \.paragraphs()\ method. The plugin wraps sentence objects to allow paragraph-level operations such as \.text()\, \.json()\, and filtering, while still permitting navigation back to sentences and terms. It is distributed as UMD and ESM builds via Rollup and includes TypeScript definitions for type safety.
plugins/paragraphs · high confidence
New compromise-speed plugin for high-performance NLP
Introduces the compromise-speed plugin, providing three new capabilities for handling large-scale or interactive text processing: workerPool for parallel sentence parsing, streamFile for memory-efficient file processing, and keyPress for caching parsed sentences during rapid input. The package includes TypeScript definitions, a Rollup build configuration, and documentation.
plugins/speed · high confidence
New compromise-speed plugin with worker-based parallel processing
The compromise-speed plugin (v0.1.2) is now available, providing high-performance NLP capabilities through parallel processing. It introduces a \workerPool\ method that distributes text analysis across multiple OS threads using Node.js worker threads, a \streamFile\ function for processing large files via streaming reads, and a \keyPress\ method that caches parsed sentences to optimize repeated operations. The plugin is distributed in CommonJS, ES Module, and minified formats.
plugins/speed/builds · high confidence
New compromise-stats plugin for TF-IDF and N-gram analysis
The plugins/stats directory now contains a new NLP statistics plugin for compromise. Users can analyze text using TF-IDF to identify characteristic words via the .tfidf() method, and extract repeating sub-phrases using n-gram methods such as .ngrams(), .unigrams(), .bigrams(), and .trigrams(). The plugin includes TypeScript definitions, a README with usage examples, and a build configuration for generating UMD and ESM bundles.
plugins/stats · high confidence
New contraction expansion and contraction API
The API now supports bidirectional transformation of contractions. Users can expand contractions (e.g., 'i've' to 'i have') using the new \contractions().expand()\ method, which restores the full text from implicit forms and handles title casing. Conversely, the \contract()\ method applies various contraction rules (such as 'we are' to 'we're', 'going to' to 'gonna', and handling of 'not' negations) to compress text. These capabilities are exposed via \View.prototype.contractions\ and \View.prototype.contract\.
src/2-two/contraction-two/api · high confidence
New date and time lexicon for natural language parsing
The dates plugin now includes a comprehensive lexicon of natural language terms for parsing dates, times, durations, holidays, and timezones. This update adds support for recognizing specific date references (e.g., 'all day'), time durations (e.g., 'hrs', 'qtr'), major global holidays (e.g., 'christmas', 'eid al fitr'), and informal timezone abbreviations (e.g., 'cst', 'bst') mapped to IANA identifiers. It also introduces a false-positive list to prevent ambiguous terms from being incorrectly tagged as timezones without sufficient context.
plugins/dates/src/model/words · high confidence
New date parsing logic for relative, holiday, and explicit dates
The dates plugin now includes a new parsing module in \plugins/dates/src/api/parse/one/02-parse\ that handles relative dates (e.g., 'today', 'yesterday', 'tomorrow', 'next week'), holidays (via \spacetime-holiday\), relative time units (e.g., 'next month', 'last year'), yearly patterns (e.g., 'summer 2002', 'q4 2020'), and explicit calendar dates (e.g., 'June 5th 2019', 'the 21st'). This module orchestrates these specific parsers to interpret natural language date expressions into structured date objects.
plugins/dates/src/api/parse/one/02-parse · high confidence
New dates plugin with debug capabilities
The dates plugin has been restructured and updated to version 3.8.1. It now includes a new debug method that logs parsed date ranges (start and end) to the console with colored formatting, aiding in development and troubleshooting. The plugin's core logic is organized into separate modules for API, computation, tags, words, and regex, and it registers its regex patterns and debug functionality within the world model.
plugins/dates/src · high confidence
New demo pages for performance, plugins, and web workers
Added three new demonstration pages to the demos directory: a performance stress-test page that benchmarks NLP processing speed, a plugin demo showing how to integrate the speech plugin, and a web-worker demo illustrating asynchronous NLP processing in a background thread.
demos · high confidence
New development and build utility scripts
Added a suite of new scripts to support the library's development workflow: \docs.js\ generates machine-readable API and tagset documentation for LLMs; \filesize.js\ checks bundle size growth against the latest npm release; \pack.js\ compresses lexicon and model data for the build; \plugins.js\ runs commands across plugin directories; \version.js\ extracts the version number; and \chunks.js\, \debug.js\, \match.js\, and \match-linter.js\ provide interactive and batch tools for testing, debugging, and linting NLP matches.
scripts · high confidence
New dictionary lookup plugin with compressed trie support
A new plugin at src/1-one/lookup/plugin.js has been added to provide dictionary lookup capabilities. It exposes an API and a library interface that allows users to pre-compile a list of matches into a compressed trie structure via the buildTrie (aliased as compile) function, optimizing subsequent lookups.
src/1-one/lookup · high confidence
New facts extraction plugin for parsing sentence structure
A new 'facts' plugin has been added to the product, introducing a system to extract structured semantic information (subjects, verbs, objects, and modifiers) from text. The implementation includes an API entry point that exposes a \facts()\ method on views, which processes sentences through a series of parsers (for nouns, verbs, adjectives, and pivot words) and a post-processing step that handles subject borrowing across clauses. This allows users to programmatically analyze the grammatical and factual components of statements within the application.
src/4-four/facts · high confidence
New fuzzy matching and comprehensive term matching logic
The term matching system now supports fuzzy matching for text terms using a Damerau-Levenshtein edit distance algorithm, allowing users to match terms even with minor spelling variations when the fuzzy flag is enabled. Additionally, the core matching logic has been expanded to handle a wider variety of term attributes, including matching by ID, tags, methods, pre/post whitespace, regex patterns, chunks, switches, machine forms, and senses, providing more granular control over how terms are identified and matched within the system.
src/1-one/match/methods/match/term · high confidence
New lexicon files for adverbs, conjunctions, currencies, determiners, expressions, and prepositions
Added new data files in data/lexicon/misc to provide explicit lists of adverbs, conjunctions, currencies, determiners, expressions, and prepositions. These files contain static arrays of words and phrases (e.g., 'a lot', 'although', '$', 'the', 'oh', 'in') that support the language processing capabilities of the application.
data/lexicon/misc · high confidence
New misc plugin for text formatting utilities
A new 'misc' plugin has been added to the src/3-three directory, providing a suite of text formatting utilities. This plugin exposes API methods for handling acronyms, parentheses, possessives, quotations, selections, and slashes. Specifically, the possessives module allows users to identify and strip possessive markers (like 's) from text, expanding matches to include associated persons, places, or organizations.
src/3-three/misc · high confidence
New money parsing capability for currency symbols and amounts
This change introduces a new \Money\ view plugin and a currency symbol mapping file, enabling the system to parse monetary values from text. The \currencies.js\ file defines a mapping of 20 currency symbols (such as $, €, £) to their respective currency codes. The \api.js\ file implements a \Money\ class with methods to extract the numeric amount, the currency code (resolving it from symbols if not explicitly named), and structured JSON output, allowing users to query money-related data within documents.
src/3-three/numbers/money · high confidence
New morphological transformation models for verbs and adjectives
The pre-tagging model now includes comprehensive rules for converting between verb forms (past, present, gerund, participle) and adjective degrees (comparative, superlative), as well as converting adjectives to nouns. These transformations are powered by new compressed data files and the 'suffix-thumb' library, enabling the tagger to handle irregular forms and complex suffix patterns for grammatical analysis.
src/2-two/preTagger/model/models · high confidence
New n-gram and TF-IDF statistics capabilities in the stats plugin
The stats plugin now provides new text analysis methods on the View object. Users can extract n-grams (unigrams, bigrams, trigrams, and custom sizes) via \ngrams\, \unigrams\, \bigrams\, and \trigrams\, as well as start/end/edge grams. Additionally, the plugin introduces TF-IDF scoring via \tfidf\ and a method to build custom IDF models via \buildIDF\, allowing for statistical weighting of terms based on frequency and document distribution.
plugins/stats/src · high confidence
New noun post-tagging rules for entities and locations
The post-tagging model now includes new rule sets for identifying specific noun categories. The nouns module adds patterns to disambiguate words like 'more' and 'sense' based on context, and introduces extensive rules for tagging 'Actor' roles (e.g., job titles like 'chief design officer', 'co-founder', and 'dance coach'). The organizations module adds rules to identify entities such as 'University of \[Place\]', 'government of \[Place\]', and sports teams (e.g., '\[Place\] united', '\[Place\] fc'). The places module adds rules for tagging regions (including US state abbreviations), addresses, and disambiguating 'Turkey' as either a country or food based on surrounding verbs.
src/2-two/postTagger/model/nouns · high confidence
New noun-finding logic and plugin entry point
The nouns module now includes a new \find.js\ file that implements the \findNouns\ function, which extracts noun phrases from document clauses by applying a series of split and filter operations (handling commas, expressions, pronouns, determiners, and adjectives). A corresponding \plugin.js\ file has been added to expose this functionality via the module's API.
src/3-three/nouns · high confidence
New number parsing rules for fractions, money, and units
The post-tagger now recognizes a broader range of numeric expressions. It can identify fractions (e.g., 'half a penny', 'a fifth', 'three out of five'), monetary values (e.g., '6 dollars and 5 cents', 'ten bucks'), and various units (e.g., '5 kg', '12 miles per hour', 'twelve percent'). This improves the accuracy of tagging for complex numeric phrases in text.
src/2-two/postTagger/model/numbers · high confidence
New numbers module with parsing, formatting, and unit filtering
A new \numbers\ module has been added to the codebase, introducing a \Numbers\ view class that provides methods to parse, format, and filter numeric values within text. Users can now convert numbers between text and numeric forms (e.g., 'eight' to '8' or '8th' to 'eighth') using methods like \toNumber\, \toText\, \toCardinal\, and \toOrdinal\. The module also supports locale-aware formatting via \toLocaleString\, filtering numbers by value ranges (\greaterThan\, \lessThan\, \between\), and checking for specific measurement units with \isUnit\. Under the hood, it includes a new parser (\find.js\) to correctly segment consecutive number words and a utility (\\_toString.js\) to handle large exponential numbers.
src/3-three/numbers/numbers · high confidence
New numbers plugin aggregates fractions, numbers, and money APIs
A new plugin entry point has been added to the numbers module that consolidates the fractions, numbers, and money sub-APIs into a single interface. Users can now import this plugin to access all three numerical capabilities through one unified registration, simplifying integration with the View system.
src/3-three/numbers · high confidence
New output formats and debugging tools
The API now supports additional text output formats including 'normal', 'machine', 'root', and 'implicit', alongside a new 'spec' format designed for LLM round-tripping. A new 'hash' method (MD5) is available for text, and the 'debug' method has been expanded with color-coded console views for tags, chunks, and highlights, as well as a client-side table view.
src/1-one/output/api · high confidence
New paragraphs plugin for text segmentation
A new plugin has been added to the system that introduces a \paragraphs\ method to the View prototype. This feature allows users to segment text content into paragraphs based on double newline separators, returning a collection of paragraph views that support operations such as extracting text, JSON serialization, regex matching, filtering, and iteration.
plugins/paragraphs/src · high confidence
New payload plugin for attaching and managing data on text matches
A new payload plugin has been added to the system, introducing capabilities to attach arbitrary data to specific text matches within the view. Users can now use the new \addPayload\ method to store static values or function-generated results associated with selected text ranges, and retrieve them via \getPayloads\. The plugin also provides a \clearPayloads\ method to remove stored data for the current selection and includes a \debug\ utility for inspecting stored payloads in the console. This functionality is implemented through a new \plugin.js\ and \debug.js\ file in the \plugins/payload/src\ directory.
plugins/payload/src · high confidence
New pre-tagging methods for plural detection and sentence splitting
The pre-tagging module now includes a \looksPlural\ function to identify plural forms based on suffix rules and exception lists, and a \quickSplit\ function to chunk documents into sentences by intelligently splitting on commas and semicolons while avoiding splits within dates, places, or adjectives. These methods are exported via the \methods/index.js\ entry point to support downstream tagging logic.
src/2-two/preTagger/methods · high confidence
New second-pass tagging rules for years, prefixes, suffixes, and case
The pre-tagger now applies a new set of second-pass rules to refine term classification. It can identify four-digit years (1400–2100) by analyzing surrounding context, such as preceding prepositions or following nouns. It also handles title-case words as proper nouns (excluding specific tags like Date or Month), recognizes Roman numerals, and applies tags based on word prefixes (e.g., 'overwork' inherits 'work's tag) and suffixes. Additionally, it supports query-based switch terms and their prefixed variants (e.g., 'restrike' -\> 'strike').
src/2-two/preTagger/compute/tagger/2nd-pass · high confidence
New sense disambiguation module for verb, noun, and adjective ambiguity
The src/4-four/sense directory introduces a new capability to disambiguate word senses within the text processing pipeline. This module provides a model containing sense definitions for verbs (e.g., 'plug', 'strike'), nouns (e.g., 'chip', 'pitch'), and adjectives (e.g., 'ill', 'cold'), along with a compute engine that resolves the correct sense by analyzing surrounding context words within a 20-word window. The API exposes a \sense\ method to apply these resolved senses to document terms, enabling downstream features to distinguish between multiple meanings of ambiguous words based on their grammatical tag and local linguistic context.
src/4-four/sense · high confidence
New sentence analysis and conjugation API
The \src/3-three/sentences\ module now exposes a new plugin and API for analyzing and transforming sentences. Users can identify sentence types (questions, exclamations, statements) via methods like \isQuestion\, \isExclamation\, and \isStatement\. The API also provides verb conjugation capabilities, allowing sentences to be converted to past, present, future, infinitive, negative, or positive forms. Additionally, a \json\ method is available to extract structured data including subject, verb, predicate, and grammar information for each sentence.
src/3-three/sentences · high confidence
New sentence conjugation utilities for tense and polarity transformations
Added new modules in the conjugation directory to transform sentences into specific grammatical forms. Users can now convert sentences to future, past, present, infinitive, negative, or positive states. The tense converters (toFuture, toPast, toPresent) handle basic conjugation of the first verb and attempt to synchronize subsequent verbs based on linguistic rules (e.g., handling copulas, gerunds, and infinitives). The polarity converters (toNegative, toPositive) modify the first verb's negation status.
src/3-three/sentences/conjugate · high confidence
New speech plugin for pronunciation metadata
The new speech plugin adds pronunciation metadata capabilities to the library. Users can now estimate spoken phonemes and syllables using the \.syllables()\ method, or generate normalized pronunciation strings via \.soundsLike()\. Both methods support direct invocation on document terms or can be combined with other properties using \.compute()\. The plugin is distributed as a separate package (\compromise-speech\) and includes TypeScript definitions for type safety.
plugins/speech · high confidence
New streamFile utility for processing large text files
A new \streamFile\ function has been added to the speed plugin, enabling users to process large text files via Node.js streams. This utility reads a file in chunks, uses the NLP library's sentence tokenizer to split the content, and applies a provided callback function to each segment, allowing for memory-efficient text analysis on large inputs.
plugins/speed/src/stream · high confidence
New sweep matching engine with indexed hook system
The sweep module now uses a new buildNet system that parses match definitions into structured needs and wants, then indexes them into a hooks map for efficient lookup. This replaces the previous matching logic with a system that supports AND/OR operators, optional/negative matches, and fast/slow OR resolution, while caching requirements to optimize sentence processing.
src/1-one/sweep/methods/buildNet · high confidence
New sweep methods module exports
A new module at src/1-one/sweep/methods/index.js has been added to the codebase. This file serves as an entry point that aggregates and exports three core components: buildNet, bulkMatch, and bulkTagger, making them available for import from this location.
src/1-one/sweep/methods · high confidence
New text manipulation API for modifying documents
This release introduces a comprehensive set of methods for modifying text content, including case transformations (toLowerCase, toUpperCase, toTitleCase, toCamelCase), structural changes (insert, remove, replace, concat), and sorting (sort, reverse, unique). Users can now easily alter the casing of matched terms, insert or remove text at specific points, replace content with optional preservation of tags and punctuation, and sort or deduplicate matches. Additional utilities allow for precise control over whitespace and punctuation (pre, post, trim, hyphenate, toQuotations, toParentheses), and a harden/soften mechanism provides more stable pointer management during mutations.
src/1-one/change/api · high confidence
New text manipulation methods: join, split, and lookaround
The match API now includes new methods for restructuring and inspecting text. Users can merge adjacent words using \join\ and \joinIf\, split text around matches with \split\, \splitBefore\, and \splitAfter\, and inspect surrounding context using \before\, \after\, \grow\, \growLeft\, and \growRight\. These additions expand the library's ability to modify and query sentence structure beyond simple matching.
src/1-one/match/api · high confidence
New text manipulation utilities for sorting, inserting, and removing terms
The library now includes three new modules to handle core text editing operations. The \\_sort.js\ module provides comparison functions for sorting matches by alphabetical order, length, word count, sequence, and document frequency. The \insert.js\ module introduces \cleanPrepend\ and \cleanAppend\ functions that insert new terms into a document while managing spacing, title case adjustments, and sentence-ending punctuation. The \remove.js\ module exports \pluckOut\ to delete terms from the document, automatically repairing punctuation and cleaning up empty sentences.
src/1-one/change/api/lib · high confidence
New text normalization plugin with configurable processing levels
A new normalization plugin has been added to the three.js ecosystem, providing a \View.prototype.normalize\ method that applies text processing transformations to document terms. Users can select from three preset levels—light (unicode, punctuation, whitespace, acronyms), medium (adds case, contractions, parentheses, quotations, emoji, honorifics, debullet), or heavy (adds possessives, adverbs, nouns, verbs)—to control the extent of text cleaning, such as lowercasing, removing punctuation, expanding contractions, stripping emojis, and singularizing nouns.
src/3-three/normalize · high confidence
New toNumber parser for converting English number words to numeric values
Added a new \toNumber\ module in \src/3-three/numbers/numbers/parse/toNumber\ that converts English number words into numeric values. The implementation includes a data dictionary for ones, teens, tens, and large multiples (up to septillion), along with logic to handle casual forms (e.g., 'a dozen'), modifiers (e.g., 'half-million', 'negative'), decimals (e.g., 'point five'), and improper fractions. It also supports numeric string cleaning (removing commas, currency symbols, ordinals) and validates input sequences to prevent invalid combinations like 'seven eleven'.
src/3-three/numbers/numbers/parse/toNumber · high confidence
New tokenization compute utilities for aliases, frequency, and indexing
The tokenize/compute module now includes new standalone utilities that enhance token analysis: alias expansion (supporting slash-separated terms and known symbol aliases), machine-normalized text generation (stripping punctuation and normalizing contractions), word frequency counting, character offset tracking, document indexing, and word counting. These functions are exposed via a central index and can be applied to token views to enrich terms with additional metadata for downstream processing.
src/1-one/tokenize/compute · high confidence
New topics API plugin for querying combined entities
A new topics API plugin has been added to the system, introducing a \topics()\ method on the View prototype. This method aggregates results from people, places, and organizations while excluding generic pronouns (such as 'someone', 'man', 'woman', etc.), sorts them by sequence, and supports retrieving specific items via an index. The plugin is registered in \plugin.js\ and includes specific API definitions for organizations in \orgs/api.js\.
src/3-three/topics · high confidence
New verb conjugation and tense detection grammar
The verb parsing module now includes a comprehensive grammar system that detects and classifies verb forms, including simple, progressive, perfect, and passive tenses across present, past, and future contexts, as well as conditional and modal constructions. This change introduces pattern-matching rules for identifying specific conjugations (e.g., 'he walks', 'he is walking', 'he has walked') and integrates cleanup logic to handle adverbs, negatives, and phrasal verb particles before classification, enabling more accurate grammatical analysis of verb phrases.
src/3-three/verbs/api/parse/grammar · high confidence
New verb conjugation methods for future, gerund, negative, participle, past, and present forms
The conjugation API now exposes dedicated methods to transform verbs into specific grammatical forms. \toFuture\ converts verbs to future tense (e.g., 'walk' → 'will walk', 'is walking' → 'will be walking'). \toGerund\ creates the -ing form (e.g., 'walk' → 'is walking'). \toNegative\ applies negation (e.g., 'walks' → 'does not walk', 'is cool' → 'is not cool'). \toParticiple\ generates past participles with auxiliary verbs (e.g., 'walk' → 'has walked'). \toPast\ converts to past tense (e.g., 'walks' → 'walked', 'will walk' → 'walked'). \toPresent\ converts to present tense (e.g., 'walked' → 'walks', 'will walk' → 'walks'). These methods handle various sentence structures including progressive, perfect, passive, and conditional forms.
src/3-three/verbs/api/conjugate · high confidence
New verb lexicon data files added
The system now includes dedicated data files for verb forms, specifically adding lists of infinitives, modals, participles, phrasal verbs, and exceptions to the verb lexicon. This provides the underlying data for verb conjugation and tagging capabilities.
data/lexicon/verbs · high confidence
New verb-finding logic and plugin entry point
The verbs module now includes a new \find.js\ component that implements specific heuristics for locating verb phrases in text, handling cases like conjunctions, prepositions, gerunds, and various tense combinations (past-tense, past-past, auxiliary verbs). A corresponding \plugin.js\ file exposes this functionality via an API, making the verb-finding capability available to the rest of the application.
src/3-three/verbs · high confidence
New whitespace tokenization method for handling punctuation
A new tokenization method located at src/1-one/tokenize/methods/03-whitespace has been added to handle text parsing by stripping leading and trailing punctuation. The implementation normalizes punctuation based on a provided model, preserving specific characters like emoticons, acronyms (e.g., F.B.I.), and number-related symbols (e.g., parentheses in phone numbers, apostrophes in contractions). This change introduces a dedicated module for cleaning up punctuation as whitespace during the tokenization process.
src/1-one/tokenize/methods/03-whitespace · high confidence
Support for natural language date and time ranges
The dates plugin now parses natural language ranges, allowing users to specify start and end points using patterns like '3pm to 4pm', 'january 5th to 7th', 'between march and january', and 'before june'. The parser handles complex scenarios including midnight crossings, month-to-month spans, and relative time bounds, automatically correcting reversed dates and applying inclusive end-dates for date ranges.
plugins/dates/src/api/parse/range · high confidence
Support for parsing multiple discrete dates from natural language ranges
The date plugin now recognizes and parses phrases that indicate multiple separate dates rather than a single continuous range. Users can now input expressions like 'january or march 1999', 'jan 5 or 8', '5 or 8 of jan', or 'june or july 2019', and the system will return a list of distinct start/end date objects for each specified date, correctly handling shared months, years, and conjunctions like 'or' or 'and'.
plugins/dates/src/api/parse/range/combos · high confidence
Tokenize plugin initialization structure
The tokenize module now exposes a plugin object that aggregates compute, methods, and model components and registers hooks for alias, machine, index, and id processing. This establishes the entry point for the tokenize functionality within the application's plugin system.
src/1-one/tokenize · high confidence
Wikipedia plugin data generation scripts added
Added a suite of scripts in \plugins/wikipedia/scripts/generate\ to build the plugin's internal entity model. The process downloads Wikipedia pageview dumps, filters them by language and project settings (using a minimum pageview threshold and a blocklist of generic terms), and compresses the resulting list of entities into a compact model file for use by the plugin.
plugins/wikipedia/scripts · high confidence
Architecture
Refactored pre-tagging clues into modular, composable rule sets
The pre-tagging model's clue definitions have been restructured from a monolithic format into a modular system of specialized modules (e.g., \\_adj.js\, \\_noun.js\, \\_verb.js\) and composite clue files (e.g., \adj-gerund.js\, \person-noun.js\). This change introduces a composable architecture where base lexical and syntactic patterns are defined in individual modules and then combined using \Object.assign\ to handle ambiguous contexts (such as distinguishing between a gerund and an adjective, or a person and a noun). The \index.js\ entry point now exports a unified map of these composite clues, enabling more granular and maintainable configuration of the tagger's behavior for specific word-class interactions.
src/2-two/preTagger/model/clues · high confidence
Behavioural changes
Added colon and hyphen punctuation tagging rules
The pre-tagging system now includes specific rules for handling punctuation in the first pass. A new module detects colons following the first word to tag expressions (e.g., 'edit: foo'), while another module identifies hyphenated word pairs (e.g., 'bone-headed') and tags them as 'Hyphenated'. These additions enhance the granularity of the initial text analysis by recognizing specific punctuation-based patterns.
src/2-two/preTagger/compute/tagger/1st-pass · high confidence
Added compiled distribution files for the speech plugin
The speech plugin now includes pre-built distribution files (CommonJS, ES Module, and minified versions) in the builds directory. This change ensures that the plugin's phonetic processing capabilities, such as syllable counting and 'sounds-like' matching, are available in standard JavaScript formats for direct consumption by applications, rather than requiring end-users to build the plugin source themselves.
plugins/speech/builds · high confidence
Added irregular noun pluralization mappings
The system now supports accurate pluralization for a wide range of irregular nouns through a new lookup table in the pre-tagging model. This includes Latin-derived forms (e.g., criterion→criteria, phenomenon→phenomena), Germanic irregulars (e.g., man→men, mouse→mice), and various other irregular patterns (e.g., leaf→leaves, child→children), ensuring that the noun inflection logic handles these cases correctly instead of applying generic rules.
src/2-two/preTagger/model/irregulars · high confidence
Added pattern testing and manual corpus for verb-phrase conjugation fixes
A new pattern testing tool has been added to the scripts/patterns directory to help validate and debug pattern matching, specifically supporting recent fixes to verb-phrase conjugation. The update includes a new manual.js file containing a curated list of text examples (such as 'u r cool', 'walking is cool', and various date/time phrases) used as a test corpus, alongside a tester.js script that runs these examples against the NLP engine to identify unused or empty patterns. This allows users and developers to verify that the conjugation logic changes are correctly applied across a broader set of inputs.
scripts/patterns · high confidence
Adjective transformation methods now use configurable models
The adjective transformation logic in the preTagger has been refactored to use a centralized \inflect.js\ module that delegates to a \convert\ utility. This change introduces support for configurable transformation models (superlative, comparative, and noun conversion) via the \model.two.models\ configuration object, allowing these linguistic transformations to be customized rather than relying on hardcoded rules.
src/2-two/preTagger/methods/transform/adjectives · high confidence
Automated version file generation in build scripts
The build process now automatically generates a dedicated source file (\_version.js) containing the application version string. This script reads the version from package.json and writes it to the source directory, ensuring the version is available as an importable constant without requiring the entire package.json to be loaded at runtime.
plugins/speed/scripts · high confidence
Cache now includes root terms and tag prefixes in sweep matches
The cache mechanism for document terms has been updated to include additional data points during the sweep process. Specifically, the system now caches the root form of terms (when available), prefixes all term tags with a '\#' character, and continues to cache switch statuses, implicit words, machine terms, and aliases. This change ensures that lookups and matches within the sweep cache can now leverage root forms and structured tag identifiers, improving the comprehensiveness of cached term matching.
src/1-one/cache/methods · high confidence
Centralized transformation method exports
The transform module now provides a unified entry point that aggregates noun, verb, and adjective transformation logic, allowing consumers to import all part-of-speech transformation capabilities from a single location rather than accessing each category individually.
src/2-two/preTagger/methods/transform · high confidence
Enhanced person name disambiguation and phrase recognition
The person tagger now includes new logic to resolve ambiguous names and recognize complex person phrases. The new \ambig-name.js\ module handles cases where words overlap categories, such as distinguishing 'Sydney Harbour' as a place versus 'Sydney' as a person, or identifying 'Will' as a name in 'Will Pharell' versus a modal verb in 'will go'. The \person-phrase.js\ module expands recognition of full names, including titles (e.g., 'Pope Francis'), honorifics (e.g., 'Dr. John'), nicknames (e.g., 'Dwayne "the rock" Johnson'), and international naming conventions (e.g., 'van der', 'bin Laden').
src/2-two/postTagger/model/person · high confidence
Enhanced place entity parsing and new shorthand API
The places topic now includes a new API method, \View.prototype.places(n)\, which allows users to retrieve the nth place entity from the document using a shorthand syntax. Under the hood, the \find\ logic has been updated to improve the accuracy of place detection by intelligently handling comma-separated values; it now correctly splits entities like 'europe, china' while preserving compound locations such as 'paris, france' by recognizing specific city/region/country patterns.
src/3-three/topics/places · high confidence
Expanded adjective and adverb tagging rules for the post-tagger
The post-tagger model in the adjective module now includes new rule sets to improve the accuracy of part-of-speech tagging for adjectives and adverbs. Specifically, \adj-adverb.js\ adds patterns to handle adverbial modifiers of adjectives (e.g., 'dark green'), adverbs following copulas (e.g., 'was still in'), and adverbs in verb phrases (e.g., 'shops direct'). \adj-gerund.js\ introduces rules to distinguish gerunds used as adjectives (e.g., 'amusing', 'looking annoying') from pure gerunds. \adj-noun.js\ adds logic to identify adjectives modifying nouns (e.g., 'her favourite sport') and nouns used adjectivally (e.g., 'brewing giant'). \adj-verb.js\ expands coverage for adjectives following copulas or perception verbs (e.g., 'seem confused', 'felt loved') and handles hyphenated compounds. \adjective.js\ adds rules for hyphenated adjectives (e.g., 'self-driving', 'faith-based') and specific adjective prefixes (e.g., 'un-skilled'). These changes enhance the tagger's ability to correctly classify words in complex syntactic contexts.
src/2-two/postTagger/model/adjective · high confidence
Expanded contraction expansion and number parsing logic
The contraction processing module now handles a broader range of linguistic and numeric patterns. It expands English contractions like 'ain't' to 'is not' and intelligently disambiguates 'd' (did/had/would) based on surrounding context. French contractions (j', l', d') are expanded with basic gender agreement logic. Additionally, the system now splits number-unit combinations (e.g., '4km') into separate number and unit tokens, and expands number ranges (e.g., '5-9pm') into '5 to 9pm' while applying appropriate tags.
src/1-one/contraction-one/compute · high confidence
Expanded lexicon switches for ambiguous word classes
Added new lexicon switch files to handle words that function as multiple parts of speech or entities, improving the accuracy of grammatical tagging. The new data includes actor-verb pairs (e.g., 'coach'), adjective-gerund and noun-gerund combinations, adjective-noun switches, and past/present adjective forms. It also introduces specific mappings for person-related ambiguities, such as person-adj, person-date (e.g., 'April'), person-noun, person-place, and person-verb, alongside unit-noun mappings for measurement units.
data/lexicon/switches · high confidence
Expanded verb tagging rules for imperatives, passives, and auxiliary constructions
The post-tagging model in the verbs module has been significantly expanded with new rule sets for identifying specific verb forms. The \imperative.js\ file introduces patterns to detect commands (e.g., 'do not go', 'please go', 'shut the door') and reflexive constructions. \passive.js\ adds rules to identify passive voice structures using copulas and participles (e.g., 'was being walked', 'had been eaten'). \auxiliary.js\ defines heuristics for auxiliary verbs and modal combinations (e.g., 'will have', 'would be walking', 'about to go'). Additionally, \verb-noun.js\ and \noun-gerund.js\ refine the distinction between verbal and nominal uses of gerunds and infinitives, while \phrasal.js\ and \adj-gerund.js\ handle phrasal verbs and adjective-gerund interactions. These changes improve the accuracy of grammatical tagging for complex sentence structures.
src/2-two/postTagger/model/verbs · high confidence
Improved contraction expansion and possessive detection
The system now more accurately expands contractions like 'ain't', 'how'd', and 'what'd' into their full forms, and better distinguishes between possessive 's (e.g., 'Bob's book') and the contraction 'is' (e.g., 'Bob's here') by analyzing surrounding context and part-of-speech tags.
src/2-two/contraction-two/compute · high confidence
Improved date extraction logic to reduce false positives and handle ranges
The date-finding API now applies stricter filtering to exclude duration-only phrases (like '20 minutes') and specific patterns that previously caused false positives, such as 'one saturday' or monetary values. It also introduces a new splitting mechanism that correctly separates multiple dates in a list (e.g., 'june 5, june 10' or 'tuesday, wednesday') while preserving date ranges introduced by 'between' or 'within'.
plugins/dates/src/api/find · high confidence
Improved text normalization and punctuation handling
The library now better supports periods in email addresses and after quotations, and normalizes text by removing zero-width characters and handling unicode spaces more robustly. Honorifics are stripped during normalization, and possessive forms are correctly processed. These changes improve the accuracy of text cleaning and parsing for common edge cases involving punctuation and special characters.
(repo-wide) · high confidence
Improved tokenization of hyphenated terms, ranges, and slash-separated phrases
The term tokenization logic in the \02-terms\ module has been updated to better handle complex word structures. It now intelligently splits hyphenated words (e.g., 'x-ray', 'aug-20') while preserving specific prefixes and suffixes, combines slash-separated phrases like 'he / she' into single tokens, and merges numeric ranges such as '2 - 5' into '2-5'. These changes result in more accurate and context-aware word splitting for users processing text with these common patterns.
src/1-one/tokenize/methods/02-terms · high confidence
Introduce date parsing unit classes for day, time, week, and year
The date parsing logic in the \plugins/dates\ plugin has been restructured to use a class-based hierarchy for handling different time units. A new \Unit\ base class and specific subclasses (\Day\, \WeekDay\, \Hour\, \Minute\, \Month\, \Year\, etc.) now manage the parsing, shifting, and formatting of date components. This change introduces support for day-of-week relative logic (e.g., 'next Tuesday'), quarter and season handling, and specific time normalization (e.g., defaulting to 10am for 'middle' of a day).
plugins/dates/src/api/parse/one/units · high confidence
Introduce deterministic, collision-resistant term ID generation
The compute module now includes a new \uuid.js\ utility that generates unique and ordered IDs for document terms based on time, sentence index, and term position, rather than relying on Date logic. This change ensures IDs remain stable and do not overflow under high-volume processing (e.g., novels or infinite jest-length texts) by using a base-36 encoding scheme with randomization. The \compute\ module's \id\ function now assigns these IDs to terms that lack them, improving consistency and preventing ID collisions in large documents.
src/1-one/change/compute · high confidence
Introduce multi-stage chunking pipeline with rule-based matching
The chunker in src/3-three/chunker/compute has been refactored from a single-pass approach into a five-stage pipeline (easyMode, byNeighbour, matcher, fallback, fixUp). This change introduces a new matcher stage that uses a defined set of pattern rules (e.g., '\#Copula \#Adverb+? \[\#Adjective\]') to identify chunks, replacing or augmenting previous heuristic logic. The pipeline now processes document terms through these stages sequentially to assign chunk labels (Noun, Verb, Adjective, Pivot), with a fallback mechanism to ensure all terms are chunked and a final fix-up stage to validate verb phrases.
src/3-three/chunker/compute · high confidence
Lexicon data structure and content overhaul
The lexicon data file has been completely replaced with a new generated structure containing compressed, tokenized entries for a wide range of linguistic categories. This update introduces or significantly expands support for complex grammatical forms including comparative and superlative adjectives, various verb tenses (present, past, infinitive, gerund, participle), and specific syntactic roles like actors, expressions, and phrasal verbs. It also enhances entity recognition with detailed lists for proper nouns, organizations, places, and names, alongside numerical and temporal units.
src/2-two/preTagger/model/lexicon · high confidence
Library restructured into modular plugin layers with lazy parsing support
The library has been reorganized into a modular, multi-layer architecture (src/one.js through src/four.js) where core functionality is split across distinct plugin sets: layer one provides basic tokenization, matching, and caching; layer two adds pre/post-taggers, contractions, and a new lazy parsing feature that tokenizes input first and only tags sentences containing matched words to improve efficiency; layer three introduces linguistic features like adjectives, adverbs, chunking, coreference, and normalization; and layer four adds sense and facts plugins. The main entry point (src/nlp.js) now exposes version 14.17.0 and provides access to internal world, model, methods, and hooks, while the lazy parsing capability is exposed via the lib.lazy API for users who want to optimize parsing performance by deferring full tagging until necessary.
src · high confidence
New centralized lexicon structure with switch-based tagging
The data/lexicon module has been restructured to import and aggregate lexical data from organized subdirectories (nouns, verbs, places, etc.) into a single flat index. This change introduces a new 'switches' category for compound tags (e.g., Actor\|Verb, Person\|Adj) and explicitly defines reflexive pronouns and specific grammatical forms (like comparatives and gerunds) in the misc module, enabling more granular linguistic tagging for users.
data/lexicon · high confidence
New date and time recognition logic in the dates plugin
The \plugins/dates/src/compute\ directory has been replaced with a new implementation that introduces specific taggers for years, time ranges, timezones, and post-processing fixups. This change enables the system to recognize and tag a wider variety of temporal expressions, including specific year ranges (e.g., '1998', '2020'), time ranges (e.g., '3-4pm', 'from 9 to 5'), timezone abbreviations (e.g., 'PST', 'UTC-5'), and complex date shifts (e.g., 'three days before'). The new logic also includes cleanup rules to prevent false positives, such as untagging 'march' in 'the soldiers march tomorrow' or 'about' in 'about thanksgiving'.
plugins/dates/src/compute · high confidence
New modular match-syntax parser with fuzzy, root-inflection, and hyphen-splitting support
The match-syntax parser in src/1-one/match/methods/parseMatch has been restructured into a five-step pipeline (parseBlocks, parseToken, splitHyphens, inflectRoot, postProcess) that changes how match expressions are interpreted. Users can now use fuzzy matching by wrapping terms in tildes (e.g., \~word\~), which applies a configurable similarity threshold (default 0.85). Root matching with curly braces (e.g., {walk}) automatically expands to all conjugations of verbs, nouns, and adjectives when the transformation methods are available, falling back to a direct machine lookup if not. Hyphenated words are split into separate tokens unless the first part is a known prefix, and named capture groups are supported using angle brackets (e.g., \<name\>). The parser also supports optional terms (?), greedy quantifiers (+, \*), range quantifiers ({min,max}), and case-sensitive regex matching via the new caseSensitive option.
src/1-one/match/methods/parseMatch · high confidence
New number parsing logic with disabled multiplier suffixes
A new parsing module has been introduced at src/3-three/numbers/numbers/parse/index.js to handle numeric values, including support for comma-separated formats (e.g., '3,123') and fractions. The implementation explicitly disables the multiplication of numbers by 'm' (million) or 'k' (thousand) suffixes, ensuring these characters are treated as literal text rather than multipliers. The parser also strips ordinal suffixes (st, nd, rd, th) and integrates with existing fraction parsing logic.
src/3-three/numbers/numbers/parse · high confidence
New regex-based pre-tagging rules for entities and slang
The pre-tagging model now includes new regular expression patterns to identify specific text entities and informal language. Users will see improved detection for URLs, email addresses, timezones, and various date/time formats (including ISO and local styles). Financial values, phone numbers, and numeric ranges are now tagged more precisely. Additionally, the system can now recognize social media elements like hashtags and mentions, as well as slang expressions, emojis, and specific linguistic patterns such as gerunds and possessives.
src/2-two/preTagger/model/regex · high confidence
New suffix, prefix, and context-based tagging patterns for the pre-tagger
The pre-tagging model now includes new pattern files (endsWith, neighbours, prefixes, suffixes) that enable more accurate part-of-speech and entity tagging based on word endings, surrounding context, and prefixes. Users will see improved detection of past tense verbs, adjectives, nouns, actors, places, and other entities through regex-based suffix matching, neighbor-word context rules, and prefix mappings.
src/2-two/preTagger/model/patterns · high confidence
New tag manipulation API for document terms
The tag API module now exposes methods to manage tags on document terms directly. Users can add tags via \tag()\ (with optional verbose logging and safety checks via \tagSafe()\), remove them with \unTag()\, and filter terms that are compatible with a specific tag using \canBe()\. These operations automatically invalidate caches to ensure consistency.
src/1-one/tag/api · high confidence
New tag validation and formatting pipeline
The tag processing logic in src/1-one/tag/methods/addTags has been replaced with a new modular pipeline. This change introduces validation that supports deprecated 'isA' and 'notA' properties, automatically infers implicit 'is' and 'not' relationships, and ensures bi-directional linking for 'not' tags. It also adds a formatting step that resolves node colors based on type (e.g., Nouns are blue, Dates are red) and consolidates children of excluded tags. The system now distinguishes between user-generated and internal tags to apply specific processing rules.
src/1-one/tag/methods/addTags · high confidence
New text normalization pipeline for tokenization
The \src/1-one/tokenize/compute/normal\ module now implements a structured normalization pipeline that cleans and standardizes input text before tokenization. The new \01-cleanup.js\ step reduces noise by lowercasing, trimming, removing trailing punctuation, coercing Unicode ellipses and dashes to ASCII equivalents, stripping zero-width characters, and removing commas within numbers. The \02-acronyms.js\ step detects and processes various acronym formats (e.g., 'N.D.A', 'NDA', 'c.e.o') by removing periods. The main \index.js\ orchestrates these steps, calling the cleanup and acronym functions, and also integrates with an existing \killUnicode\ method for ASCII transliteration, storing the final normalized string in \term.normal\.
src/1-one/tokenize/compute/normal · high confidence
New tokenization pipeline methods for sentence, term, and whitespace splitting
The tokenization module now exposes a structured set of methods for processing text input. Users can now explicitly invoke functions to split text into sentences, break sentences into terms, and handle whitespace separation. A new 'killUnicode' utility is also available to normalize characters by replacing unicode variants with their ASCII equivalents (e.g., 'Björk' to 'Bjork'). These methods are orchestrated by a new 'parse' function that chains these steps together, normalizing terms as part of the document generation process.
src/1-one/tokenize/methods · high confidence
New verb parsing logic for subjects, adverbs, and phrasal verbs
The \src/3-three/verbs/api/parse\ module has been replaced with a new implementation that provides more granular analysis of verb phrases. The parser now explicitly extracts adverbs into pre- and post-root categories, identifies the grammatical subject (including handling subordinate clauses and pronoun detection), and splits phrasal verbs into their verb and particle components. It also handles contractions, auxiliary verbs, and negation, ensuring a subject is always returned even if inferred from context.
src/3-three/verbs/api/parse · high confidence
Noun pluralization now handles uncountable nouns and irregular forms
The noun inflection logic in the pre-tagging module has been updated to correctly handle uncountable nouns (which remain unchanged) and irregular plurals (looked up from a model list) before applying suffix rules. This change ensures that words like 'sheep' or 'children' are processed according to specific dictionary entries rather than generic suffix patterns, improving the accuracy of pluralization for non-standard cases.
src/2-two/preTagger/methods/transform/nouns/toPlural · high confidence
Post-tagger model restructured into modular rule sets
The post-tagger's rule engine has been refactored from a single monolithic file into a modular structure located in src/2-two/postTagger/model. New dedicated modules now handle specific linguistic patterns: adverb usage (e.g., 'still good', 'way hotter'), conjunctions and prepositions (e.g., 'to the store', 'like the time'), miscellaneous idioms (e.g., 'u r', 'there is'), and conversational expressions (e.g., 'holy shit', 'come on'). The main index.js file orchestrates these modules, applying them in a defined priority order to ensure correct tag assignment for complex phrases.
src/2-two/postTagger/model · high confidence
Refactored NLP core with new View class and plugin system
The API layer has been restructured to use a new \View\ class that manages document pointers and caching, replacing the previous internal structure. This change introduces a more robust plugin system via \extend.js\, which now supports arrays of plugins, irregular verb forms, and custom hooks. Input handling in \inputs.js\ has been updated to support multiple formats including JSON, pre-tokenized arrays, and numbers, while \world.js\ centralizes the model and method definitions.
src/API · high confidence
Refactored date parsing into modular tokenization steps
The date parsing logic in the dates plugin has been restructured from a monolithic implementation into a series of specialized, pure functions. This change introduces distinct modules for handling specific date components: shifts (e.g., '2 weeks ago'), counters (e.g., '7th week'), time parsing (e.g., 'quarter past two'), relative indicators (e.g., 'next Monday'), temporal sections (e.g., 'start of June'), timezones (e.g., 'UTC-5'), and weekdays. The main tokenizer now orchestrates these steps sequentially, allowing for more precise extraction and cleanup of date-related tokens from natural language input.
plugins/dates/src/api/parse/one/01-tokenize · high confidence
Refactored match engine with new modular components and negative conditional support
The match logic in src/1-one/match/methods/match has been restructured into a modular set of files (01-failFast.js, 02-from-here.js, 03-getGroup.js, 03-notIf.js, \_lib.js, and index.js) to improve maintainability and performance. This change introduces a 'failFast' optimization that checks word and tag caches before attempting full matches, and adds support for 'notIf' conditions, allowing users to exclude matches that contain specific patterns. The refactoring also standardizes pointer handling and group extraction, ensuring consistent result structures across different match scenarios.
src/1-one/match/methods/match · high confidence
Refactored pattern matching engine with new step-based logic
The matching engine in the \steps\ directory has been restructured into a modular, step-based architecture to improve handling of complex patterns. This change introduces dedicated handlers for specific matching behaviors, including \and-block\ and \or-block\ for logical grouping, \greedy-match\ and \astrix\ for quantifier support, and \contraction-skip\ to correctly process implicit terms in contractions. Additionally, new logic files (\and-or.js\, \greedy.js\, \negative-greedy.js\) centralize the core algorithms for choice resolution, greedy consumption, and negative lookahead, replacing the previous monolithic implementation.
src/1-one/match/methods/match/steps · high confidence
Refactored sweep matching into modular pipeline with new filtering logic
The sweep matching logic has been restructured into a modular pipeline (getHooks, trimDown, runMatch) to improve maintainability and performance. This change introduces stricter filtering for matches: it now enforces 'needs' (all must be present), 'ifNo' (none must be present), and 'wants' (at least 'minWant' must be present) conditions before attempting actual term matching. It also adds a minimum word count check ('tooSmall') and supports an 'opts.matchOne' flag to stop after the first successful match, optimizing for single-match scenarios.
src/1-one/sweep/methods/sweep · high confidence
Refactored tag management with conflict and dependency handling
The tag handling logic in the 'one' module has been restructured to support complex tag relationships. The new \setTag\ method now automatically removes conflicting tags (defined via 'not' dependencies) and adds parent tags when a known tag is applied. It also supports multi-tag syntax (e.g., '\#Noun . \#Adjective') and respects frozen terms to prevent modification. Conversely, \unTag\ now removes child tags when a parent tag is cleared and also respects frozen terms. A new \canBe\ helper checks if a tag can be applied based on existing conflicts, and \addTags\ is integrated into the main methods index.
src/1-one/tag/methods · high confidence
Restructured pre-tagging pipeline with optimized term tagging
The pre-tagging logic in the tagger module has been reorganized into a three-pass system (first, second, and third) to handle punctuation, individual term properties, and neighbor-based context respectively. A new internal \fastTag\ utility was introduced to efficiently apply tags to terms, including a check to skip processing for frozen terms. The main \preTagger\ entry point now orchestrates these passes, utilizing \quickSplit\ for sentence segmentation and applying specific heuristics for case handling, suffixes, prefixes, years, and organization/place word detection.
src/2-two/preTagger/compute/tagger · high confidence
Rewritten sentence tokenizer with improved internationalization and quote/parenthesis handling
The sentence splitting logic in the tokenization module has been replaced with a new, multi-stage pipeline that significantly improves accuracy for international text and complex punctuation. The new implementation supports CJK full-stops (。!?) and a wide range of Unicode sentence terminators (Devanagari, Arabic, Urdu, Armenian, Ethiopic, Burmese, Khmer) without requiring whitespace, while correctly handling nested brackets in CJK text. It also introduces smarter merging rules to keep embedded quotes and parenthetical remarks attached to their parent sentences, and better filters out non-sentence chunks like acronyms, ellipses, and leading initials.
src/1-one/tokenize/methods/01-sentences · high confidence
Tagger now supports safe tagging, freezing, and untagging
The tagger in the sweep method now handles more complex tagging scenarios. It introduces a 'safe' mode that checks for tag conflicts before applying a tag, preventing inconsistent states. It also supports 'freezing' matches to prevent further modification and allows for 'untagging' existing tags. Additionally, it includes a basic plural detection feature for nouns.
src/1-one/sweep/methods/tagger · high confidence
Tokenization model now includes aliases, punctuation rules, and Unicode transliteration
The tokenization model has been expanded to handle more complex text patterns. It now supports character aliases (e.g., '&' to 'and'), defines specific pre- and post-word punctuation characters (like '@' and '%'), and recognizes common emoticons. Additionally, the model includes lists of valid prefixes (e.g., 'pre-', 'anti-') and suffixes (e.g., '-ish', '-less') for word splitting, and introduces a Unicode transliteration map to convert non-ASCII characters (such as accented letters and fullwidth forms) into their ASCII equivalents.
src/1-one/tokenize/model · high confidence
Typeahead now pre-generates and caches all valid prefixes
The typeahead component now proactively computes all possible prefixes for the provided word list up-front, rather than relying on substring checks at query time. This new behavior ensures that only valid, non-ambiguous prefixes (filtered by a configurable minimum length and optional lexicon safety checks) are registered in the model, improving lookup consistency and performance by avoiding runtime string slicing during user input.
src/1-one/typeahead/lib · high confidence
Updated compromise-dates plugin build
The compiled CommonJS and minified JavaScript builds for the compromise-dates plugin have been regenerated. This update includes the latest parsing logic for date ranges, weekdays, and time zones, ensuring the plugin correctly handles formats like 'between \#Date' and various month/weekday combinations in the distributed bundle.
plugins/dates/builds · high confidence
Upgrade to Compromise 14.17.0 with improved match and text handling
The bundled library has been updated to version 14.17.0. This release includes fixes for the \replaceWith\ method to better preserve surrounding context and punctuation, improvements to the sweep cache to support root matching and partial documents, and optimizations to the match engine for handling optional and negative patterns more efficiently.
builds/three · high confidence
Verb conjugation now supports phrasal verbs and participle validation
The conjugation logic in the preTagger now handles phrasal verbs by parsing and reattaching particles (e.g., 'fall over') to all generated tense forms. It also validates past participles against the lexicon to ensure they are recognized as valid participles or adjectives, while adding a specific exception for the verb 'play'.
src/2-two/preTagger/methods/transform/verbs/conjugate · high confidence
Wikipedia plugin build now includes packed trie data for symbol resolution
The compiled plugin file now contains a new internal encoding system (base-36 alpha codes) and a packed trie data structure used to resolve Wikipedia-style symbols and references. This change introduces the logic to unpack and process these symbols, enabling the plugin to correctly interpret and handle reference links within text.
plugins/wikipedia/builds · high confidence
Wikipedia plugin model updated with new character encoding
The Wikipedia plugin's internal model file has been replaced with a new version that uses a different character encoding scheme. This change is likely intended to improve compatibility with specific text formats or to resolve encoding-related issues in the plugin's data processing.
plugins/wikipedia/src · medium confidence
compromise-dates plugin release 3.9.0
The dates plugin has been updated to version 3.9.0, introducing support for repeating dates (returning a \repeat\ object), nth-weekday parsing (e.g., 'the second monday of february'), two-digit years, and 'quarter to' time formats. This release also includes numerous bug fixes for holiday calculations, overnight ranges, and timezone handling, alongside TypeScript type definitions and a new README.
plugins/dates · high confidence
Fixes
Fix date parsing regex and define date-related tags
This change introduces a new regex configuration file to correctly parse date formats, specifically addressing issues with day and month limits in MM/DD format and supporting the DMY option. It also adds a new tags definition file that establishes semantic categories for date-related entities such as FinancialQuarter, Season, Year, Holiday, and DateShift, ensuring they are correctly identified as Date types while excluding conflicting tags like Fraction or RomanNumeral where appropriate.
plugins/dates/src/model · high confidence
Test coverage
Added TypeScript type-checking tests for compromise package exports; Added comprehensive test coverage for number parsing and manipulation; Added comprehensive test coverage for verb conjugation and parsing; Added comprehensive test suite for contraction handling; Added comprehensive test suite for the 'three' NLP module; Added comprehensive test suite for the tagger; Added test coverage for NLP output formats and spec parsing; Added test coverage for NLP tag matching in gerunds, entities, and verb tenses; Added test coverage for adjective transformations; Added test coverage for build metrics and known tagging issues; Added test coverage for chunker and clause parsing; Added test coverage for clone, replace, and swap transformations; Added test coverage for date tagger ambiguity and chunking; Added test coverage for dictionary lookup functionality; Added test coverage for document manipulation, input handling, and text normalization; Added test coverage for document matching, fuzzy logic, and sweep operations; Added test coverage for named capture groups and multi-match scenarios; Added test coverage for noun inflection and parsing; Added test coverage for number, money, and fraction parsing features; Added test coverage for people detection, parsing, and gender inference; Added test coverage for pointer set operations; Added test coverage for sentence normalization and conjugation; Added test coverage for speech plugin syllable and sounds-like functionality; Added test coverage for text manipulation and structural operations; Added test coverage for text normalization and sentence manipulation; Added test coverage for text normalization options and presets; Added test coverage for the 'one' NLP module's match and miss functionality; Added test coverage for the Wikipedia plugin; Added test coverage for the freeze functionality; Added test coverage for the stats plugin's n-gram functionality; Added test coverage for tokenization edge cases; Added test infrastructure and coverage scripts; Added test suite for the 'two' NLP plugin; Added tests for caching, cache invalidation, and offset computation; Added tests for hash and HTML output formatting; Added tests for lexicon side-loading and apostrophe handling; Added tests for payload attachment and management; Added tests for tag hierarchy and lexicon persistence; Added tests for the paragraphs plugin; Added tests for the streamFile plugin functionality; Comprehensive test suite for the dates plugin; Expanded coreference resolution test coverage; Expanded test coverage for core NLP utilities; Expanded test coverage for the v2 matching engine.
Dependencies
Update core and plugin dependencies
The core compromise package and its plugins have updated their dependencies. The main package now relies on efrt 2.7.0, grad-school 0.0.5, and suffix-thumb 5.0.3, with dev dependencies including rollup 4.63.1, typescript 5.9.3, and eslint 10.10.0. The dates plugin now uses spacetime ^7.12.1 and spacetime-holiday 0.3.0, while the markdown experiment plugin adds several mdast and micromark utilities. Other plugins like stats and wikipedia have updated their efrt dependency to ^2.5.0.
(dependencies) · high confidence
Housekeeping
Bump version to 14.17.0; Updated to version 14.17.0.
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
This is the PUBLIC form of this artifact. Findings are listed in full, but the details of SECURITY findings — which rule fired, in which file, on which line, and how to fix it — are deliberately withheld, and any secret-scanner results are excluded entirely. Where detail is absent here it was REMOVED FOR PUBLICATION; it is not missing from the analysis. The complete artifact is available from the repository owner.
Score
- CAI 58 → 60 (+1.7)
- Rubric changed (rubric-2026.09.12 → rubric-2026.09.18) — scores are not directly comparable.
Lenses
- Code Health 56 → 56 (+0.0)
- Architecture 78 → 80 (+1.5)
- Maturity 64 → 64 (+0.0)
- Readiness 57 → 57 (+0.5)
- Security 60 → 63 (+3.5)
- Performance 100 (new)
Resolved (4)
- Documentation: no installation or build instructions (README.md)
- Documentation: no usage examples (README.md)
- Hotspot: src/1-one/tokenize/methods/03-whitespace/tokenize.js (src/1-one/tokenize/methods/03-whitespace/tokenize.js)
- Hotspot: src/3-three/numbers/numbers/api.js (src/3-three/numbers/numbers/api.js)
New (11)
- High CVE: [GHSA redacted] (pnpm-lock.yaml)
- High CVE: [GHSA redacted] (pnpm-lock.yaml)
- Hotspot: src/1-one/change/api/replace.js (src/1-one/change/api/replace.js)
- Medium: security finding (details withheld)
- Medium: security finding (details withheld)
- Medium: security finding (details withheld)
- No direct assertions: svo main clause after a preposition (tests/three/sentences/svo.test.js)
- Outdated (npm): efrt
- Outdated (npm): grad-school
- Outdated (npm): suffix-thumb
- Projects may be oversized for their cohesion
Changes since last survey
- 3 commits — 3 feature/other, 0 fixes
By area
- (repo) — 3 commits
Notable commits
- change: Merge branch 'master' into dualfroz/relative-clause-subject
- change: Merge pull request #1224 from dualfroz/dualfroz/main-clause-openers
- change: Merge pull request #1225 from dualfroz/dualfroz/relative-clause-subject
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
spencermountain/compromise was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 2 October 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 2a88b3ba57be6985fb82f9aa9cec7a5be9ccc15e — the exact code this score is about.
- Scored under rubric-2026.09.18 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-e569280dd5e2.