explosion/spaCy
63.7
Adequate · 18 September 2026
90.1k
lines of production code
Python
with JavaScript
1
measurement over time
What this system is
This system is a natural language processing library that provides foundational language support, including tokenization, lemmatization, and morphological analysis, for over 60 languages. It offers a modular pipeline architecture for tasks such as named entity recognition, dependency parsing, text classification, and entity linking, alongside rule-based matching capabilities. The system includes a command-line interface for configuration, training, and debugging, as well as visualization tools for linguistic structures.
How it got here
2014–2017 — Global language expansion and CLI rewrite
59 changes.
This period focused on significantly expanding spaCy's multilingual capabilities by adding initial support for numerous languages, including Bengali, Japanese, and Russian, alongside restructuring core language modules for better maintainability. Concurrently, the project overhauled its architecture by rewriting the CLI with Typer, refactoring the tokens subpackage, and migrating the documentation website to Next.js.
2018–2019 — Global language expansion and documentation overhaul
48 changes.
This period focused on significantly broadening spaCy's internationalization by adding initial language support for dozens of new locales, including Persian, Arabic, Vietnamese, and various Asian and European languages. Concurrently, the project overhauled its documentation website with a new modular design system and React components, while also introducing advanced matching capabilities and restructuring internal pipeline architecture.
2020–2022 — language expansion and performance optimization
46 changes.
This period focused on significantly expanding spaCy's multilingual coverage by adding support for numerous new languages, including Basque, Armenian, Malayalam, and Latin, alongside comprehensive tokenizer testing. Concurrently, the codebase underwent major internal refactoring, featuring a modular machine learning architecture, a Cython-based rewrite of parser internals for performance, and the introduction of a new KnowledgeBase module for entity linking.
2023–2025 — Language support expansion and website migration
11 changes.
The project significantly expanded its multilingual capabilities by adding initial support and test coverage for several new languages, including Malay, Norwegian Nynorsk, Scottish Gaelic, Kurdish Kurmanji, Tibetan, and Haitian Creole. Concurrently, the website infrastructure was migrated from Gatsby to Next.js to modernize the documentation platform and improve rendering performance.
Features
Add 'xx' language ID for multi-language support
Introduces a new 'xx' language identifier and a corresponding MultiLanguage class in spaCy, enabling models to support multiple languages. This change includes a new module at spacy/lang/xx with example sentences aggregated from various languages (German, English, Spanish, French, Italian, Dutch, Polish, Portuguese, Russian) to facilitate testing and validation of multi-lingual capabilities.
spacy/lang/xx · high confidence
Add Amharic language support
Introduces initial support for the Amharic language (code 'am') in spaCy. This includes a new language class with tokenizer suffixes, stop words, and specific tokenizer exceptions for abbreviations and compound words, as well as lexical attributes to recognize Amharic numerals and ordinals.
spacy/lang/am · high confidence
Add Ancient Greek (grc) language support
Introduces a new language module for Ancient Greek (ISO code 'grc'), enabling tokenization and basic linguistic processing for this language. The implementation includes a dedicated \AncientGreek\ class with defaults for tokenizer prefixes, suffixes, and infixes, a comprehensive list of stop words, lexical attributes for number recognition, and specific tokenizer exceptions to handle common Ancient Greek contractions (such as preposition+article combinations). Example sentences are also provided for testing.
spacy/lang/grc · high confidence
Add Arabic language support
Introduces a new Arabic language model (\spacy.lang.ar\) for use in spaCy. This includes configuration for right-to-left writing direction, tokenizer suffixes and exceptions (handling Arabic abbreviations and punctuation), stop words, lexical attributes (such as number recognition), and example sentences for testing.
spacy/lang/ar · high confidence
Add Armenian (hy) language support
Introduces initial support for the Armenian language in spaCy. This includes the \Armenian\ language class, a set of Armenian stop words, example sentences for testing, and lexical attributes that enable number recognition (LIKE\_NUM) for both numeric digits and Armenian number words.
spacy/lang/hy · high confidence
Add Azerbaijani language support
spaCy now supports the Azerbaijani language (code 'az'). Users can initialize an Azerbaijani pipeline using \spacy.blank('az')\ or \spacy.load('en\_core\_web\_sm')\-style loading for pre-trained models. This addition includes core language defaults such as stop words, lexical attributes for number recognition (cardinal and ordinal), and example sentences for testing.
spacy/lang/az · high confidence
Add Gujarati language support
Users can now process text in Gujarati using spaCy. This change introduces the Gujarati language class, including a set of stop words and example sentences for testing, enabling basic NLP capabilities for the 'gu' locale.
spacy/lang/gu · high confidence
Add Irish (ga) language support with lemmatization and tokenizer rules
Introduces a new Irish language model (\ga\) to spaCy, including a lookup-based lemmatizer that handles Irish-specific mutations (eclipsis and lenition) and a set of tokenizer exceptions for common contractions and abbreviations.
spacy/lang/ga · high confidence
Add Kurdish Kurmanji language support
Added support for the Kurdish Kurmanji language (code 'kmr') to spaCy. This includes the language class registration, a set of stop words, lexical attributes for recognizing numbers and ordinals, and example sentences for testing.
spacy/lang/kmr · high confidence
Add Kyrgyz (ky) language support
Introduces a new Kyrgyz language subclass in spaCy, providing essential localization data including stopwords, tokenizer exceptions for abbreviations (e.g., days, months, 'etc.'), punctuation rules, and number recognition attributes. This enables users to initialize a Kyrgyz pipeline (\spacy.load('ky')\) for tokenization and basic linguistic processing.
spacy/lang/ky · high confidence
Add Latin language support
Introduces a new Latin language model (code 'la') to spaCy, enabling tokenization, stop-word filtering, and noun-chunk parsing for Latin text. The implementation includes specific tokenizer exceptions for enclitics (e.g., 'mecum') and abbreviations, handles Roman numerals and Latin number words for numeric detection, and provides a set of example sentences for testing.
spacy/lang/la · high confidence
Add Ligurian (lij) language support
Added a new language module for Ligurian (code 'lij'), including tokenizer configuration with specific infixes for elisions, a set of stop words, tokenizer exceptions for common contractions (such as 'co-a', 'da-e', and prefix combinations like 'sott'a-o'), and example sentences for testing.
spacy/lang/lij · high confidence
Add Lithuanian language support
Adds the Lithuanian language module to spaCy, enabling tokenization and basic language processing for the 'lt' locale. This includes a new Lithuanian class with default settings for tokenizer infixes, suffixes, exceptions, stop words, and lexical attributes, along with example sentences for testing.
spacy/lang/lt · high confidence
Add Lower Sorbian language support
Users can now process text in Lower Sorbian (dsb) using spaCy. This change introduces the \LowerSorbian\ language class and its associated defaults, including a stop-word list, lexical attributes for recognizing numbers and ordinals, and example sentences for testing.
spacy/lang/dsb · high confidence
Add Luganda (lg) language support
Added a new language extension for Luganda (ISO code 'lg'), enabling tokenization, stop-word filtering, and number detection for this language. The implementation includes a \Luganda\ class inheriting from spaCy's base language, along with specific configuration for tokenizer infixes, a set of Luganda stop words, and a custom \like\_num\ function that recognizes Luganda number words (e.g., 'emu', 'bbiri') and numeric formats.
spacy/lang/lg · high confidence
Add Macedonian language support
Spacy now includes a built-in Macedonian language model (mk). This addition provides tokenization with specific exception handling for abbreviations, stop words, and lemmatization rules tailored to the Macedonian language.
spacy/lang/mk · high confidence
Add Malay (ms) language support
Introduces a new language module for Malay (ISO code 'ms'), enabling tokenization, stop-word filtering, and noun-chunk detection for this language. The implementation includes specific tokenizer rules for Malay punctuation and units, a comprehensive list of stop words, and examples for testing.
spacy/lang/ms · high confidence
Add Malayalam language support
Introduces initial support for the Malayalam language (code 'ml') in spaCy. This includes the core language class configuration, a set of stop words, number recognition logic for Malayalam numerals, and example sentences for testing.
spacy/lang/ml · high confidence
Add Marathi language support
Users can now initialize a spaCy pipeline for Marathi (language code 'mr') using the new Marathi class. This addition includes a default stop words list sourced from the stopwords-iso project, enabling basic tokenization and stop-word filtering for Marathi text.
spacy/lang/mr · high confidence
Add Persian (Farsi) language support
Introduces a new Persian language module (\spacy.lang.fa\) enabling tokenization, stop-word filtering, and noun-chunk detection for Farsi text. The implementation includes right-to-left writing system metadata, rule-based lemmatization via a generated verb exception list, and tokenizer exceptions to handle common Persian suffixes and clitics.
spacy/lang/fa · high confidence
Add Polish language support with lemmatizer and tokenizer
Introduces a new Polish language class (spacy.lang.pl) that provides tokenizer rules (prefixes, infixes, suffixes), stop words, and lexical attributes (like number detection). It also registers a Polish-specific lemmatizer that uses POS-based lookup tables with custom logic for handling prefixes like 'nie' and 'naj' in verbs and adjectives, and case-sensitive lookups for nouns.
spacy/lang/pl · high confidence
Add Scottish Gaelic (gd) language support
Added support for Scottish Gaelic (language code 'gd') to spaCy, including the core language class, a comprehensive list of stop words, and tokenizer exceptions based on Gaelic Orthographic Conventions and the Annotated Reference Corpus of Scottish Gaelic.
spacy/lang/gd · high confidence
Add Slovak language support
Added support for the Slovak language (code 'sk') to spaCy. This includes the core language class, a set of stop words, lexical attribute definitions for number recognition, and example sentences for testing.
spacy/lang/sk · high confidence
Add Slovenian language support
Added the Slovenian (sl) language module to spaCy, including the base language class, tokenizer configuration (prefixes, suffixes, infixes), stop words, lexical attributes (for numbers, ordinals, and currencies), and tokenizer exceptions for common abbreviations and company types. Example sentences are also provided for testing.
spacy/lang/sl · high confidence
Add Tamil language support
Introduces initial support for the Tamil language (code 'ta') in spaCy. This includes the core language class with default settings for stop words and lexical attributes, a custom \like\_num\ implementation to recognize Tamil numerals and number words, a set of Tamil stop words, and example sentences for testing.
spacy/lang/ta · high confidence
Add Tatar language support
Introduces a new Tatar language model (code 'tt') to spaCy, enabling tokenization, stop-word filtering, and number recognition for Tatar text. This includes configuration for tokenizer infixes and exceptions (such as weekday and month abbreviations), a list of Tatar stopwords, and example sentences for testing.
spacy/lang/tt · high confidence
Add Tibetan (bo) language support
Adds initial support for the Tibetan language (ISO code 'bo') to spaCy. This includes the core language class configuration, a set of stop words sourced from Zenodo, lexical attributes for recognizing Tibetan numerals, and example sentences for testing.
spacy/lang/bo · high confidence
Add Tigrinya language support
Users can now process text in Tigrinya (code 'ti') using spaCy. This change introduces the Tigrinya language class with configuration for tokenization, including specific suffixes, stop words, and tokenizer exceptions for abbreviations. It also provides lexical attributes to recognize Tigrinya number words and ordinals, along with example sentences for testing.
spacy/lang/ti · high confidence
Add Ukrainian language support
Added the Ukrainian language module to spaCy, including the \Ukrainian\ language class, tokenizer exceptions for common abbreviations, stop words, and lexical attributes. The module includes a lemmatizer that defaults to \pymorphy3\ (with \pymorphy2\ support) and provides example sentences for testing.
spacy/lang/uk · high confidence
Add Upper Sorbian language support
Users can now process text in Upper Sorbian (hsb) using spaCy. This change introduces the necessary language-specific components, including a tokenizer with specific exception handling for abbreviations like 'mil.' and 'wob.', a stop-word list, and lexical attributes for recognizing numbers and ordinals in the language.
spacy/lang/hsb · high confidence
Add Urdu language support
Introduces initial support for the Urdu language (ur) in spaCy, including tokenizer suffixes, a stop-word list, and lexical attribute getters for recognizing Urdu numerals and ordinals.
spacy/lang/ur · high confidence
Add Vietnamese language support
Introduces a new Vietnamese language model (\spacy.lang.vi\) that includes a custom tokenizer leveraging the external \pyvi\ library for word segmentation, along with default stop words, lexeme attributes (such as number recognition), and example sentences for testing.
spacy/lang/vi · high confidence
Add alpha support for Tagalog language
Spacy now includes initial support for the Tagalog language (code 'tl'). This addition provides the core language data required for basic processing, including a stop-word list, tokenizer exceptions for common contractions (such as 'tayo'y' and 'isa'y'), and lexical attributes to recognize number words and numeric formats.
spacy/lang/tl · high confidence
Add base language support for Afrikaans, Estonian, Icelandic, and Latvian
Users can now initialize spaCy pipelines for Afrikaans (af), Estonian (et), Icelandic (is), and Latvian (lv). This change introduces the base language classes and corresponding stop-word lists for these four languages, enabling basic tokenization and stop-word filtering capabilities.
(repo-wide) · high confidence
Add basic Telugu language support
Introduces initial support for the Telugu language (code 'te') in spaCy. This includes a new language class with default configurations for lexeme attributes and stop words, a function to recognize Telugu number words and numeric formats, and a set of example sentences for testing.
spacy/lang/te · high confidence
Add initial Albanian language support
Introduces the Albanian (sq) language module, providing the core Language class, a set of Albanian stop words, and example sentences for testing.
spacy/lang/sq · high confidence
Add initial Kannada language support
Introduces basic support for Kannada (code 'kn') in spaCy, including a language class, a set of stop words, and example sentences for testing.
spacy/lang/kn · high confidence
Add initial Setswana (tn) language support
Added the Setswana language module to spaCy, enabling tokenization and basic language-specific rules for the 'tn' locale. This includes configuration for tokenizer infixes, a set of stop words, lexical attributes for number recognition (including Setswana number words and ordinals), and example sentences for testing.
spacy/lang/tn · high confidence
Add initial Yoruba language support
Introduces the Yoruba language (\yo\) to spaCy by adding the \spacy/lang/yo\ module. This includes the core language class, a list of stop words, lexical attribute definitions for number recognition, and example sentences for testing.
spacy/lang/yo · high confidence
Add initial support for the Czech language
spaCy now includes a new language module for Czech (cs), enabling users to load and process Czech text. This addition provides the core language data required for pipeline components, including a list of stop words, lexical attribute getters for number recognition, and a set of example sentences for testing.
spacy/lang/cs · high confidence
Add initial support for the Sinhala language (si)
Introduces basic language data for Sinhala, enabling users to initialize a spaCy pipeline for the 'si' locale. This includes a set of stop words, lexical attribute getters for recognizing number words, and example sentences for testing.
spacy/lang/si · high confidence
Add support for Sanskrit language
Users can now process Sanskrit text using the language code 'sa'. This change introduces the core language configuration, including a stop-word list, lexical attributes for number recognition, and example sentences for testing.
spacy/lang/sa · high confidence
Add support for the Nepali language
Introduces a new language module for Nepali (code 'ne'), enabling users to initialize a spaCy pipeline with \spacy.load('ne\_core\news\\*')\. This includes basic language defaults such as a stop-words list, lexical attribute getters for normalization and number detection, and example sentences for testing.
spacy/lang/ne · high confidence
Added Haitian Creole (ht) language support
Users can now process text in Haitian Creole using the \spacy.lang.ht\ module. This new language package includes a tokenizer with rules for local punctuation and contractions, a stop-word list, a lemmatizer, and noun-chunk detection, enabling basic NLP pipelines for this language.
spacy/lang/ht · high confidence
Added NER and text classification example datasets
New example data files have been added to support Named Entity Recognition (NER) and text classification (textcat) training workflows. The NER examples in \extra/example\_data/ner\_example\_data\ include IOB and JSON formats (with and without POS tags) suitable for conversion via \spacy convert\, accompanied by updated documentation. The text classification examples in \extra/example\_data/textcat\_example\_data\ provide JSON training data for mutually exclusive and multi-label scenarios, sourced from Cooking StackExchange and the Jigsaw Toxic Comments dataset, along with their respective Creative Commons license texts.
_extra/example\data · high confidence
Added Norwegian Nynorsk (nn) language support
Introduced a new language extension for Norwegian Nynorsk (nn), providing the necessary tokenizer rules, punctuation handling, and exception lists to enable tokenization for this language variant. The implementation includes specific configuration for prefixes, infixes, and suffixes, as well as a set of example sentences for testing.
spacy/lang/nn · high confidence
Added quickstart widget generation script
A new Python script (jinja\_to\_js.py) and shell wrapper (setup.sh) have been added to the website/setup directory to compile Jinja2 templates into JavaScript functions. This tool generates the training quickstart widget code, supporting various JavaScript module formats (AMD, CommonJS, ES6) and including specific enhancements like list concatenation and empty dict rendering.
website/setup · high confidence
Added release automation scripts
New shell scripts have been added to the bin directory to streamline the release process. get-package.sh and get-version.sh extract the package name and version from spacy/about.py, while push-tag.sh and release.sh automate the creation and pushing of git tags for releases, ensuring the repository is clean before proceeding.
bin · high confidence
Basque language support added to spaCy
Users can now initialize a Basque language model using the 'eu' code. This change introduces the core language data for Basque, including tokenizer suffixes, a list of stop words, and lexical attributes for recognizing numbers and ordinals, enabling basic tokenization and preprocessing for Basque text.
spacy/lang/eu · high confidence
Initial Bengali language support with rule-based lemmatization
Adds the Bengali (bn) language model to spaCy, enabling tokenization, stop-word filtering, and rule-based lemmatization for Bengali text. The implementation includes specific tokenizer rules for Bengali punctuation and digits, a set of Bengali stop words, and tokenizer exceptions for common abbreviations (such as titles and units) to normalize them to their full forms.
spacy/lang/bn · high confidence
Initial Bulgarian language support added
spaCy now includes a new Bulgarian language model (\spacy.lang.bg\). This addition provides the core language class, configuration, and data required for processing Bulgarian text, including a comprehensive list of stop words, tokenizer exceptions for common abbreviations (such as titles, measurements, and academic fields), and lexical attributes for recognizing number words.
spacy/lang/bg · high confidence
Initial Danish language support added to spaCy
This change introduces the Danish language module (spacy/lang/da), providing the foundational components required to process Danish text. It includes a tokenizer with specific exceptions for common abbreviations (such as months, weekdays, and titles), stop words, punctuation rules, and syntax iterators for noun chunks. Additionally, it implements lexicographical attributes like \like\_num\ to recognize numeric words and ordinals, along with example sentences for testing.
spacy/lang/da · high confidence
Initial Dutch language support with rule-based lemmatization and noun chunking
This change introduces the Dutch (nl) language module to spaCy, providing the foundational components for processing Dutch text. It includes a custom DutchLemmatizer that uses rule-based lookups and exceptions to determine base forms, a noun\_chunks syntax iterator for detecting noun phrases, and a tokenizer configuration with extensive abbreviation exceptions to improve sentence segmentation. The module also provides Dutch-specific stop words, punctuation rules, and example sentences for testing.
spacy/lang/nl · high confidence
Initial Finnish language support with tokenizer and noun chunker
Added the Finnish language module (\spacy.lang.fi\), providing the foundational language data required for processing Finnish text. This includes a custom tokenizer with specific infixes, suffixes, and exception rules for common abbreviations and conjunction contractions (e.g., 'mutt' -\> 'mutta'). The module also introduces a \like\_num\ lexical attribute that recognizes Finnish number words (e.g., 'yksi', 'kaksi') and a dedicated noun chunker to detect base noun phrases based on dependency labels. Standard stop words and example sentences are included to support model testing and usage.
spacy/lang/fi · high confidence
Initial Hebrew language support added
Added the Hebrew language module to spaCy, enabling basic NLP capabilities for Hebrew text. This includes configuration for right-to-left writing direction, a list of Hebrew stop words, lexical attributes for recognizing Hebrew number words and ordinals, and a set of example sentences for testing.
spacy/lang/he · high confidence
Initial Indonesian language support
Added the Indonesian (id) language model to spaCy, including tokenizer rules (prefixes, suffixes, infixes), stop words, and noun chunk detection. The implementation provides specific handling for Indonesian morphology, such as reduplication and affixes, and includes a set of example sentences for testing.
spacy/lang/id · high confidence
Initial Italian language support with POS-aware lemmatization and noun chunks
This change introduces the Italian language module to spaCy, providing the foundational components for processing Italian text. It includes a new POS-aware lemmatizer that uses morphological lookup tables to determine lemmas based on part-of-speech tags, improving accuracy for words with multiple forms. The module also adds Italian-specific tokenizer rules for handling elisions (e.g., 'l'art.') and abbreviations, a comprehensive list of stop words, and a noun chunking syntax iterator to identify base noun phrases. Additionally, example sentences are provided for testing the language model.
spacy/lang/it · high confidence
Initial Japanese language support with SudachiPy tokenizer
Added the \spacy.lang.ja\ module, introducing a new Japanese tokenizer powered by SudachiPy that supports configurable split modes (A, B, C) and serialization. This change provides foundational language data including stop words, syntax iterators for noun chunks, and mappings for part-of-speech tags and morphological features (Inflection, Reading). Users can now process Japanese text with basic tokenization, POS tagging, and lemmatization capabilities.
spacy/lang/ja · high confidence
Initial Korean language support with Mecab-based tokenizer
Adds the Korean language model (spacy.lang.ko) to spaCy, enabling tokenization, part-of-speech tagging, and lemmatization for Korean text. The implementation uses the natto-py library to interface with mecab-ko for morphological analysis, mapping Korean-specific POS tags to Universal Dependencies standards via a provided tag map. It includes default configurations for stop words, punctuation infixes, and number detection, along with example sentences for testing.
spacy/lang/ko · high confidence
Initial Romanian language support added
This change introduces the foundational language data for Romanian (ro) in spaCy. It includes a tokenizer configuration with specific prefixes, suffixes, and infixes tailored for Romanian morphology, a comprehensive list of stop words, and tokenizer exceptions for common abbreviations. Additionally, it provides a custom \like\_num\ implementation to recognize Romanian number words and ordinals, along with example sentences for testing.
spacy/lang/ro · high confidence
Initial Serbian language support added
Added the Serbian language module (code 'sr') to spaCy, providing the foundational components for tokenization and normalization. This includes configuration for tokenizer infixes and suffixes, a set of stop words, and specific tokenizer exceptions for common abbreviations (such as days, months, and titles) and slang to ensure correct normalization. The module also defines lexical attributes, such as recognizing number words, and includes example sentences for testing.
spacy/lang/sr · high confidence
Initial Swedish language support with rule-based lemmatization and noun chunking
Adds the Swedish (sv) language module to spaCy, enabling tokenization, stop words, and syntax iterators for the first time. The module includes a rule-based lemmatizer factory configured as the default, allowing users to access lemmatization without a neural model. It also provides a custom noun chunking implementation based on dependency labels, specific tokenizer exceptions for Swedish abbreviations (e.g., 'jan.', 'mån.') and verbs, and logic to recognize Swedish number words as numeric tokens.
spacy/lang/sv · high confidence
Initial Thai language support with custom tokenizer
Added the Thai language module to spaCy, introducing a new \Thai\ language class and a custom \ThaiTokenizer\ that relies on the PyThaiNLP library for word segmentation. The package includes configuration for Thai-specific defaults, a list of stop words, lexeme attributes for number detection, and a set of tokenizer exceptions for common Thai abbreviations and titles.
spacy/lang/th · high confidence
Initial Turkish language support added to spaCy
This change introduces the Turkish (tr) language model to spaCy, enabling tokenization, lemmatization, and syntactic analysis for Turkish text. The implementation includes a custom tokenizer with rules for handling Turkish-specific abbreviations (such as military ranks and titles), number formats, and inflections. It also provides a language-specific noun chunker that respects Turkish dependency structures, a list of stop words, and logic to recognize cardinal and ordinal numbers (like 'birinci', 'yüzüncü') via the LIKE\_NUM attribute. Example sentences are included for testing purposes.
spacy/lang/tr · high confidence
Initial support for Catalan language processing
This change introduces the first official support for the Catalan language (code 'ca') in spaCy. It adds a complete language module including a custom lemmatizer with rule-based and lookup capabilities, tokenizer configurations for prefixes, suffixes, and infixes, a list of stop words, and syntax iterators for noun chunk detection. Users can now initialize a Catalan pipeline and process text with these baseline linguistic resources.
spacy/lang/ca · high confidence
Initial support for Croatian language
Added the Croatian (hr) language model to spaCy, including the core language class, a list of stop words, example sentences for testing, and a lemma lookup license file.
spacy/lang/hr · high confidence
Initial support for Hindi language processing
Added the Hindi language module to spaCy, enabling tokenization and basic linguistic analysis for Hindi text. This includes a language class with stop words, a stemmer that normalizes word forms, and support for recognizing numeric values (including Indian numbering system terms and ordinals) via the LIKE\_NUM attribute.
spacy/lang/hi · high confidence
Initial support for Luxembourgish (lb)
Adds a new language module for Luxembourgish, enabling basic tokenization and text processing capabilities. The implementation includes configuration for tokenizer infixes and exceptions (handling specific abbreviations and apostrophes), a list of stop words, and rules for identifying numeric text (including Luxembourgish number words). Example sentences are also provided for testing purposes.
spacy/lang/lb · high confidence
Initial support for Norwegian Bokmål (nb)
Added the Norwegian Bokmål language module, providing a tokenizer with language-specific prefixes, suffixes, and infixes, a stop-word list, and a rule-based lemmatizer. The release includes a custom noun-chunk syntax iterator and example sentences for testing.
spacy/lang/nb · high confidence
Initial support for the Greek language
Adds the Greek language model (code 'el') to spaCy, enabling tokenization, lemmatization, and noun chunk detection for Greek text. This includes a rule-based lemmatizer, Greek-specific tokenizer prefixes, suffixes, and infixes, a list of stop words, and syntax iterators for noun phrases. The entry also provides example sentences for testing and a utility script to generate POS data from Wiktionary.
spacy/lang/el · high confidence
Introduce DependencyMatcher and fuzzy matching capabilities
The matcher module now includes a new DependencyMatcher class that allows matching patterns against dependency parse trees using relational operators (such as parent, child, and sibling relationships), expanding rule-based matching beyond linear token sequences. Additionally, fuzzy text matching is enabled in the standard Matcher via a new Levenshtein-based comparison function, allowing patterns to match tokens with slight spelling variations. These changes are implemented in the new \spacy/matcher/dependencymatcher.pyx\ and \spacy/matcher/levenshtein.pyx\ files, alongside updated type stubs and the main matcher implementation.
spacy/matcher · high confidence
Introduce KnowledgeBase and InMemoryLookupKB for entity linking
The \spacy.kb\ module now provides a \KnowledgeBase\ abstract class and an \InMemoryLookupKB\ implementation to support entity linking. This allows the Entity Linker component to resolve named entity mentions to real-world concepts by storing entity identifiers, textual aliases, and prior probabilities. The module exposes \Candidate\ objects and helper functions (\get\_candidates\, \get\_candidates\_batch\) to retrieve potential entity matches for text spans, enabling users to disambiguate mentions using a local in-memory knowledge base.
spacy/kb · high confidence
Introduce displaCy span visualizer and dependency parser enhancements
The displaCy visualization suite now includes a new 'span' style, allowing users to render named entity spans directly within the text flow using the \displacy.render(..., style='span')\ API. For the existing dependency visualizer, users can now enable fine-grained part-of-speech tagging via the \fine\_grained\ option and optionally display lemmas using the \add\_lemma\ option. Additionally, the dependency parser supports noun phrase collapsing through the \collapse\_phrases\ option, which merges noun chunks into single tokens for a cleaner view.
spacy/displacy · high confidence
New Russian language support with pymorphy3 lemmatizer
Adds a new Russian language model (\spacy.lang.ru\) that includes tokenizer exceptions for abbreviations (days, months, academic titles), a comprehensive stop-word list, and numeric word attributes. The default lemmatizer is now \pymorphy3\, replacing the previous \pymorphy2\ implementation, and supports both direct morphological analysis and lookup modes.
spacy/lang/ru · high confidence
New machine learning layer components and NVTX profiling support
This change introduces several new machine learning layers to the spacy.ml module, including CharacterEmbed for character-level embeddings, StaticVectors for projecting vocab vectors with optional dropout, FeatureExtractor for extracting token attributes, extract\_ngrams for n-gram frequency extraction, extract\_spans for retrieving span subsequences, and PrecomputableAffine for efficient affine transformations with padding. It also adds NVTX (NVIDIA Tools Extension) profiling support via callbacks that wrap model nodes and pipe methods (such as pipe, predict, and update) to enable GPU performance tracing, and includes a new parser model framework (TransitionModel) that integrates with the internal C-based parser implementation.
spacy/ml · high confidence
New reusable UI component library for the documentation site
The website now includes a comprehensive set of new React components in the \website/src/components\ directory, providing a unified design system for the documentation. This includes layout primitives like \Grid\, \Main\, and \Landing\ blocks, interactive elements such as \Accordion\, \Dropdown\, \Button\, and \Alert\, and specialized content renderers like \CodeBlock\ (with dynamic loading), \Embed\ (supporting YouTube, SoundCloud, and Google Sheets), and \Juniper\ (for executing Python code in Jupyter kernels). These components replace previous ad-hoc implementations, ensuring consistent styling, accessibility (e.g., ARIA attributes in \Accordion\ and \Icon\), and behavior across the site.
website/src/components · high confidence
New website metadata and configuration files
The website now uses a new set of metadata files in the \website/meta\ directory to drive its configuration and content. This includes \site.json\ for global site settings (domain, theme, navigation, footer), \languages.json\ for supported language data, \sidebars.json\ for documentation sidebar structure, \universe.json\ for the spaCy Universe project listings, and \type-annotations.json\ for API type links. Several TypeScript/JavaScript modules (\dynamicMeta.mjs\, \languageSorted.tsx\, \recordLanguages.tsx\, \recordSections.tsx\, \recordUniverse.tsx\, \sidebarFlat.tsx\) have been added to process and export this metadata for use by the website build system.
website/meta · high confidence
New website widget components for documentation and quickstart guides
The website now includes a suite of new React components in the widgets directory to enhance the documentation experience. These include a Changelog widget that dynamically fetches and displays stable and pre-release versions from the GitHub API, a Features widget that summarizes spaCy's capabilities using data from language metadata, and an Integration widget for displaying partner logos. Additionally, new Quickstart widgets (Install, Models, Training) provide interactive, configurable code generation for installation, model loading, and training configuration based on user selections for OS, hardware, and language. Supporting widgets for Languages, Projects, and a Styleguide for colors and patterns are also added.
website/src/widgets · high confidence
Portuguese language module initialization and configuration
The Portuguese language support in spaCy has been initialized with a new module structure. This includes the definition of the Portuguese language class and its defaults, such as tokenizer exceptions for common abbreviations (e.g., 'Dr.', 'etc.'), punctuation rules, stop words, and syntax iterators for noun chunk detection. Additionally, lexical attributes are configured to recognize numeric words and ordinals, and example sentences are provided for testing purposes.
spacy/lang/pt · high confidence
Architecture
New modular machine learning model architecture in spacy.ml.models
The \spacy/ml/models\ package has been restructured into a new modular layout, introducing dedicated modules for each pipeline component (entity\_linker, multi\_task, parser, span\_finder, spancat, tagger, textcat, tok2vec). This change centralizes the definition of neural network architectures—such as the transition-based parser, span categorizer, and various text classification models—making the underlying model construction logic more explicit and easier to customize or extend.
spacy/ml/models · high confidence
Pipeline components reorganized into spacy/pipeline package
All pipeline components (such as AttributeRuler, DependencyParser, EntityRuler, EntityLinker, and others) have been moved into the new spacy/pipeline subpackage. This change centralizes the component implementations and their factory registrations, providing a cleaner internal structure for the library while maintaining the same public API for users.
spacy/pipeline · high confidence
Refactor spacy.tokens into a structured subpackage
The monolithic tokens module has been split into a dedicated subpackage (spacy/tokens) with separate files for Doc, Token, Span, and serialization logic. This reorganization introduces new public types including DocBin for efficient binary serialization, SpanGroups for managing arbitrary span annotations, and MorphAnalysis for morphological data. The change also adds a Retokenizer context manager to handle document modifications like merging and splitting tokens, and updates the serialization format to use gzipped msgpack for better performance and smaller file sizes.
spacy/tokens · high confidence
Behavioural changes
Centralized global language data and tokenizer rules
The \spacy/lang\ module now provides a unified set of global configuration files for tokenization and lexical attributes. This includes \char\_classes.py\ for defining Unicode character ranges (e.g., Latin, CJK, Cyrillic), \lex\_attrs.py\ for global attribute functions like \like\_url\ and \like\_num\, \norm\_exceptions.py\ for normalizing punctuation and currency symbols, \punctuation.py\ for default tokenizer prefixes/suffixes/infices, and \tokenizer\_exceptions.py\ for base exceptions and URL patterns. This change consolidates language-specific logic into shared defaults that individual language models can extend or override.
spacy/lang · high confidence
Chinese language module restructured with configurable tokenizers and scoring
The Chinese language module has been reorganized to support multiple tokenization strategies (character-based, jieba, and pkuseg) via a new \ChineseTokenizer\ class and a standardized configuration system. Users can now select their preferred segmenter and, when using pkuseg, configure custom user dictionaries and models. The module also introduces tokenization scoring capabilities for evaluation and includes updated lexical attributes for number recognition and a comprehensive stop-word list.
spacy/lang/zh · high confidence
Converters now yield documents via generators
The training data converters (CoNLL NER, CoNLL-U, IOB, and JSON) have been refactored to return generators instead of lists. This means that when converting training files into spaCy Doc objects, the system now yields documents one by one rather than loading the entire dataset into memory at once, which improves memory efficiency for large training sets.
spacy/training/converters · high confidence
Expanded quickstart templates with new pipeline components and language-specific transformer recommendations
The quickstart configuration generator now supports additional pipeline components, including spancat, spancat\_singlelabel, trainable\_lemmatizer, entity\_linker, and span\_finder, allowing users to generate configs for these tasks directly. The template logic has been updated to omit tok2vec/transformer layers for TextCatBOW when efficiency is optimized, and to use core language models as vectors by default. Additionally, the recommendations file has been updated with specific transformer model names for a wider range of languages (e.g., Arabic, Bulgarian, Bangla, Catalan, Danish, etc.) to improve out-of-the-box performance for those locales.
spacy/cli/templates · high confidence
German language module restructured with improved tokenization and noun chunking
The German language module has been reorganized to provide more accurate tokenization and syntactic analysis. Tokenization now includes specific rules for German contractions (such as 'auf'm' expanding to 'auf dem') and a comprehensive list of abbreviation norms (e.g., 'z.B.' to 'zum Beispiel'). A new, language-specific noun chunk iterator has been added to better detect base noun phrases, addressing previous issues with overlapping chunks. Additionally, the module now includes a curated set of German stop words and example sentences for testing.
spacy/lang/de · high confidence
Hungarian tokenizer rules and data restructured
The Hungarian language module has been reorganized to improve tokenizer accuracy and maintainability. Tokenization rules (prefixes, suffixes, and infixes) are now defined in a dedicated \punctuation.py\ file, with specific adjustments such as excluding the degree symbol (°) from icon concatenation to preserve tokens like '99°' and removing the percent sign (%) from unit suffixes to fix percentage tokenization. A comprehensive list of tokenizer exceptions for abbreviations and special forms is provided in \tokenizer\_exceptions.py\, and stop words are managed in \stop\_words.py\. Additionally, example sentences for testing have been added to \examples.py\.
spacy/lang/hu · high confidence
Introduce rule-based lemmatizer and comprehensive language data for French
The French language module now includes a dedicated rule-based lemmatizer (\FrenchLemmatizer\) that handles lemmatization for nouns, verbs, adjectives, and other parts of speech using lookup tables and suffix-stripping rules, falling back to a lookup table for out-of-vocabulary words. This change also adds essential language data including a comprehensive list of tokenizer exceptions for hyphenated words and abbreviations, updated stop words, punctuation rules for French-specific elisions and hyphens, and a \like\_num\ implementation to recognize numeric words and ordinals. Additionally, example sentences and syntax iterators for noun chunks are provided to support testing and parsing.
spacy/lang/fr · high confidence
New Cython-based parser and NER internals
The parser and NER components have been rewritten in Cython to improve performance and maintainability. This change introduces a new internal module \spacy.pipeline.\_parser\_internals\ containing core state management (\\_state\, \stateclass\), transition systems (\transition\_system\, \arc\_eager\, \ner\), beam search utilities (\\_beam\_utils\), and projectivization logic (\nonproj\). The new implementation uses C-level data structures for faster state transitions and memory management, and includes specific optimizations such as cycle detection during projectivization and constant-time head lookups.
_spacy/pipeline/\_parser\internals · high confidence
New modular SASS design system for the documentation website
The website's styling has been refactored from a monolithic approach into a modular component-based system using SASS modules. This change introduces a comprehensive set of new stylesheets (e.g., \accordion.module.sass\, \alert.module.sass\, \code.module.sass\, \layout.sass\) that define the visual appearance of UI elements like navigation, sidebars, code blocks, and landing page cards. It also establishes a unified design token system in \layout.sass\ for colors, fonts, and spacing, and adds specific styles for features like Algolia DocSearch integration and responsive table scrolling.
website/src/styles · high confidence
Project CLI commands now delegate to Weasel
The \spacy project\ subcommands (assets, clone, document, dvc, pull, push, remote\_storage, and run) no longer contain their own implementation logic. Instead, they now import and re-export all functionality from the \weasel.cli\ module. This means that the behavior, arguments, and output of these project management commands are now determined by the Weasel library rather than spaCy's internal code.
spacy/cli/project · high confidence
Refactored English language data and lemmatizer into modular components
The English language module has been restructured to improve maintainability and performance. Language-specific data such as tokenizer exceptions, stop words, syntax iterators, and punctuation rules are now defined in separate, dedicated modules (e.g., \tokenizer\_exceptions.py\, \stop\_words.py\) rather than being embedded in the main language class. A new \EnglishLemmatizer\ component has been introduced, which uses a dedicated \is\_base\_form\ logic to optimize lemmatization by skipping tokens that are already in their base form. Additionally, the \lex\_attrs.py\ module now includes enhanced logic for detecting numeric words and ordinals, improving the accuracy of the \LIKE\_NUM\ attribute for English text.
spacy/lang/en · high confidence
Refactored training alignment and data augmentation infrastructure
The training module has been restructured to improve performance and flexibility. Alignment logic has been moved to optimized Cython implementations (align.pyx, alignment\_array.pyx) using a simplified ragged array type for faster indexing, and is now exposed via a new Python Alignment class. Data augmentation capabilities have been expanded with new combined, whitespace, and orth-variant augmenters in augment.py, allowing for more sophisticated training data generation. Additionally, the training loop and corpus readers have been updated to support these new augmentation strategies and improved alignment caching.
spacy/training · high confidence
Spanish language module restructured with rule-based lemmatizer and improved tokenization
The Spanish language module has been reorganized into a modular structure, introducing a new rule-based lemmatizer that uses morphological features to determine lemmas, improving accuracy for Spanish text. Tokenization has been enhanced with updated infixes, suffixes, and exception rules to better handle Spanish-specific patterns like abbreviations and time formats. Additionally, the module now includes example sentences for testing, a refined stop word list, and improved noun chunk detection logic to prevent overlapping phrases.
spacy/lang/es · high confidence
Updated system architecture diagram
The public-facing architecture diagram in the documentation has been updated to reflect the current system design, ensuring that visual representations of components and data flows match the actual implementation.
website/public · high confidence
Website migrated from Gatsby to Next.js
The website has been rebuilt using Next.js, replacing the previous Gatsby-based architecture. This migration introduces a new project structure with Next-specific configuration files (next.config.mjs, tsconfig.json), a new ESLint setup extending Next's core web vitals, and a Dockerfile tailored for the Next.js environment. The site now utilizes MDX for content, supported by custom remark plugins for code block handling, find-and-replace logic, and section wrapping. Additionally, the new build includes PWA capabilities via next-pwa, Netlify-specific optimizations for image caching and headers, and a standardized Node 18 runtime environment.
website, website/src/templates · high confidence
Website pages migrated to Next.js
The website's page structure has been rewritten using Next.js, replacing the previous Gatsby-based implementation. This change introduces Next.js-specific routing patterns, including dynamic catch-all routes for documentation pages (\[...listPathPage\].tsx) and static generation for models and universe categories. The migration also includes a new custom \_app.tsx for global layout and analytics (Plausible), a \_document.tsx for HTML structure, and dedicated pages for the 404 error, landing, and specific content sections, all leveraging Next.js's getStaticPaths and getStaticProps for pre-rendering.
website/pages · high confidence
spaCy CLI rewritten with Typer and new subcommand structure
The command-line interface has been completely rewritten using the Typer framework, replacing the previous implementation. This introduces a new hierarchical command structure with dedicated subcommand groups for debugging (debug config, debug data, debug diff, debug model), benchmarking (benchmark speed), and initialization (init config, init pipeline). The rewrite also adds new capabilities such as the \spacy apply\ command for running inference on documents, the \spacy assemble\ command for building pipelines from config files, and the \spacy find-threshold\ command for tuning multi-label classifier thresholds. Additionally, the CLI now supports config overrides via environment variables and the \--opt=value\ syntax, and model symlinks are deprecated in favor of loading packages by full name or directory path.
spacy/cli · high confidence
Fixes
Fix model downloading in environments without pip on PATH
Resolves an issue where the \spacy download\ command failed in environments where the \pip\ executable is not available on the system PATH but is present as a Python module (such as in certain virtual environments or containers). The command now correctly invokes pip via the Python module interface to ensure model downloads succeed in these constrained setups.
(repo-wide) · high confidence
Test coverage
Added German language test suite; Added Hungarian tokenizer tests; Added Italian language test suite; Added Serbian (sr) tokenizer and exception tests; Added comprehensive English tokenizer and language-specific test suite; Added comprehensive serialization test suite; Added comprehensive test suite for Danish language processing; Added comprehensive test suite for parser and NER internals; Added language-specific test suite for initialization, lemmatizers, and attributes; Added test coverage for Arabic tokenizer; Added test coverage for Catalan tokenizer; Added test coverage for Dutch language processing; Added test coverage for Finnish language processing; Added test coverage for Greek language tokenization; Added test coverage for Indonesian language tokenizer and lexer; Added test coverage for Korean language processing; Added test coverage for Latin language support; Added test coverage for Luxembourgish (lb) tokenizer; Added test coverage for Russian tokenizer, lemmatizer, and text attributes; Added test coverage for Slovenian language tokenization; Added test coverage for Spanish and Norwegian tokenization and noun chunking; Added test coverage for Swedish (sv) language processing; Added test coverage for Thai language tokenizer serialization and tokenization; Added test coverage for Tigrinya tokenizer; Added test coverage for Vietnamese tokenizer serialization and tokenization; Added test suite for Japanese language processing; Added test suite for Turkish language processing; Added test suite for Ukrainian language processing; Added test suite for the Matcher module; Added test suite for training components; Added test suite for vocab, vectors, and lookups components; Added test to verify consistency of package dependencies; Added tests for Ancient Greek (grc) tokenizer and text processing; Added tests for Armenian language tokenizer and lex attributes; Added tests for Basque language tokenizer; Added tests for Bulgarian tokenizer and text attributes; Added tests for Chinese tokenizer serialization, tokenization, and configuration; Added tests for Czech number-like token detection; Added tests for Gujarati tokenizer; Added tests for Haitian Creole (ht) language support; Added tests for Hebrew tokenizer and lexical attributes; Added tests for Irish (ga) tokenizer exception handling; Added tests for Kurdish Kurmanji language support; Added tests for Lithuanian tokenizer behavior; Added tests for Malayalam tokenizer; Added tests for Morphology serialization and feature conversion; Added tests for Nepali language tokenizer; Added tests for Persian noun chunking behavior; Added tests for Polish tokenizer and text attributes; Added tests for Sanskrit tokenizer behavior; Added tests for Tatar language tokenizer; Added tests for Urdu tokenizer behavior; Added tokenizer tests for Afrikaans, Croatian, Icelandic, Latvian, and Slovak; Added tokenizer tests for Albanian (sq) and multilingual (xx) languages; Added tokenizer tests for Bengali (bn); Added tokenizer tests for Faroese and Norwegian Nynorsk; Added tokenizer tests for Kyrgyz (ky); Added unit tests for French tokenizer and noun chunking; Comprehensive test suite for Doc, Span, and Token behavior; Comprehensive test suite for pipeline components; Expanded tokenizer test coverage for edge cases and debugging; Initial test suite structure and language-specific coverage.
Dependencies
Migrate from Pydantic v1 to v2
The library has upgraded its data validation and serialization engine from Pydantic v1 to v2. This change requires users to have Pydantic version 2.0.0 or higher installed. As part of this migration, the internal configuration schemas and model metadata validation now use the Pydantic v2 API, which may affect custom components or scripts that rely on the previous Pydantic v1 internals.
spacy · high confidence
Website migration to Next.js and updated Python dependencies
The documentation website has been migrated from Gatsby to Next.js, introducing new JavaScript dependencies such as Next 13.0.2, React 18.2.0, and associated tooling in the website package. On the Python side, the project has updated its core dependencies to require Thinc \>=8.3.12, Pydantic \>=2.0.0, and Numpy \>=2.0.0, while also adding Confection and updating Weasel to version 1.0.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Baseline
- First survey — no prior run to compare against. CAI 64.
Lenses
- Code Health 97
- Architecture 100
- Maturity 69
- Readiness 52
- Security 81
- Domain Modelling 100
- Accessibility 66
Changes since last survey
- 300 commits — 248 feature/other, 52 fixes
By area
- (root) — 66 commits
- .github/workflows — 36 commits
- website/docs — 32 commits
- spacy/about.py — 28 commits
- website/meta — 21 commits
- spacy/cli — 20 commits
- spacy/tests — 20 commits
- (repo) — 18 commits
- spacy/lang — 15 commits
- spacy/language.py — 8 commits
- spacy/pipeline — 8 commits
- spacy/tokens — 4 commits
- website/src — 4 commits
- .github/FUNDING.yml — 3 commits
- spacy/init.py — 2 commits
- spacy/displacy — 2 commits
- spacy/strings.pyx — 2 commits
- website/netlify.toml — 2 commits
- .github/spacy_universe_alert.py — 1 commit
- bin/release.sh — 1 commit
Notable commits
- fix: Fix 'issue template' link in CONTRIBUTING.md (#13587) [ci skip]
- fix: Fix --require-parent default
- fix: Fix CI (#13469)
- fix: Fix CI: bump mypy pin for numpy 2.5 stubs, sync confection pin
- fix: Fix LLM docs on task factories.
- fix: Fix allocation of non-transient strings in StringStore (#13713)
- fix: Fix build directory [ci skip]
- fix: Fix cdef declaration for cython 3
- fix: Fix click dependency issue (#13971) (#13973)
- fix: Fix dependencies
- fix: Fix displacy span stacking (#13068)
- fix: Fix environment variable for test
- fix: Fix import sorting for ruff isort compliance
- fix: Fix inverted cli arg
- fix: Fix landing banner links [ci skip]
- fix: Fix matrix in tests
- fix: Fix memory zones
- fix: Fix misspelling (#13631) [ci skip]
- fix: Fix numpy constant
- fix: Fix numpy floats in meta.json
- …and 280 more
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
explosion/spaCy was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 18 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 26b4d1dc04a812f426e4bef3e8a1b6f159d6f048 — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-5d04157a340d.