google/langextract
70.1
Strong · 26 September 2026
26.2k
lines of production code
Python
primary language
4
measurements over time
What this system is
LangExtract is a Python library designed to extract structured information from text using large language models. It provides a modular architecture with a plugin-based provider system that supports backends like Gemini, OpenAI, and Ollama, allowing for flexible model selection and custom integrations. The system includes tools for benchmarking extraction quality, validating output schemas, and generating structured data from various text sources.
Features
Add Ollama integration example with Docker support and timeout configuration
Users can now run LangExtract with local LLMs via Ollama using the new \examples/ollama\ directory. This includes a comprehensive demo script (\demo\_ollama.py\) that showcases extraction capabilities, a \Dockerfile\ and \docker-compose.yml\ for containerized execution with health checks, and updated documentation explaining how to configure timeout settings for slower models or larger prompts.
examples/ollama · high confidence
Developer tooling and documentation updates for v1.7.0
This release introduces a pre-commit configuration (\.pre-commit-config.yaml\) integrating isort and pyink for consistent code formatting, alongside a comprehensive \.pylintrc\ to standardize linting rules. The project now includes a \CITATION.cff\ file with a Zenodo DOI for proper academic citation, a \COMMUNITY\_PROVIDERS.md\ registry listing third-party plugin integrations (such as AWS Bedrock, LiteLLM, and vLLM), and a production Dockerfile. Additionally, the \README.md\ has been updated to reflect the new default model (\gemini-3.5-flash\), add sections for custom providers and OpenAI usage, and include a live demo link.
(repo-wide) · high confidence
New Agent Skill for LangExtract usage
A new 'langextract-usage' Agent Skill has been added to the skills directory, providing documentation and runnable examples for the LangExtract library. This skill teaches AI coding assistants how to use LangExtract to extract structured information from text, covering installation, API key configuration, basic and relationship extraction, and handling multiple documents. It includes detailed references on provider selection (Gemini, OpenAI, Ollama), prompt validation levels, and resolver parameters for fine-tuning alignment. The skill is designed to be installed via copy or symlink into compatible agent tools like Google Antigravity, Anthropic Claude Code, OpenAI Codex, and GitHub Copilot.
skills · high confidence
New LangExtract benchmark suite for performance and quality testing
A new benchmark suite has been added to the \benchmarks\ directory to measure tokenization speed and extraction quality across multiple languages and text types. The suite includes a main runner (\benchmark.py\) that supports cloud models (Gemini) and local models (Ollama), automatically downloading test texts from Project Gutenberg for English, Japanese, French, and Spanish. It also features a dedicated fuzzy alignment benchmark (\fuzzy\_benchmark.py\) to evaluate the correctness and performance of the resolver's alignment logic, and visualization tools (\plotting.py\) to generate comparative charts of tokenization throughput and extraction metrics.
benchmarks · high confidence
New Romeo and Juliet extraction notebook with interactive visualization
A new example notebook demonstrates extracting characters, emotions, and relationships from Shakespeare's Romeo and Juliet using LangExtract. The notebook includes setup instructions for the Gemini API, defines an extraction task with few-shot examples, and shows how to run extraction using the gemini-3.5-flash model. It also features an interactive visualization component that saves results to JSONL and generates an HTML file for exploring extracted entities and their attributes.
examples/notebooks · high confidence
New custom provider plugin example for LangExtract
Added an example demonstrating how to create and register a custom provider plugin for LangExtract. This includes a sample \CustomGeminiProvider\ implementation that wraps the Gemini API, a \CustomProviderSchema\ for generating structured output constraints from examples, and a test script to verify the setup. The example shows how to use entry points for automatic provider discovery and how to configure model inference with custom schema constraints.
_examples/custom\_provider\plugin · high confidence
New provider system with router, lazy loading, and plugin support
LangExtract now uses a new provider system that maps model IDs to specific LLM backends (Gemini, Ollama, OpenAI) via a pattern-based router. This system supports lazy loading of provider dependencies to avoid unnecessary imports and includes automatic discovery of third-party provider plugins via Python entry points. Users can now explicitly select providers or rely on auto-detection, and the system provides structured output schemas for Gemini and OpenAI.
langextract/providers · high confidence
New scripts for creating and validating community provider plugins
Added \scripts/create\_provider\_plugin.py\ to automate the generation of new LangExtract provider plugins, including package structure, entry points, and boilerplate code. Added \scripts/validate\_community\_providers.py\ to enforce formatting and content rules for the community provider registry table in \COMMUNITY\_PROVIDERS.md\, ensuring valid PyPI names, GitHub links, and alphabetical sorting.
scripts · high confidence
Removals
Removal of Kokoro test script
The Kokoro test script (kokoro/test.sh) has been removed from the repository. This script previously handled the setup of the Python virtual environment, installation of dependencies, type checking with pytype, and execution of pytest tests for the langextract project within the Kokoro CI system. Its deletion indicates that this specific CI configuration is no longer maintained or used.
kokoro · high confidence
Architecture
LangExtract core library restructured into a clean, layered architecture
The \langextract/core\ package has been reorganized into a modular, layered architecture to improve maintainability and fine-grained dependency management. This change introduces a new \FormatHandler\ class that centralizes prompt formatting and output parsing for JSON and YAML, including support for code fences and wrapper objects. It also adds a framework for user-provided \output\_schema\ validation (enforcing LangExtract's raw JSON envelope) and integrates it with \BaseLanguageModel\ for providers like Gemini and OpenAI. Additionally, the core now includes a multi-language tokenizer (Regex and Unicode), a comprehensive set of domain data classes (Document, Extraction, CharInterval), and debug utilities with automatic redaction of sensitive keys in logs.
langextract/core · high confidence
Behavioural changes
LangExtract v2.0.0 backward compatibility layer for deprecated imports
LangExtract introduces a backward compatibility layer (\langextract.\_compat\) to support the upcoming v2.0.0 release. This layer provides shims for deprecated imports from \langextract.inference\, \langextract.schema\, \langextract.exceptions\, and \langextract.registry\, redirecting them to their new canonical locations in \langextract.core\ and \langextract.providers\ while emitting \FutureWarning\ deprecation notices. Users relying on the old import paths will continue to function but are advised to update their code to the new module structure before v2.0.0.
langextract · high confidence
Test coverage
Added comprehensive test suite for LangExtract core features
Added a new test suite in the \tests/\ directory covering the provider plugin generator, \extract()\ parameter precedence, schema integration with Gemini and Ollama, factory model creation, format handling, fuzzy alignment correctness, Gemini retry logic, and I/O progress bar cleanup.
tests · high confidence
Dependencies
LangExtract 1.7.0 release with dependency updates and plugin infrastructure
This release updates the library to version 1.7.0 and raises the minimum Python requirement to 3.10. It significantly upgrades the Google GenAI SDK dependency from \>=0.1.0 to \>=1.39.0 and adds Google Cloud Storage as a new dependency, while removing the previously required langfun and python-magic dependencies. The package now includes a provider plugin system via entry points for Gemini, Ollama, and OpenAI, and adds optional dependencies for OpenAI (\>=1.50.0), development tooling (pyink, isort, pre-commit), and Jupyter notebooks. An example custom provider plugin is also included to demonstrate how users can register third-party providers.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
This is the PUBLIC form of this artifact. Findings are listed in full, but the details of SECURITY findings — which rule fired, in which file, on which line, and how to fix it — are deliberately withheld, and any secret-scanner results are excluded entirely. Where detail is absent here it was REMOVED FOR PUBLICATION; it is not missing from the analysis. The complete artifact is available from the repository owner.
Score
- CAI 45 → 70 (+25.5)
- Rubric changed (rubric-2026.08.15 → rubric-2026.09.15) — scores are not directly comparable.
Lenses
- Code Health 90 → 84 (-6.8)
- Architecture 94 → 99 (+4.6)
- Maturity 63 → 67 (+3.5)
- Readiness 25 → 67 (+42.4)
- Security 48 → 70 (+21.7)
Resolved (55)
- Coverage not measured — test suite did not build
- Dimension evaluation failed
- Duplicated block (12 lines × 2) (tests/openai_batch_test.py)
- Duplicated block (12 lines × 2) (tests/provider_schema_test.py)
- Duplicated block (13 lines × 2) (tests/test_live_api.py)
- Duplicated block (14 lines × 2) (tests/annotation_test.py)
- Duplicated block (14 lines × 2) (tests/chunking_test.py)
- Duplicated block (16 lines × 2) (tests/provider_plugin_test.py)
- Duplicated block (17 lines × 2) (tests/inference_test.py)
- Duplicated block (18 lines × 2) (langextract/providers/ollama.py)
- Duplicated block (18 lines × 2) (tests/annotation_test.py)
- Duplicated block (21 lines × 2) (tests/schema_test.py)
- Duplicated block (6 lines × 3) (tests/test_gemini_batch_api.py)
- Duplicated block (8 lines × 2) (tests/prompting_test.py)
- Duplicated block (9 lines × 2) (tests/resolver_test.py)
- Duplicated block (9 lines × 2) (tests/test_gemini_batch_api.py)
- High IaC: DS-0002 (examples/ollama/Dockerfile)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- …and 35 more
New (126)
- Ambiguous resolution methods. resolve takes a model_id and resolve_provider takes a provider_name. It is unclear if these are mutually exclusive or if resolve can also take a provider name. The naming suggests different intents but they likely perform similar lookup operations.
- Annotator._annotate_documents_sequential_passes (cognitive 16) (langextract/annotation.py)
- Annotator._annotate_documents_single_pass (cognitive 31) (langextract/annotation.py)
- ChunkIterator.next (cognitive 19) (langextract/chunking.py)
- Conflicting high-level entry points. langextract.extraction.extract is a standalone function that performs extraction. langextract.annotation.Annotator is a class with annotate_documents and annotate_text methods that also perform extraction. It is unclear when a user should use the functional API vs the class-based API, as they likely share significant implementation logic.
- Dependency hygiene PARTLY measured — Python dependencies read, no exact pin to grade for currency
- Duplicate type definition with different semantics. langextract.core.data.CharInterval is used in the public Extraction and TextChunk models, while langextract.core.tokenizer.CharInterval is used in the internal Token model. They have identical fields (start_pos, end_pos) but exist in different modules, creating confusion about which type to use for character intervals in different contexts.
- Duplicated block (10–11 lines × 2) (langextract/resolver.py)
- Duplicated block (15 lines × 2) (examples/ollama/demo_ollama.py)
- Duplicated block (15 lines × 2) (scripts/create_provider_plugin.py)
- Duplicated block (17–18 lines × 2) (langextract/providers/gemini.py)
- Duplicated block (18–31 lines × 2) (langextract/_compat/inference.py)
- Duplicated block (21 lines × 2) (langextract/providers/schemas/gemini.py)
- Duplicated block (21–22 lines × 2) (langextract/providers/gemini.py)
- Duplicated block (21–27 lines × 2) (langextract/providers/schemas/gemini.py)
- Duplicated block (32–35 lines × 2) (langextract/providers/ollama.py)
- Edited copy of a member (40 corresponding lines) (langextract/providers/ollama.py)
- FileTooLong: langextract/resolver.py (langextract/resolver.py)
- FileTooLong: providers/gemini_batch.py (langextract/providers/gemini_batch.py)
- FileTooLong: scripts/create_provider_plugin.py (scripts/create_provider_plugin.py)
- …and 106 more
Changes since last survey
- 9 commits — 7 feature/other, 2 fixes
By area
- langextract/providers — 3 commits
- (root) — 2 commits
- docs/examples — 1 commit
- langextract/annotation.py — 1 commit
- langextract/extraction.py — 1 commit
- langextract/factory.py — 1 commit
Notable commits
- fix: Fix Vertex batch live tests (#538)
- fix: Fix provider plugin generation (#513)
- change: Clarify suppressed parsing errors (#521)
- change: Forward Gemini generation settings
- change: Prepare v1.7.0 release (#539)
- change: Raise on model responses without text (#534)
- change: Release completed prompts before building the next batch (#543)
- change: Route OpenAI reasoning and GPT-3.5 models (#494)
- change: Surface blocked Gemini batch items instead of returning empty output (#537)
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
google/langextract was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 26 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 62b933a2c757fd2bbb100498571b8d1692db4344 — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-d0929f7ac71f.