Skip to content
CAI
Software that uses CAICheck a score

stanford-oval/storm

49.3

Weak · 18 September 2026

12.6k

lines of production code

Python

primary language

1

measurement over time

CAI band scale
CAI lens gauges

What this system is

This system is a Python-based framework for automated, multi-agent collaborative research and long-form article generation. It orchestrates complex workflows using DSPy and LiteLLM to support a wide variety of language models and retrieval sources, including vector stores and web search APIs. The platform provides both a programmatic pipeline for knowledge synthesis and a minimal Streamlit interface for managing the creation and browsing of generated articles.

Features

Add article creation and management utilities

The demo application now includes utility modules for creating and managing articles. CreateNewArticle.py implements a multi-step workflow for generating new articles, handling user input, triggering research and writing processes, and displaying results. MyArticles.py provides a page for browsing existing articles with pagination support and allows users to select and view previously generated content.

_frontend/demo\_light/pages\util · high confidence

Introduce Co-STORM collaborative research engine

Adds the Co-STORM module to the knowledge\_storm package, providing a new engine for collaborative, multi-agent research workflows. This includes a configurable LM setup supporting OpenAI, Azure, and Together AI providers via LiteLLM, along with specialized agents (SimulatedUser, PureRAGAgent, Moderator, CoStormExpert) and utilities for warm-starting hierarchical chats and expert generation.

_knowledge\_storm/collaborative\storm · high confidence

Introduce Co-STORM collaborative research modules

Added a new set of modules under \knowledge\_storm/collaborative\_storm/modules\ that implement the Co-STORM framework for collaborative, multi-agent knowledge synthesis. This includes agents for expert role-playing, simulated users, and a moderator to guide discussions, alongside utilities for grounded question answering, information insertion into a hierarchical knowledge base, and article generation. The addition introduces a callback system for monitoring pipeline stages and integrates with the DSPy framework for language model interactions.

_knowledge\_storm/collaborative\storm/modules · high confidence

Introduce STORM wiki pipeline modules for article generation and polishing

The \knowledge\_storm/storm\_wiki/modules\ directory now contains the core implementation for the STORM wiki pipeline, including modules for knowledge curation, outline generation, article generation, and article polishing. These modules leverage the DSPy framework to orchestrate LLM interactions for tasks such as persona-based information gathering, section writing, and lead section generation. The codebase introduces specific DSPy signatures and modules (e.g., \ConvToSection\, \PolishPageModule\) and integrates with a callback system to handle pipeline events. Additionally, the retriever module includes a list of generally unreliable and deprecated sources according to Wikipedia standards to filter search results.

_knowledge\_storm/storm\wiki/modules · high confidence

Introduce minimal STORM Wiki user interface

A new minimal Streamlit-based UI for STORMWikiRunner is added to the \frontend/demo\_light\ directory. This interface allows users to create new articles, view intermediate generation steps in real-time, and display written articles alongside their references. It includes a 'My Articles' page to browse previously created content and provides setup instructions for configuring API keys and dependencies.

_frontend/demo\light · high confidence

Knowledge Storm v1.1.0 release with LiteLLM integration and new model support

This release introduces Knowledge Storm version 1.1.0, featuring a comprehensive integration of LiteLLM to unify access to various language models and embedding services. The library now supports a wider range of providers, including DeepSeek, Gemini, and Groq, alongside existing OpenAI and Azure capabilities. The \Encoder\ class has been encapsulated to leverage LiteLLM for embedding generation with parallel processing and local disk caching. Additionally, the release includes new retrieval methods for Serper, Tavily, Brave Search, and SearXNG, and refactors the dataclass and interface structures to support datatype sharing between STORM and Co-STORM workflows.

_knowledge\storm · high confidence

New example scripts for Co-STORM, Azure AI Search, and custom corpora

The examples directory now includes a new Co-STORM pipeline script (run\_costorm\_gpt.py) that supports both OpenAI and Azure OpenAI models, along with multiple search retrievers. The existing STORM Wiki examples have been updated to support Azure AI Search as a retrieval option, and new scripts have been added to demonstrate running STORM with Claude, DeepSeek, Gemini, and Groq models. Additionally, a helper script and documentation have been added to show how to ground STORM on a custom corpus using a local Qdrant vector store.

examples · high confidence

Removals

Removal of legacy DeepSearchRunner engine

The legacy \src/engine.py\ file containing the \DeepSearchRunner\ class and its associated argument structures has been deleted. This removes the previous implementation of the topic research pipeline, including its support for general and persona-guided conversational question asking, from the source code.

src · high confidence

Removal of legacy evaluation scripts and metrics

The \eval\ directory has been cleaned up by removing several legacy evaluation scripts and their associated resources. Specifically, \citation\_quality.py\ (AutoAIS), \eval\_article\_quality.py\ (Prometheus-based rubric grading), \eval\_outline\_quality.py\, \evaluation\_prometheus.py\, \evaluation\_trim\_length.py\, and \metrics.py\ have been deleted, along with the \eval\_rubric\_5.json\ configuration file. This removes the ability to run these specific citation, article, and outline quality evaluations.

eval · high confidence

Removal of specific 2022 event data files from FreshWiki

The FreshWiki dataset has removed the JSON data files for several specific 2022 events, including the 2022 AFL Grand Final, the 2022 Crimean Bridge explosion, the 2022 Hungarian parliamentary election, the 2022 Istanbul bombing, the 2022 Luzon earthquake, and the 2022 New York City Subway attack. This cleanup aligns with the removal of STORM paper-specific experiment directories, indicating a reduction in the dataset's scope or a reorganization of stored historical records.

FreshWiki · high confidence

Removed legacy Wikipedia source list module

The static HTML file for the Wikipedia 'Reliable sources/Perennial sources' module has been deleted from the \src/modules\ directory. This change removes the locally cached snapshot of the Wikipedia page listing perennial sources, eliminating the legacy module that previously provided this specific reference data.

src/modules · high confidence

Removed legacy pre-writing and writing scripts

The standalone entry-point scripts \src/scripts/run\_prewriting.py\ and \src/scripts/run\_writing.py\ have been deleted. These files previously provided direct CLI interfaces for the STORM pipeline's pre-writing (research and outline generation) and final-writing (article generation) stages, including argument parsing for LLM configuration and input sources. Users can no longer invoke these specific legacy scripts directly.

src/scripts · high confidence

Behavioural changes

STORM Wiki engine now uses LiteLLm for LLM integration

The STORM Wiki engine has been updated to integrate LiteLLm as the underlying interface for language model interactions. This change allows the system to support a wider variety of LLM providers and models through LiteLLm's unified API, replacing previous direct integrations. Users can now configure different LLMs for specific pipeline stages (such as conversation simulation, question asking, outline generation, and article polishing) using LiteLLm-compatible model identifiers, enabling greater flexibility in model selection and deployment.

_knowledge\_storm/storm\wiki · high confidence

STORM v1.1.0 release with LiteLLM integration and Python 3.10+ requirement

The \knowledge-storm\ package has been updated to version 1.1.0, introducing integration with LiteLLM to support a wider range of language and embedding models. The project now enforces a minimum Python version of 3.10 and restructures the source code into a proper Python package (\knowledge\_storm/\) with a \setup.py\ configuration. Additionally, a pre-commit hook using Black is added to enforce code formatting, and the documentation has been updated to reflect the new Co-STORM capabilities and installation instructions.

(repo-wide) · high confidence

Dependencies

Updated Python dependencies for core and demo components

The project's Python dependencies have been updated across two manifest files. In the root requirements.txt, the stack has shifted from older NLP libraries (Flair, NLTK, scikit-learn) to a LangChain-based architecture, introducing packages such as langchain-text-splitters, langchain-huggingface, langchain-qdrant, qdrant-client, litellm, and trafilatura, while upgrading dspy-ai to version 2.4.9. Additionally, a new requirements.txt file was created for the frontend/demo\_light directory, pinning Streamlit to version 1.31.1 and adding supporting UI components like streamlit-card, extra-streamlit-components, and st-pages.

(dependencies) · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

Baseline

  • First survey — no prior run to compare against. CAI 49.

Lenses

  • Code Health 95
  • Architecture 100
  • Maturity 46
  • Readiness 27
  • Security 88

Changes since last survey

  • 238 commits — 208 feature/other, 30 fixes

By area

  • (root) — 63 commits
  • (repo) — 61 commits
  • knowledge_storm/rm.py — 18 commits
  • knowledge_storm/collaborative_storm — 10 commits
  • knowledge_storm/storm_wiki — 10 commits
  • knowledge_storm/utils.py — 7 commits
  • examples/run_storm_wiki_claude.py — 6 commits
  • frontend/demo_light — 6 commits
  • knowledge_storm/lm.py — 6 commits
  • .github/workflows — 4 commits
  • examples/run_storm_wiki_gpt_with_VectorRM.py — 4 commits
  • src/storm_wiki — 4 commits
  • examples/helper — 3 commits
  • examples/run_storm_wiki_gpt.py — 3 commits
  • examples/storm_examples — 3 commits
  • src/lm.py — 3 commits
  • FreshWiki/json — 2 commits
  • eval/citation_quality.py — 2 commits
  • eval/eval_article_quality.py — 2 commits
  • examples/README.md — 2 commits

Notable commits

  • fix: Fix #140.
  • fix: Fix AzureOpenAIModel.
  • fix: Fix ClaudeModel support.
  • fix: Fix AzureOpenAIModel to match the latest Azure API.
  • fix: Fix a type in readme.md
  • fix: Fix examples: args.max_thread_num is not passed into STORMRunnerArguments.
  • fix: Fix format issue.
  • fix: Fix import order and refine code comments.
  • fix: Fix organization name for prometheus model
  • fix: Fix python code format.
  • fix: Fix the Python complain when anthropic is not installed.
  • fix: Fixed line-too-long
  • fix: Fixed bug with collected_results grabbing the wrong data for title and link when knowledge graph to be None
  • fix: Merge pull request #148 from stanford-oval/dev-fix-vllm
  • fix: Merge pull request #208 from itinance/fix/typo-reogranize
  • fix: Merge pull request #218 from montasaurus/adam/fix-broken-readme-link
  • fix: Revert unnecessary change.
  • fix: STORM code typo fix
  • fix: Update README.md — fix secrets.toml syntax
  • fix: Update README.md — fixes You.com API key env variable
  • …and 218 more

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

stanford-oval/storm was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 18 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit fb951af7744dab086e34962e9bc6fe878e145f83 — the exact code this score is about.
  • Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-5d04157a340d.