microsoft/qlib
60.7
Adequate · 26 September 2026
102.4k
lines of production code
Python
with C++
4
measurements over time
What this system is
Qlib is an open-source, AI-oriented quantitative investment platform that provides a comprehensive infrastructure for quantitative research, model development, and backtesting. It supports a wide range of machine learning models, including deep learning and gradient boosting, enabling users to build, train, and evaluate alpha factors across various market frequencies and regions. The system facilitates end-to-end workflows from data ingestion and feature engineering to rolling model management, online serving, and reinforcement learning-based order execution.
How it got here
2020 — Initial repository structure and core module development
34 changes.
This period established the foundational structure of the Qlib repository, introducing essential build configurations, security measures, and a comprehensive test suite. It focused on building core infrastructure, including a new initialization system, modular data handling pipelines, and high-performance Cython operators, while expanding the model library with new baselines and meta-learning capabilities. The work also delivered critical user-facing features such as online trading simulation, interactive reporting, automated hyperparameter tuning, and extensive benchmark examples for various machine learning models.
2021 — online serving and high-frequency infrastructure
28 changes.
This period focused on establishing robust online model management and high-frequency data processing capabilities. Key developments included the introduction of an online serving module for dynamic model updates, a new file-based storage backend, and extensive support for high-frequency trading workflows with custom operators and rolling data handlers.
2022–2025 — Reinforcement learning and meta-learning infrastructure
21 changes.
This period focused on establishing a comprehensive Reinforcement Learning framework, including data handling, order execution environments, and training APIs, alongside introducing meta-learning model infrastructure. It also expanded benchmarking capabilities with new examples for concept drift and graph-based forecasting, while adding Point-in-Time data collection and rolling evaluation modules to support advanced quantitative research workflows.
Features
Add CSI300/CSI100/CSI500 index data collector
A new data collector script (collector.py) and documentation (README.md) have been added to the cn\_index directory to fetch constituent changes for CSI300, CSI100, and CSI500 indices. The tool parses company additions and removals from CSIndex Excel files and HTML announcements, normalizes stock symbols, and outputs instrument data compatible with the Qlib framework.
_scripts/data\_collector/cn\index · high confidence
Add DDG-DA and Rolling Benchmark examples for concept drift adaptation
Introduces two new benchmarking examples in the dynamic benchmarks directory: DDG-DA, which implements the Data Distribution Generation for Predictable Concept Drift Adaptation method to forecast and adapt to data distribution changes, and a Rolling Retrain (RR) baseline that periodically retraining forecasting models to handle market dynamics. Users can run these examples via \workflow.py\ and \rolling\_benchmark.py\ respectively, supporting both Linear and LightGBM models with provided configuration files and a visualization script for DDG-DA results.
_examples/benchmarks\dynamic/DDG-DA · high confidence
Add GRU benchmark examples for Alpha158 and Alpha360 datasets
New benchmark examples for the Gated Recurrent Unit (GRU) model have been added to the documentation, including a README and configuration files for both the Alpha158 and Alpha360 datasets. These configurations allow users to reproduce experiments using the PyTorch-based GRU implementation with specific data handlers, preprocessing pipelines, and backtesting strategies defined for the Chinese stock market.
examples/benchmarks/GRU · high confidence
Add HIST benchmark example for stock trend forecasting
Added a new benchmark example for the HIST (Graph-based Framework for Stock Trend Forecasting) model. This includes a configuration file (\workflow\_config\_hist\_Alpha360.yaml\) that sets up the HIST model using the Alpha360 data handler and a TopkDropoutStrategy for backtesting on the CSI300 market, along with a README linking to the original code and paper.
examples/benchmarks/HIST · high confidence
Add MultiSegRecord and SignalMseRecord for advanced signal evaluation
New record classes are now available in the workflow module to enhance signal analysis capabilities. MultiSegRecord allows users to generate predictions and compute Information Coefficient (IC) and Rank IC metrics across multiple data segments, supporting detailed performance breakdowns. SignalMseRecord provides mean squared error (MSE) and root mean squared error (RMSE) calculations for signal evaluation, automatically logging these metrics and saving the results as artifacts for experiment tracking.
qlib/contrib/workflow · high confidence
Add Point-in-Time (PIT) data collector for Chinese stock fundamentals
A new script and documentation have been added to the \scripts/data\_collector/pit\ directory to support collecting and normalizing Point-in-Time fundamental data for Chinese stocks. The \collector.py\ module integrates with the Baostock API to fetch quarterly and annual financial metrics, including Return on Equity (ROE) from performance express reports and profit data, as well as year-over-year net income growth forecasts. Users can download this data, normalize it, and dump it into the Qlib PIT format using the provided command-line interface.
_scripts/data\collector/pit · high confidence
Add TFT benchmark data formatting infrastructure
Introduced the \data\_formatters\ package for the Temporal Fusion Transformer (TFT) benchmark, providing the base classes and specific implementations required to prepare datasets for model training. The \base.py\ module defines the \GenericDataFormatter\ abstract class and input type enumerations, establishing the contract for data transformation, normalization, and splitting. The \qlib\_Alpha158.py\ module implements this interface for the Alpha158 dataset, handling column definitions, scaler calibration using scikit-learn, and train/validation/test splits based on temporal boundaries.
_examples/benchmarks/TFT/data\formatters · high confidence
Add TFT benchmark experiment settings for Alpha158
New configuration files have been added to the TFT benchmark examples to support the Alpha158 experiment. This includes an \\_\init\\_.py\ file and a \configs.py\ module that define the directory structure for data, serialized models, and results, as well as the specific data formatter (\data\_formatters.qlib\_Alpha158\) and default hyperparameter iteration counts required to run this benchmark.
_examples/benchmarks/TFT/expt\settings · high confidence
Add TFT benchmark library components
Added the core library files for the Temporal Fusion Transformer (TFT) benchmark example, including the model architecture implementation, hyperparameter optimization managers, and utility functions for data handling and TensorFlow configuration.
examples/benchmarks/TFT/libs · high confidence
Add Temporal Fusion Transformers (TFT) benchmark example
Added a new benchmark example for the Temporal Fusion Transformers (TFT) model, including the \tft.py\ implementation, a YAML workflow configuration for the Alpha158 dataset, and documentation. This enables users to run TFT-based time series forecasting and backtesting within the Qlib framework, with specific requirements for Python 3.6-3.7 and GPU support.
examples/benchmarks/TFT · high confidence
Add Temporal Routing Adaptor (TRA) benchmark examples
Added a new benchmark directory for the Temporal Routing Adaptor (TRA) model, including a README with usage instructions and performance tables, a Jupyter notebook for reproducing paper results, and configuration files for running TRA with Alpha158 and Alpha360 datasets via the Qlib workflow. The entry also includes the example Python script, shell scripts for execution, and the core model and dataset implementation files (src/model.py, src/dataset.py) that define the MTSDatasetH and TRAModel classes used by these configurations.
examples/benchmarks/TRA · high confidence
Add US index data collector for SP500, NASDAQ100, DJIA, and SP400
Users can now collect historical company composition data for major US stock indices (SP500, NASDAQ100, DJIA, SP400) using the new \scripts/data\_collector/us\_index/collector.py\ script. This tool parses instrument lists from Wikipedia and NASDAQ OMX, supporting both current and historical index constituents, and outputs data compatible with the Qlib framework's instrument format.
_scripts/data\_collector/us\index · high confidence
Add crypto data collector for CoinGecko
Introduces a new data collection tool for cryptocurrency market data sourced from CoinGecko. The change adds a Python collector script, a requirements file specifying dependencies like pycoingecko and pandas, and documentation. Users can now download, normalize, and dump crypto data (prices, volumes, market caps) for use in the system, though the dataset is limited to data retrieval and does not support backtesting due to the lack of OHLC data.
_scripts/data\collector/crypto · high confidence
Add detailed quantitative research workflow tutorial
A new Jupyter notebook tutorial has been added to the examples directory, guiding advanced users through a step-by-step quantitative research workflow using Qlib. The notebook demonstrates how to initialize the library, download market data, inspect raw features such as calendar and stock prices, and visualize this data using Plotly candlestick charts, providing a more granular alternative to the existing quick-start guides.
examples/tutorial · high confidence
Add fund data collector for CN market
Introduces a new data collection tool for Chinese mutual fund data sourced from East Money. The change adds a \collector.py\ script and \README.md\ documentation, enabling users to download, normalize, and dump daily fund data (specifically fields like DWJZ and LJJZ) into Qlib format for the CN region.
_scripts/data\collector/fund · high confidence
Add model interpretation interface for LightGBM
A new model interpretation module has been introduced, providing a base interface for feature importance analysis and a concrete implementation for LightGBM models. Users can now access feature importance scores via the \LightGBMFInt\ class, which wraps the LightGBM booster's native feature importance method and returns a sorted pandas Series of feature names and their corresponding importance values.
qlib/model/interpret · high confidence
Add nested decision execution example
Added a new example demonstrating nested decision execution in backtesting, allowing users to combine strategies across different frequencies. The example shows how to use a DropoutTopkStrategy for weekly portfolio generation and an SBBStrategyEMA for daily order execution, as well as a high-frequency variant with daily portfolio generation and minutely order execution.
_examples/nested\_decision\execution · high confidence
Add portfolio optimization example with EnhancedIndexingStrategy
Users can now run an end-to-end portfolio optimization workflow using the new \EnhancedIndexingStrategy\ to maximize returns while minimizing tracking error against a benchmark. This example includes a configuration file (\config\_enhanced\_indexing.yaml\) demonstrating the setup, a script (\prepare\_riskdata.py\) to generate statistical risk model data, and documentation (\README.md\) guiding users through the preparation of CSI300 weight and risk data.
examples/portfolio · high confidence
Add simple reinforcement learning example notebook
A new Jupyter notebook (\examples/rl/simple\_example.ipynb\) has been added to demonstrate how to build and run a basic reinforcement learning workflow using Qlib RL. The notebook provides a step-by-step guide for users to define a primitive simulator, a policy, and a reward function, and then execute training and backtesting workflows based on these components.
examples/rl · high confidence
Added LSTM benchmark examples for Alpha158 and Alpha360 datasets
New benchmark examples for the Long Short-Term Memory (LSTM) model have been added to the examples/benchmarks/LSTM directory. This includes a README and two workflow configuration files: one for the Alpha158 dataset (using TSDatasetH with specific feature filters) and one for the Alpha360 dataset (using DatasetH). Both configurations provide complete, runnable setups for training, validation, and backtesting the LSTM model on Chinese stock market data (CSI300 benchmark) using PyTorch.
examples/benchmarks/LSTM · high confidence
Added LightGBM hyperparameter optimization examples for Alpha158 and Alpha360
New example scripts (hyperparameter\_158.py and hyperparameter\_360.py) and documentation have been added to the examples/hyperparameter/LightGBM directory. These scripts enable users to perform hyperparameter optimization for LightGBM models using Optuna, specifically targeting the Alpha158 and Alpha360 datasets. The examples demonstrate how to configure the LGBModel with various tunable parameters (such as colsample\_bytree, learning\_rate, and num\_leaves) and integrate with the existing Qlib dataset configurations.
examples/hyperparameter · high confidence
Initial release of the report module with interactive visualization support
This change introduces the \qlib.contrib.report\ package, providing a new capability to generate analysis reports for quantitative models. The module includes \graph.py\, which leverages Plotly to create interactive charts (scatter, bar, distribution, heatmap, histogram) and handles notebook rendering, including specific support for Google Colab. It also provides \utils.py\ with helper functions for generating subplots and handling datetime index gaps in time-series plots. The \\_\init\\_.py\ file defines the list of available report graph types, such as position analysis, score IC, cumulative returns, risk analysis, and model performance.
qlib/contrib/report · high confidence
Initial repository structure and build configuration
This change establishes the foundational structure of the Qlib repository by adding essential configuration and build files. It includes a Makefile for managing development environments, linting, and documentation, alongside configuration files for pre-commit hooks (Black, flake8), commit message linting, and static type checking (mypy). The setup.py file is introduced to handle the compilation of Cython extensions for high-performance data operations, and a Dockerfile is added to support containerized deployment. Additionally, standard repository metadata files such as LICENSE, README, CHANGELOG, and SECURITY.md are included to define usage terms, provide documentation, and outline security reporting procedures.
(repo-wide) · high confidence
Introduce Qlib RL data handling module
Added a new \qlib/rl/data\ package that provides the data infrastructure for the Reinforcement Learning backtesting framework. This includes abstract base classes for intraday backtest and processed data, a provider interface, and concrete implementations for loading and processing market data from both Qlib's native handler format and external pickle-styled files (used in research projects). The module also features an integration layer to initialize Qlib resources for NeuTrader compatibility and implements safe, restricted pickle deserialization to address security concerns when loading untrusted data files.
qlib/rl/data · high confidence
Introduce Qlib RL framework for single-asset order execution
Adds a new \qlib.rl\ module providing the core infrastructure for reinforcement learning-based trading, specifically targeting single-asset order execution. The release includes a \Simulator\ class to manage trading state transitions, \Interpreter\ classes to bridge between simulator states and RL policy observations/actions, and \Reward\ components for defining and combining reward signals. This establishes the foundational environment and interface layer for integrating RL agents into Qlib's execution logic.
qlib/rl · high confidence
Introduce RL training and backtesting API with parallel execution
The \qlib.rl.trainer\ module now provides a structured API for training and backtesting reinforcement learning policies. Users can utilize the \train\ and \backtest\ functions to execute workflows using a \Trainer\ and \TrainingVessel\ architecture, which supports parallel simulation via configurable concurrency and finite environment types. The module includes built-in callbacks for checkpointing, early stopping based on monitored metrics, and metrics logging, enabling robust and observable RL model development.
qlib/rl/trainer · high confidence
Introduce SepDataFrame to optimize memory usage for separate data processing
A new SepDataFrame class has been added to the data utilities to handle scenarios where multiple DataFrames (such as features, labels, and weights) are processed together but accessed separately. This change allows users to avoid the high memory cost of concatenating and then splitting data by keeping the DataFrames in a dictionary structure while mimicking standard DataFrame behavior for indexing and operations.
qlib/contrib/data/utils · high confidence
Introduce automated hyperparameter tuning pipeline
Added a new tuner module in qlib/contrib/tuner that enables automated hyperparameter optimization for estimators. The component includes a CLI launcher (launcher.py) to start tuning jobs via YAML configuration, a configuration manager (config.py) to parse experiment, optimization, and data settings, and a pipeline orchestrator (pipeline.py) that iterates through tuner configurations. The core logic uses Hyperopt's Tree-Parzen Estimator (tpe) to search parameter spaces defined in space.py, executing estimator subprocesses and evaluating results via restricted pickle loading (tuner.py) to ensure security while optimizing metrics like information ratio or model scores.
qlib/contrib/tuner · high confidence
Introduce file-based storage backend for calendars, instruments, and features
The \qlib/data/storage\ module now includes a new file-based storage implementation (\FileStorageMixin\, \FileCalendarStorage\, \FileInstrumentStorage\, \FileFeatureStorage\) that persists data to local text files. This adds a concrete backend for managing trading calendars, instrument definitions, and feature data, allowing users to store and retrieve this information from the filesystem rather than relying solely on in-memory or other storage mechanisms. The implementation supports frequency-based data organization, caching for performance, and standard list-like operations for calendar management.
qlib/data/storage · high confidence
Introduce model performance analysis and reporting module
Adds a new \analysis\_model\ package under \qlib/contrib/report\ that provides tools for visualizing and analyzing quantitative model performance. The module exposes a \model\_performance\_graph\ function which generates Plotly-based charts, including cumulative return curves, long-short/long-average distribution plots, and Information Coefficient (IC) metrics (bar charts, heatmaps, and Q-Q plots) to help users evaluate model accuracy and stability.
_qlib/contrib/report/analysis\model · high confidence
Introduce modular dataset, handler, loader, and processor components
The data layer in \qlib/data/dataset\ has been restructured into distinct, reusable components: \DataHandler\ and \DataHandlerLP\ for managing data fetching and preprocessing, \DataLoader\ (including \QlibDataLoader\ and \DLWParser\) for loading raw data from sources, \Processor\ classes (such as \DropnaProcessor\, \Fillna\, and \TanhProcess\) for applying transformations, and \BaseHandlerStorage\ with \HashingStockStorage\ for optimized data access. This change provides a more flexible and configurable interface for preparing data for model training and inference, allowing users to easily customize data loading, filtering, and processing pipelines.
qlib/data/dataset · high confidence
Introduce new evaluation, portfolio analysis, and PyTorch utility modules
The \qlib/contrib\ package now includes three new modules that expand the library's capabilities for strategy evaluation and model integration. \evaluate.py\ provides a \risk\_analysis\ function that supports both arithmetic and geometric accumulation modes for calculating annualized returns, along with \indicator\_analysis\ for trading statistics. \evaluate\_portfolio.py\ adds functions to calculate position values, daily return series, annualized returns, Sharpe ratios, and maximum drawdown directly from position data. Additionally, \torch.py\ introduces a \data\_to\_tensor\ utility to seamlessly convert pandas DataFrames, NumPy arrays, and lists into PyTorch tensors for device placement.
qlib/contrib · high confidence
Introduce online trading module for managing user portfolios
Adds a new \qlib.contrib.online\ package that provides infrastructure for online trading simulations. This includes a \UserManager\ to handle user data persistence (loading/saving strategies, models, and accounts), an \Operator\ class to orchestrate the online workflow (adding/removing users, generating trade orders, and executing them via a simulator), and supporting components like \User\, \ScoreFileModel\, and utility functions for safe pickle loading and exchange configuration.
qlib/contrib/online · high confidence
Introduce risk model covariance estimators
Added a new risk model module in \qlib/model/riskmodel\ that provides several covariance estimation strategies for stock returns. The module includes a base \RiskModel\ class handling data preprocessing (NaN masking/filling, centering, scaling) and exposes three specific estimators: \POETCovEstimator\ (Principal Orthogonal Complement Thresholding), \ShrinkCovEstimator\ (Ledoit-Wolf and Oracle Approximating Shrinkage towards constant variance, constant correlation, or single-factor targets), and \StructuredCovEstimator\ (using PCA or Factor Analysis to model latent factors). These classes allow users to estimate covariance matrices with more robust statistical methods than the default empirical approach.
qlib/model/riskmodel · high confidence
Introduce rolling model execution module
Adds a new \qlib.contrib.rolling\ package that provides a general-purpose framework for offline rolling model evaluation. The module includes a \Rolling\ base class to manage task configuration, data handling, and experiment tracking, along with a \DDGDA\ implementation for distribution-aware rolling strategies. Users can now execute rolling workflows via the command line using \python -m qlib.contrib.rolling\.
qlib/contrib/rolling · high confidence
Introduce structured experiment and recorder workflow APIs
The \qlib/workflow\ module now provides a structured, MLflow-based interface for managing experiments and recording results. Users can use the global \QlibRecorder\ (R) to start and end experiments with context managers, automatically handling status updates and error recovery. The new \Experiment\ and \Recorder\ classes offer intuitive methods for logging parameters, metrics, and artifacts, as well as searching and retrieving past records. Additionally, \RecordTemp\ and its subclasses like \SignalRecord\ provide templates for generating and saving analysis results such as IC and backtest reports, simplifying the workflow from model training to evaluation.
qlib/workflow · high confidence
Introduce unified data collection framework and new market collectors
The data collection scripts have been refactored to use a new base architecture (BaseCollector, BaseNormalize, BaseRun) that standardizes how data is fetched, normalized, and saved, allowing for easier creation of custom collectors. This change adds support for collecting 1-minute frequency data from Yahoo Finance, introduces a new collector for the Brazilian IBOVESPA index (br\_index), and implements a future calendar collector for Chinese trading days using Baostock. Additionally, the utility functions now use Akshare for retrieving the Chinese trading calendar and include a method to generate a calendar list based on a data coverage threshold.
_scripts/data\collector · high confidence
Introduction of Cython-accelerated rolling and expanding statistical operators
The data module now includes high-performance Cython implementations for rolling and expanding window calculations, specifically slope, R-squared, and residuals. These new operators (accessible via \rolling\_slope\, \rolling\_rsquare\, \rolling\_resi\, \expanding\_slope\, \expanding\_rsquare\, and \expanding\_resi\) are integrated into the expression engine to significantly speed up feature computation for quantitative models compared to previous pure-Python approaches.
qlib/data · high confidence
Introduction of Online Serving Module for Dynamic Model Management
A new online serving module has been added to the workflow, introducing \OnlineManager\ to orchestrate the lifecycle of online models and \OnlineStrategy\ (including \RollingStrategy\) to define how tasks are generated and models are updated over time. This module enables dynamic switching of decisive models based on time, supporting both real-time online trading and historical simulation modes. It includes \OnlineToolR\ for managing model tags (online/offline) within recorders and \RecordUpdater\ (specifically \DSBasedUpdater\) to refresh predictions and datasets as new data arrives, allowing users to maintain up-to-date model outputs without retraining from scratch.
qlib/workflow/online · high confidence
Introduction of core model base classes and trainer infrastructure
This change introduces the foundational components for the model module, including the \Model\ and \ModelFT\ base classes in \qlib/model/base.py\ which define the standard interface for training, prediction, and fine-tuning. It also adds \qlib/model/trainer.py\, providing \Trainer\ and \DelayTrainer\ classes to manage task execution and model lifecycle, along with utility classes like \ConcatDataset\ and \IndexSampler\ in \utils.py\ to support data handling.
qlib/model · high confidence
Introduction of meta-learning model infrastructure
The \qlib/model/meta\ package has been added, introducing a new set of classes to support meta-learning workflows. This includes \MetaTask\ for defining individual tasks with specific processing modes (full, test, transfer), \MetaTaskDataset\ for managing collections of meta-tasks, and base classes \MetaModel\, \MetaTaskModel\, and \MetaGuideModel\ to structure how meta-models guide or define base forecasting tasks. This provides the foundational components for training models that can transfer knowledge across different datasets or task definitions.
qlib/model/meta · high confidence
Introduction of new strategy base classes and infrastructure integration
The \qlib/strategy\ module now provides foundational classes (\BaseStrategy\, \RLStrategy\, \RLIntStrategy\) that integrate strategies with the backtesting infrastructure. These classes manage trade execution via \Exchange\, handle position tracking through \BasePosition\, and coordinate with the \TradeCalendarManager\. A key behavioral addition is the \get\_data\_cal\_avail\_range\ method, which allows strategies to determine the valid data range by considering both the internal data calendar and limitations imposed by an outer strategy's trade decisions. This enables more robust nested execution and cross-level communication within the backtesting framework.
qlib/strategy · high confidence
New CLI entry points for data retrieval and workflow execution
Added new command-line interface modules (\qlib/cli/data.py\ and \qlib/cli/run.py\) that expose \GetData\ for fetching datasets and a \workflow\ function for executing quant research workflows via configuration files. The workflow CLI supports Jinja2 template rendering for environment variable injection, base configuration inheritance, and automatic system path configuration, allowing users to run end-to-end experiments directly from the command line using the \fire\ library.
qlib/cli · high confidence
New RL contrib module for backtesting and training pipelines
Added a new \qlib.rl.contrib\ package providing utilities for reinforcement learning workflows, including a backtest engine (\backtest.py\) that executes orders via a simulator and generates performance reports, a configuration parser (\naive\_config\_parser.py\) for loading YAML/JSON/Python backtest settings, a training script (\train\_onpolicy.py\) for on-policy RL training with lazy data loading and checkpointing, and utility functions (\utils.py\) for reading order files in various formats.
qlib/rl/contrib · high confidence
New RL order execution example with training and backtest workflows
This change introduces a complete Reinforcement Learning example for order execution, providing scripts to generate pickle-style market data and training orders, along with configuration files for training (OPDS and PPO policies) and backtesting (including a TWAP baseline). The example demonstrates how to use the RL framework for financial trading strategies, including data preparation, model training, and performance evaluation.
_examples/rl\_order\execution · high confidence
New RL utility components for data handling, environment wrapping, and logging
Added a new \qlib.rl.utils\ package providing core infrastructure for the Reinforcement Learning framework. This includes \DataQueue\ for multi-process data loading from datasets, \EnvWrapper\ to integrate Qlib simulators and interpreters into standard Gym environments, and \FiniteVectorEnv\ (with dummy, subproc, and shmem variants) to manage finite data sources across parallel workers. Additionally, \LogCollector\ and \LogWriter\ classes are introduced to aggregate and output training metrics from distributed environment workers.
qlib/rl/utils · high confidence
New Yahoo Finance data collector with multi-region and 1-minute support
The Yahoo Finance data collector has been significantly expanded to support 1-minute frequency data in addition to daily data, and now covers multiple regions including China (CN), US, India (IN), and Brazil (BR). The collector uses the \yahooquery\ library to fetch data, with specific handling for 1-minute intervals which are limited to the last month by Yahoo's API. The implementation includes new normalization logic for 1-minute data that relies on local 1-day data for price adjustments, and provides automated update capabilities for daily data. Documentation has been updated to reflect these new regions, frequencies, and usage instructions.
_scripts/data\collector/yahoo · high confidence
New data analysis module for financial reports
Added a new \qlib.contrib.report.data\ module containing classes for analyzing feature distributions, statistics, and correlations over time. This includes analyzers for mean/std, skewness/kurtosis, missing values, infinite values, and auto-correlation, which can be used to generate detailed visualizations for model performance reports.
qlib/contrib/report/data · high confidence
New data caching and memory reuse demos
Added two new example scripts in the data\_demo folder to demonstrate how to optimize data processing performance in Qlib. The data\_cache\_demo.py script shows how to serialize a DataHandler to a pickle file and load it via a file:// URI to avoid redundant preprocessing across runs. The data\_mem\_resuse\_demo.py script demonstrates reusing a DataHandler instance directly in memory between multiple task\_train calls to skip disk I/O and data loading overhead.
_examples/data\demo · high confidence
New data collection utilities for CN 1-minute data and future trading calendars
Added two new scripts in the data collector contrib directory: \fill\_cn\_1min\_data.py\ uses existing 1-day data to fill in missing symbols in 1-minute CSV datasets, and \future\_trading\_date\_collector.py\ fetches future trading dates from Baostock to extend Qlib calendars for both daily and 1-minute frequencies.
_scripts/data\collector/contrib · high confidence
New data management and validation scripts
The scripts directory now includes a suite of tools for handling Qlib datasets. \get\_data.py\ provides a CLI to download market data (CN, US, and simple variants) and initialize the provider. \dump\_bin.py\ and \dump\_pit.py\ handle converting CSV/Parquet sources into Qlib's binary and point-in-time formats, with support for incremental updates and backups. To ensure data integrity, \check\_dump\_bin.py\ verifies that dumped binary files match the original CSVs, while \check\_data\_health.py\ scans datasets for missing values, large step changes, and naming issues. Additionally, \collect\_info.py\ gathers system and dependency details for troubleshooting.
scripts · high confidence
New ensemble and grouping modules for model aggregation
The \qlib/model/ens\ package introduces new capabilities for combining and organizing model predictions. It provides \Ensemble\ classes (\SingleKeyEnsemble\, \RollingEnsemble\, \AverageEnsemble\) to merge multiple prediction dataframes, including logic for rolling time-series concatenation and standardized averaging. Additionally, \Group\ and \RollingGroup\ classes allow users to organize model outputs into hierarchical structures and apply ensemble reductions, supporting parallel processing for efficiency.
qlib/model/ens · high confidence
New evaluation metrics for long-short strategies and prediction autocorrelation
The \qlib/contrib/eva\ module now includes new analysis capabilities for quantitative strategies. Users can calculate precision and returns for long-short portfolios using \calc\_long\_short\_prec\ and \calc\_long\_short\_return\, which allow filtering by quantile to isolate top-performing and worst-performing assets. Additionally, \pred\_autocorr\ and \pred\_autocorr\_all\ provide tools to measure the temporal autocorrelation of prediction signals, helping users assess the stability and decay of their alpha factors over time.
qlib/contrib/eva · high confidence
New example for non-fixed-frequency orderbook data via Arctic backend
Adds an example in examples/orderbook\_data demonstrating how to import, store, and query high-frequency orderbook data (ticks, orders, transactions) that does not adhere to a fixed shared frequency. The example uses an Arctic-based backend (requiring MongoDB) to handle variable-frequency data, providing scripts to initialize the library, import sample data, and run feature extraction tests using the ArcticFeatureProvider.
_examples/orderbook\data · high confidence
New example for rolling data processing workflows
Added a new example in \examples/rolling\_process\_data\ that demonstrates how to handle data generation across different rolling windows. The example introduces a \RollingDataHandler\ and a \workflow.py\ script that use a \DataHandler-based DataLoader\ to load raw features once, then apply processors to generate rolling-window-specific features, avoiding redundant data regeneration as the training window shifts.
_examples/rolling\_process\data · high confidence
New examples for workflow, model interpretation, and benchmarking
Added new example files to the \examples\ directory: \workflow\_by\_code.py\ and \workflow\_by\_code.ipynb\ demonstrate a complete quantitative research workflow (data loading, model training, signal analysis, and backtesting) using code-based configuration; \model\_interpreter/feature.py\ provides a standalone example for extracting and printing model feature importance; \run\_all\_model.py\ introduces a script to automatically run and benchmark multiple models defined in the \benchmarks\ directory, collecting and comparing metrics like IC and annualized return; and \examples/README.md\ documents the minimal hardware requirements (16GB RAM, 5GB disk) and notes on OS-specific result variance.
examples · high confidence
New high-frequency data handlers and processors for minute-level trading
This change introduces new data handling components in \qlib/contrib/data\ to support high-frequency (1-minute) trading strategies. It adds \HighFreqHandler\ and \HighFreqGeneralHandler\ for loading and processing minute-level OHLCV data, along with \HighFreqTrans\ and \HighFreqNorm\ processors to handle normalization and type conversion specific to high-frequency time series. Additionally, a \HighFreqProvider\ is included to manage the generation and caching of training, validation, and test datasets for these high-frequency models.
qlib/contrib/data · high confidence
New high-frequency feature engineering operators
Added new operators in \qlib/contrib/ops/high\_freq.py\ to support high-frequency data processing, including \DayCumsum\ for calculating cumulative sums within specific trading hours, \DayLast\ for retrieving the last value of each day, \FFillNan\ and \BFillNan\ for forward/backward filling missing values, \Date\ for extracting dates, and \Select\ for conditional value selection. These operators enable more granular feature engineering on minute-level data.
qlib/contrib/ops · high confidence
New high-frequency trading example with custom operators and data handlers
Added a new example in \examples/highfreq\ that demonstrates high-frequency (1-minute) data processing and price-trend prediction. This includes custom data operators (\DayLast\, \FFillNan\, \BFillNan\, \Date\, \Select\, \IsNull\, \Cut\) in \highfreq\_ops.py\, specialized data handlers (\HighFreqHandler\, \HighFreqBacktestHandler\) in \highfreq\_handler.py\, and a normalization processor (\HighFreqNorm\) in \highfreq\_processor.py\. The example provides a workflow (\workflow.py\) for loading, dumping, reloading, and reinitializing datasets, along with a configuration file (\workflow\_config\_High\_Freq\_Tree\_Alpha158.yaml\) for training a LightGBM model on Alpha158 features.
examples/highfreq · high confidence
New meta-learning data selection and expanded model library
This change introduces a new meta-learning-based data selection module under \qlib/contrib/meta/data\selection\, providing \MetaTaskDS\, \MetaDatasetDS\, and \MetaModelDS\ classes that use a \PredNet\ to dynamically reweight training samples based on historical performance. Additionally, the \qlib/contrib/model\ package is expanded with new baseline models including \CatBoostModel\, \DEnsembleModel\ (Double Ensemble), \ADARNN\, and \HFLGBModel\ (High-Frequency LightGBM), while the \\\init\\_.py\ file is updated to register these new classes alongside existing PyTorch and GBDT models.
qlib/contrib/model · high confidence
New multi-frequency and configurable dataset examples for LightGBM
The LightGBM benchmark directory now includes new examples demonstrating how to use intraday (1-minute) data alongside daily data for prediction. This is enabled by new \InstProcessor\ implementations (\Resample1minProcessor\ and \ResampleNProcessor\) that allow resampling high-frequency data to daily intervals, and a custom \Avg15minHandler\ that constructs features from 15-minute averages. Additionally, a \workflow\_config\_lightgbm\_configurable\_dataset.yaml\ example is added, showing how to define custom feature expressions directly in the YAML configuration without relying on the built-in Alpha158 or Alpha360 handlers.
examples/benchmarks/LightGBM · high confidence
New online serving examples for rolling strategies and prediction updates
Added three new example scripts in the \examples/online\_srv\ directory to demonstrate online model management workflows. \rolling\_online\_management.py\ illustrates how to use \OnlineManager\ with \RollingStrategy\ and \RollingGen\ to handle initial training, routine updates, and dynamic strategy addition. \online\_management\_simulate.py\ provides a simulation example for rolling tasks, including backtesting and risk analysis. \update\_online\_pred.py\ shows how to use \OnlineToolR\ to train a model, set it as online, and update its predictions.
_examples/online\srv · high confidence
New position analysis visualization module
Added a new \analysis\_position\ subpackage under \qlib.contrib.report\ that provides interactive Plotly-based graphs for analyzing backtest positions. This includes \cumulative\_return\_graph\ for tracking buy/sell/hold cumulative returns, \score\_ic\_graph\ for visualizing Information Coefficient (IC) and Rank IC, \rank\_label\_graph\ for monitoring the ranking performance of traded stocks, \report\_graph\ for displaying comprehensive backtest reports with drawdown highlights, and \risk\_analysis\_graph\ for showing monthly risk metrics like annualized return and maximum drawdown.
_qlib/contrib/report/analysis\position · high confidence
New rolling task management example for model training
Added a new example script, \task\_manager\_rolling.py\, that demonstrates how to use \TaskManager\ with \TrainerRM\ to manage rolling tasks for model training. This example illustrates the workflow for generating rolling tasks, training models (specifically XGBoost and LightGBM) using multiprocessing support via \TrainerRM\, and collecting results with \RecorderCollector\ and \RollingGroup\. It serves as a practical guide for users implementing rolling window strategies in their machine learning pipelines.
_examples/model\rolling · high confidence
New single-asset order execution RL framework
Introduces a new reinforcement learning module for single-asset order execution, providing a complete environment for training and backtesting trading strategies. The update adds a \SingleAssetOrderExecutionSimple\ simulator for lightweight backtesting and a \SingleAssetOrderExecution\ simulator that integrates with the Qlib backtest engine. It includes state and action interpreters (such as \FullHistoryStateInterpreter\ and \TwapRelativeActionInterpreter\) to translate market data into RL observations, along with reward functions like \PAPenaltyReward\ to optimize for price advantage while penalizing excessive trading volume. The framework also provides pre-built policies, including a PPO implementation and a TWAP baseline (\AllOne\), and exposes the necessary components via the \qlib.rl.order\_execution\ package.
_qlib/rl/order\execution · high confidence
New strategy module with SoftTopk, TWAP, and optimizer support
The \qlib.contrib.strategy\ package has been introduced, providing a new set of trading strategies and optimization tools. This includes \SoftTopkStrategy\ for budget-constrained rebalancing with deterministic sell/buy phases, \TWAPStrategy\ for time-weighted average price execution, and rule-based strategies like \SBBStrategyEMA\. Additionally, an \optimizer\ subpackage is added, offering \PortfolioOptimizer\ (supporting Global Minimum Variance, Mean-Variance, Risk Parity, and Inverse Volatility methods) and \EnhancedIndexingOptimizer\ for benchmark-tracking portfolios. The module also introduces \OrderGenerator\ classes to handle order creation from target weights, integrating with the backtest exchange infrastructure.
qlib/contrib/strategy · high confidence
New task workflow module for generation, management, and collection
Introduces the \qlib.workflow.task\ package, providing a structured workflow for handling tasks. This includes \TaskGen\ and \RollingGen\ for generating task templates (e.g., rolling windows), \TaskManager\ for storing and managing task lifecycles in MongoDB, and \Collector\ classes (like \RecorderCollector\) for aggregating results from experiments.
qlib/workflow/task · high confidence
Security
Introduction of secure pickle deserialization and utility refactoring
The \qlib/utils\ module has been restructured to include a new \pickle\_utils\ module that enforces secure deserialization by restricting pickle loading to a whitelist of safe classes, mitigating arbitrary code execution risks. This change replaces unsafe \pickle.load\ calls with \restricted\_pickle\_load\ across the codebase, including in \mod.py\ for instance initialization and \objm.py\ for file-based object management. Additionally, the module introduces a \Serializable\ base class in \serial.py\ to control attribute persistence during pickling, and adds \QlibException\ and specific error types in \exceptions.py\ for better error handling.
qlib/utils · high confidence
Behavioural changes
Backtest engine refactored into modular components
The backtest engine has been restructured into distinct modules (account, exchange, executor, decision, position, and high-performance data structures) to improve code organization and maintainability. This change introduces a more explicit separation of concerns, allowing users to configure exchange parameters (such as transaction costs, limit thresholds, and volume limits) and account settings (initial cash, positions, and benchmark configuration) with greater precision. The refactoring also enhances performance through new high-performance data structures for quote data retrieval and supports more flexible order execution logic, including detailed tracking of trade indicators and portfolio metrics.
qlib/backtest · high confidence
Qlib initialization and configuration system overhaul
The \qlib\ package now uses a new initialization flow in \qlib/\_\init\\_.py\ that explicitly handles NFS URI mounting (including Windows support via \auto\mount\), clears the data cache (\H.clear()\), and validates the configuration before registering. Configuration management has shifted to \qlib/config.py\, which introduces a \QSettings\ class based on \pydantic-settings\ for environment-variable-driven defaults (e.g., \QLIB\\ prefix) and adds explicit validation for required fields like \provider\_uri\ and \region\. The logging subsystem (\qlib/log.py\) is replaced by a custom \QlibLogger\ with a manager that ensures consistent module naming and level propagation, while \qlib/constant.py\ centralizes region and epsilon constants.
qlib · high confidence
Test coverage
Added test infrastructure and configuration for Qlib; Added test suite for data mid-layer components; Added tests for CSI300 dataset data quality; Added tests for GeneralPTNN with tabular and time-series datasets; Added tests for IndexData, SepDataFrame, and utility functions; Added tests for RL framework components; Added tests for backtest strategies and execution; Added tests for element and special operators; Added tests for file-based storage components; Added tests for rolling prediction and label updates; Added tests to validate MLflow client creation performance assumptions; Initial test suite and CI configuration.
Dependencies
Standardize benchmark and tooling dependencies
This change introduces explicit, pinned dependency files for all benchmark examples (such as ADARNN, ADD, ALSTM, CatBoost, and LightGBM) to ensure reproducible environments, and adds dedicated requirements files for various data collectors (including baostock, cn\_index, us\_index, and yahoo). It also establishes a new \pyproject.toml\ for the core package, defining build requirements and dependencies like \mlflow\<3.13\ and \pandas\>=1.1\, while adding Node.js tooling (\@commitlint/cli\) for CI commit validation.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Score
- CAI 44 → 61 (+17.0)
- Rubric changed (rubric-2026.08.15 → rubric-2026.09.15) — scores are not directly comparable.
Lenses
- Code Health 90 → 87 (-3.1)
- Architecture 94 → 83 (-10.8)
- Maturity 50 → 49 (-1.3)
- Readiness 23 → 61 (+37.4)
- Security 58 → 72 (+14.0)
- Domain Modelling 100 (new)
Resolved (76)
- Coverage not measured — test suite did not build
- Dimension evaluation failed
- Duplicated block (10 lines × 3) (qlib/contrib/data/loader.py)
- Duplicated block (11 lines × 2) (qlib/contrib/strategy/rule_strategy.py)
- Duplicated block (12 lines × 3) (qlib/contrib/data/highfreq_handler.py)
- Duplicated block (12 lines × 4) (qlib/workflow/recorder.py)
- Duplicated block (13 lines × 2) (examples/nested_decision_execution/workflow.py)
- Duplicated block (13 lines × 6) (qlib/contrib/data/loader.py)
- Duplicated block (14 lines × 2) (qlib/contrib/model/pytorch_tra.py)
- Duplicated block (14 lines × 2) (qlib/contrib/online/operator.py)
- Duplicated block (15 lines × 2) (qlib/data/data.py)
- Duplicated block (18 lines × 3) (qlib/contrib/data/loader.py)
- Duplicated block (19 lines × 2) (qlib/contrib/data/loader.py)
- Duplicated block (5 lines × 2) (examples/orderbook_data/example.py)
- Duplicated block (6 lines × 2) (examples/highfreq/highfreq_handler.py)
- Duplicated block (6 lines × 3) (qlib/contrib/model/pytorch_tcts.py)
- Duplicated block (8 lines × 2) (qlib/contrib/data/highfreq_handler.py)
- Duplicated block (8 lines × 2) (qlib/contrib/model/pytorch_hist.py)
- Duplicated block (8 lines × 2) (qlib/contrib/model/pytorch_tcts.py)
- Duplicated block (8 lines × 2) (qlib/contrib/model/pytorch_tra.py)
- …and 56 more
New (464)
- ACStrategy.generate_trade_decision (cognitive 25) (qlib/contrib/strategy/rule_strategy.py)
- ADARNN.train_AdaRNN (cognitive 20) (qlib/contrib/model/pytorch_adarnn.py)
- AdaRNN.forward_pre_train (cognitive 18) (qlib/contrib/model/pytorch_adarnn.py)
- Alpha158DL.get_feature_config (cognitive 69) (qlib/contrib/data/loader.py)
- Alpha158DL.get_feature_config (cyclomatic 37) (qlib/contrib/data/loader.py)
- Ambiguous reset semantics. Account has two reset methods. reset_report implies resetting only metrics, while reset implies a full state reset. However, reset takes more arguments (including port_metr_enabled) than reset_report, suggesting reset might be a re-initializer rather than a simple reset. This inconsistency in naming ('reset' vs 'reset_report') and argument complexity is confusing.
- CSIIndex._read_change_from_url (cognitive 21) (scripts/data_collector/cn_index/collector.py)
- ClientDatasetProvider.dataset (cognitive 21) (qlib/data/data.py)
- Concentrated knowledge decay
- Critical CVE: [GHSA redacted] (scripts/data_collector/br_index/requirements.txt)
- DNNModelPytorch.init (cognitive 19) (qlib/contrib/model/pytorch_nn.py)
- DNNModelPytorch.fit (cognitive 49) (qlib/contrib/model/pytorch_nn.py)
- DNNModelPytorch.fit (cyclomatic 21) (qlib/contrib/model/pytorch_nn.py)
- DataHealthChecker.check_large_step_changes (cognitive 16) (scripts/check_data_health.py)
- DiskDatasetCache.update (cognitive 21) (qlib/data/cache.py)
- Documentation: no architecture or design documentation (README.md)
- DumpPitData._dump_pit (cognitive 38) (scripts/dump_pit.py)
- DumpPitData._dump_pit (cyclomatic 17) (scripts/dump_pit.py)
- Duplicated block (10 lines × 2) (qlib/contrib/model/pytorch_localformer.py)
- Duplicated block (10 lines × 2) (qlib/contrib/strategy/rule_strategy.py)
- …and 444 more
Changes since last survey
- 2 commits — 1 feature/other, 1 fixes
By area
- (root) — 1 commit
- .github/workflows — 1 commit
Notable commits
- fix: ci: fix dependency compatibility failures (#2308)
- change: ci: pin GitHub Actions to full-length commit SHAs (#2318)
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
microsoft/qlib was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 26 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit be725493eb1a6bbb42bf11b37aa7669f59610ff1 — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-09659c52afae.