Skip to content
CAI
Software that uses CAICheck a score

pyg-team/pytorch_geometric

69.6

Adequate · 19 September 2026

120.4k

lines of production code

Python

primary language

1

measurement over time

CAI band scale
CAI lens gauges

What this system is

This system is PyTorch Geometric, a library for deep learning on irregularly structured data such as graphs, point clouds, and molecules. It provides a comprehensive suite of neural network layers, models, and data loaders for training graph neural networks, alongside specialized tools for explainability, knowledge graph embeddings, and large-scale distributed training. The codebase also includes extensive benchmarking utilities, integration with modern frameworks like PyTorch Lightning and Hugging Face, and support for hardware acceleration across various devices.

How it got here

2017–2019 — PyG 2.9 architecture and benchmarking

41 changes.

This period focused on the PyG 2.9 release, introducing core architectural improvements such as specialized tensor subclasses, a new configuration system, and a modular package structure for convolutions, aggregations, and normalization. The work also expanded the library with extensive new datasets, neural network modules, and a comprehensive benchmark suite for evaluating model performance and runtime across various hardware platforms.

2020–2022 — GraphGym, explainability, and benchmarking infrastructure

41 changes.

This period focused on establishing robust infrastructure for experiment management, model interpretation, and performance evaluation. Key developments included the introduction of GraphGym for configurable, Lightning-based training workflows, a unified explainability API with heterogeneous graph support, and comprehensive benchmarking suites for training and inference. The work also expanded the library's capabilities with new data loading abstractions, knowledge graph embedding models, and synthetic graph generators.

2023–2026 — LLM integration and distributed training expansion

33 changes.

This period focused on integrating Large Language Models with graph neural networks through the new torch\_geometric.llm module and RAG utilities, alongside significant expansions in distributed training support via cuGraph, Kùzu, and GraphLearn. The work also introduced advanced attention mechanisms, graph coarsening operators, and experimental components in the contrib package, all supported by comprehensive test coverage and new usage examples.

Features

Add BRO and Gini functional operators for graph explainability

New functional operators \bro\ (Batch Representation Orthogonality) and \gini\ (Gini coefficient) are now available in \torch\_geometric.nn.functional\. These functions implement regularization penalties described in the paper 'Improving Molecular Graph Neural Network Explainability with Orthonormalization and Induced Sparsity', allowing users to compute orthogonality penalties for graph representations and sparsity penalties for weight matrices directly.

_torch\geometric/nn/functional · high confidence

Add C++ example for loading PyG models with TorchScatter and TorchSparse

A new C++ example has been added to demonstrate loading a PyTorch Geometric (PyG) model using the C++ API. The example includes a Python script to save a scripted GIN model and a C++ application that loads this model via \torch::jit::load\. The CMake configuration explicitly links against TorchScatter and TorchSparse libraries, enabling users to run graph neural network inference in C++ with support for scatter and sparse operations.

examples/cpp · high confidence

Add DGL runtime benchmark suite

A new benchmark suite has been added to measure the training runtime of Graph Neural Networks using the DGL library. The suite includes implementations of GCN, GAT, and RGCN models (both standard and SPMV variants) and provides a main script to benchmark these models on citation datasets (Cora, CiteSeer, PubMed) and the MUTAG dataset. It supports execution on CPU, CUDA, and Apple Silicon (MPS) devices, with a warm-up phase and timing synchronization to ensure accurate performance measurements.

benchmark/runtime/dgl · high confidence

Add GraphLearn-for-PyTorch distributed training example

This location introduces a new example demonstrating how to use GraphLearn-for-PyTorch (GLT) to perform distributed GNN training with PyG. The change adds a Python script (\dist\_train\_sage\_supervised.py\) that leverages GLT's \DistNeighborLoader\ for efficient multi-node sampling, a configuration file (\dist\_train\_sage\_sup\_config.yml\) for cluster setup, a data partitioning utility (\partition\_ogbn\_dataset.py\), and a launcher script (\launch.py\) to orchestrate training across nodes. It also includes a \README.md\ explaining the requirements and steps to run the distributed training on datasets like \ogbn-products\.

_examples/distributed/graphlearn\_for\pytorch · high confidence

Add JIT compilation examples for GCN, GAT, GIN, and FiLM models

New example scripts have been added to the \examples/jit\ directory demonstrating how to apply PyTorch JIT scripting to several Graph Neural Network models. The directory now includes \gcn.py\, \gat.py\, \gin.py\, and \film.py\, each showing a complete training loop with a model wrapped in \torch.jit.script\. A README has also been added to document these examples and their corresponding model implementations.

examples/jit · high confidence

Add Kùzu remote backend example for large-scale graph training

Users can now train GraphSAGE models on large graphs (e.g., ogbn-papers100M) that do not fit in CPU memory by using Kùzu as a remote backend. This new example in \examples/distributed/kuzu\ demonstrates how to integrate Kùzu's \FeatureStore\ and \GraphStore\ with PyG's \NeighborLoader\, including scripts to prepare the data and train the model on a single machine.

examples/distributed/kuzu · high confidence

Add MGNAN and RBCD attack models to PyG

The \torch\_geometric.contrib.nn\ subpackage now exposes the MGNAN model, an extension of Graph Neural Additive Networks that supports multivariate shape functions for joint feature processing, and the Projected/Greedy Randomized Block Coordinate Descent (PRBCD) adversarial attack models for perturbing graph adjacency matrices. These components are available via \torch\_geometric.contrib.nn.models\ for users requiring advanced interpretability or robustness evaluation capabilities.

_torch\geometric/contrib/nn · high confidence

Add PyTorch Ignite example for GIN model

A new example demonstrating how to implement a Graph Isomorphism Network (GIN) using PyTorch Ignite has been added to the examples directory. The \gin.py\ script showcases the integration of PyG with PyTorch Ignite, including training, validation, and testing loops, along with features like progress bars, TensorBoard logging, and model checkpointing based on validation accuracy.

_examples/pytorch\ignite · high confidence

Add PyTorch compile examples for GCN and GIN models

New example scripts have been added to the \examples/compile\ directory demonstrating how to use PyTorch's \torch.compile\ with PyTorch Geometric models. The GCN example (\gcn.py\) shows static compilation (\dynamic=False\) on the Planetoid dataset, while the GIN example (\gin.py\) demonstrates dynamic shape compilation (\dynamic=True\) on the TUDataset, requiring PyTorch 2.1.0 or later.

examples/compile · high confidence

Add Quiver integration examples for single- and multi-GPU GraphSage training

New example scripts (\single\_gpu\_quiver.py\ and \multi\_gpu\_quiver.py\) and documentation (\README.md\) are added to \examples/quiver\ to demonstrate how to use the Quiver library (\torch-quiver\) with PyG. These examples show how to leverage Quiver for GPU-optimized graph sampling and feature aggregation, including caching hot GNN data in GPU memory, for both single-GPU and multi-GPU distributed training scenarios using the GraphSage model on the Reddit dataset.

examples/quiver · high confidence

Add cuGraph-accelerated GAT, SAGE, and RGCN convolution layers

This change introduces new \CuGraphGATConv\, \CuGraphSAGEConv\, and \CuGraphRGCNConv\ layers in \torch\_geometric.nn.conv.cugraph\. These modules provide optimized versions of the standard GAT, SAGE, and RGCN convolutions by leveraging the \pylibcugraphops\ library to fuse message passing computations, resulting in accelerated execution and lower memory footprint on GPU. The implementation includes a base \CuGraphModule\ for handling graph construction from \EdgeIndex\ and supports both legacy and new \pylibcugraphops\ APIs.

_torch\geometric/nn/conv/cugraph · high confidence

Add examples for PGM Explainer and RBCD Attack algorithms

New example scripts have been added to the \examples/contrib\ directory to demonstrate the usage of the \PGMExplainer\ for both node and graph classification tasks, as well as the \GRBCDAttack\ and \PRBCDAttack\ algorithms for graph adversarial attacks (evasion and poisoning). These examples illustrate how to integrate these experimental \torch\_geometric.contrib\ modules with standard PyTorch Geometric datasets and models.

examples/contrib · high confidence

Add knn\_interpolate function for point cloud upsampling

The \torch\_geometric.nn.unpool\ module now exposes a \knn\_interpolate\ function that implements k-NN interpolation for transferring features from source points to target points based on spatial proximity. This allows users to upsample point cloud features by weighting the features of the k nearest neighbors inversely proportional to their squared distance, supporting batched processing and configurable worker counts for performance.

_torch\geometric/nn/unpool · high confidence

Add multi-GPU training benchmarks for CUDA and Intel XPU devices

New benchmark scripts and documentation have been added to evaluate multi-GPU training performance on both NVIDIA CUDA and Intel XPU hardware. The \training\_benchmark\_cuda.py\ script uses PyTorch's NCCL backend for distributed training, while \training\_benchmark\_xpu.py\ leverages Intel's CCL backend and IPEX optimization for XPU devices. Both scripts support heterogeneous graph datasets (like ogbn-mag) and homogeneous datasets (like ogbn-products), allowing users to benchmark scalability across multiple GPUs using models such as GAT, GCN, and PNA.

_benchmark/multi\gpu · high confidence

Add point cloud classification benchmark suite

Introduces a new benchmark directory for point cloud classification on the ModelNet10 dataset, providing evaluation scripts for MPNN, PointNet++, EdgeCNN, SplineCNN, and PointCNN. The suite includes a shared dataset loader and a training/evaluation engine that supports inference mode, performance profiling, bf16 precision, and model compilation.

benchmark/points · high confidence

Add runtime benchmark suite for PyG models

A new runtime benchmark suite has been added to the \benchmark/runtime\ directory, allowing users to measure and compare the training performance of Graph Neural Network models. The suite includes implementations of GCN, GAT, and RGCN models, along with a training script (\train.py\) that times execution on CPU, CUDA, and Apple Silicon (MPS) devices. Users can run the benchmarks via \python main.py\ to evaluate model throughput on standard datasets like Cora, CiteSeer, PubMed, and MUTAG.

benchmark/runtime · high confidence

Add utility to convert RelBench databases to HeteroData

Users can now convert RelBench database objects into PyTorch Geometric HeteroData graphs using the new \from\_relbench\ utility in \torch\_geometric.contrib.utils\. Each table in the database becomes a node type, with numeric columns converted into node features and time columns stored as time attributes. Foreign key relationships are automatically transformed into bidirectional edge types, allowing seamless integration of RelBench datasets into heterogeneous graph workflows.

_torch\geometric/contrib/utils · high confidence

Added NeighborLoader benchmark suite

A new benchmark script for the NeighborLoader has been added to the benchmark suite. This tool allows users to measure sampling performance across various configurations, including different datasets (OGB MAG and OGBN datasets), batch sizes, neighbor sizes, and subgraph types (directional vs. bidirectional). It also supports advanced features such as CPU affinity for DataLoader workers and integration with torch.profiler for detailed performance profiling.

benchmark/loader · high confidence

GraphGym contrib module auto-discovery and layer implementation

The GraphGym contrib package now automatically discovers and exposes submodules (activation, config, encoder, head, layer, loader, loss, network, optimizer, pooling, stage, train, and transform) by scanning their directories for Python files. This change introduces a new \GeneralConvLayer\ and \GeneralEdgeConvLayer\ in the layer module, which implement configurable message passing with support for self-message concatenation, edge features, and adjacency normalization, allowing users to define custom GNN architectures via configuration.

_torch\geometric/graphgym/contrib · high confidence

GraphGym custom component examples and registration

The \graphgym/custom\graphgym\ package now includes a comprehensive set of example implementations for all major GraphGym extension points, including activation functions, configuration schemas, node/edge encoders, GNN layers, network architectures, heads, pooling strategies, loss functions, optimizers, schedulers, training stages, data loaders, and transforms. These examples demonstrate how to register custom components using the \@register\\*\ decorators and provide reference implementations for users building custom GraphGym pipelines.

_graphgym/custom\graphgym · high confidence

Introduce GraphGym model components and registration system

Adds the core GraphGym model architecture, including the GNN backbone, feature encoders (Integer, Atom, Bond), prediction heads (node, edge, graph), and layer abstractions (GeneralLayer, GeneralMultiLayer). Introduces a registration system for layers, activations, heads, and pooling functions, allowing users to configure and extend model components via configuration files rather than hard-coded classes.

_torch\geometric/graphgym/models · high confidence

Introduce GraphGym utility module

A new \torch\_geometric.graphgym.utils\ package has been added, providing core helper functions for the GraphGym experiment framework. This module includes utilities for aggregating and parsing run results across seeds (\agg\_runs\, \agg\_batch\), managing computational budgets and parameter counts (\comp\_budget\), auto-selecting GPU devices (\device\), determining evaluation and checkpoint epochs (\epoch\), handling I/O operations like JSON and TensorBoard logging (\io\), visualizing embeddings (\plot\), and providing a dummy context manager (\tools\).

_torch\geometric/graphgym/utils · high confidence

Introduce Knowledge Graph Embedding (KGE) models

A new \torch\_geometric.nn.kge\ package is introduced, providing implementations of four standard Knowledge Graph Embedding models: TransE, ComplEx, DistMult, and RotatE. These models share a common \KGEModel\ base class that handles entity and relation embeddings, and includes built-in utilities for training via \KGTripletLoader\, evaluating performance (Mean Rank, MRR, Hits@k), and generating negative samples. Users can now easily integrate these specific link prediction algorithms into their PyTorch Geometric workflows.

_torch\geometric/nn/kge · high confidence

Introduce RAG utilities for LLM integration

Added new utility classes in \torch\_geometric.llm.utils\ to support Retrieval-Augmented Generation (RAG) workflows. This includes \KNNRAGFeatureStore\ for KNN-based node retrieval, \NeighborSamplingRAGGraphStore\ for neighbor-sampling based graph storage, and \DocumentRetriever\ for retrieving documents from a vector database. These components provide the backend infrastructure for integrating large language models with graph data structures.

_torch\geometric/llm/utils · high confidence

Introduce benchmark/kernel for graph classification model evaluation

The benchmark/kernel directory now provides a complete evaluation suite for graph classification tasks, featuring reference implementations of GCN, GraphSAGE, GIN, Graclus, Top-K Pooling, SAG Pooling, DiffPool, EdgePool, GlobalAttention, Set2Set, SortPool, and ASAPool. The suite supports 10-fold cross-validation on standard datasets (MUTAG, PROTEINS, IMDB-BINARY, REDDIT-BINARY) and includes a performance measurement script with optional profiling, bf16 support, and torch.compile integration.

benchmark/kernel · high confidence

Introduce inference benchmark suite with CSV export and XPU support

Adds a new inference benchmark tool (\benchmark/inference/inference\_benchmark.py\) that allows users to measure performance across various graph neural network models and datasets. The benchmark supports saving results to CSV files, enables execution on Intel XPU devices (requiring IPEX), and includes options for sparse tensor usage, full-batch inference, and configurable batch sizes and layers. Documentation is provided in \benchmark/inference/README.md\ detailing environment setup, including jemalloc configuration for performance optimization.

benchmark/inference · high confidence

A new \torch\_geometric.metrics\ package is added, exposing a suite of link prediction evaluation metrics including precision, recall, F1, MAP, NDCG, MRR, hit ratio, coverage, diversity, personalization, and average popularity. These metrics support configurable top-k evaluation, handle edge label weights (including ignoring negative weights), and provide a collection interface for aggregating multiple metrics simultaneously.

_torch\geometric/metrics · high confidence

Introduce new \`torch\_geometric.sampler\` package with unified sampling interface

A new \torch\_geometric.sampler\ package has been added, providing a unified, input-type-agnostic interface for graph sampling. This package introduces base classes (\BaseSampler\, \NodeSamplerInput\, \EdgeSamplerInput\, \SamplerOutput\) and specific implementations like \NeighborSampler\ and \HGTSampler\. It supports heterogeneous graphs, temporal sampling strategies, bidirectional subgraph sampling, and weighted sampling, serving as the underlying engine for loaders like \NeighborLoader\ and \HGTLoader\.

_torch\geometric/sampler · high confidence

Introduce structured motif generator framework for explainability benchmarks

The \torch\_geometric.datasets.motif\_generator\ package has been introduced to provide a standardized way of generating synthetic graph motifs used in explainability benchmarks like GNNExplainer. This new module includes a base \MotifGenerator\ class with a resolver for easy instantiation, alongside specific implementations for \HouseMotif\, \CycleMotif\, \GridMotif\, and \CustomMotif\. Users can now easily generate these predefined graph structures or define custom ones to create benchmark datasets for evaluating graph neural network explanations.

_torch\_geometric/datasets/motif\generator · high confidence

Introduce torch\_geometric.llm module with new LLM-based models

The new \torch\_geometric.llm\ package provides a collection of models that integrate Large Language Models with graph neural networks. This release adds the \LLM\ wrapper for HuggingFace models with automatic GPU memory management, \GRetriever\ for retrieval-augmented generation on graphs, \GLEM\ for co-training GNNs and language models, \GITMol\ for multi-modal molecular science, \MoleculeGPT\ for instruction-following molecular property prediction, \ProteinMPNN\ for protein sequence design, \TXT2KG\ for text-to-knowledge-graph generation, \LLMJudge\ for evaluating answers via NVIDIA NIMs, and utility classes like \LargeGraphIndexer\ and \RAGQueryLoader\.

_torch\geometric/llm/models · high confidence

Introduce torch\_geometric.profile package for GNN benchmarking and profiling

The new \torch\_geometric.profile\ package provides a suite of tools for benchmarking and profiling GNN models. It includes a \benchmark\ function to compare multiple functions with support for forward/backward passes and progress bars, a \profileit\ decorator for detailed runtime and memory statistics on CUDA and XPU devices, and a \Profiler\ class for layer-by-layer profiling using PyTorch's profiler. The package also exposes utility functions for counting parameters, measuring model and data sizes, and retrieving CPU/GPU memory usage via garbage collection, \nvidia-smi\, or Intel IPEX.

_torch\geometric/profile · high confidence

Introduce torch\_geometric.transforms package

The \torch\_geometric.transforms\ module is now a dedicated package, consolidating graph and data transformation utilities into a single importable namespace. This change introduces a comprehensive suite of transforms—including graph structure modifiers (e.g., \AddSelfLoops\, \ToUndirected\), positional encodings (e.g., \AddLaplacianEigenvectorPE\, \AddGPSE\), and spatial/point-cloud processors (e.g., \Delaunay\, \SamplePoints\)—along with composition tools like \Compose\ and \ComposeFilters\. Additionally, \RandomTranslate\ is deprecated in favor of \RandomJitter\ to align with updated naming conventions.

_torch\geometric/transforms · high confidence

Introduce training benchmark suite with multi-device and sparse tensor support

A new training benchmark script and documentation have been added to the \benchmark/training\ directory, allowing users to measure performance metrics for graph neural network training. The benchmark supports multiple devices including CPU, CUDA, MPS, and XPU (via IPEX), and enables the use of SparseTensors for optimized data handling. It provides configurable options for models, datasets, batch sizes, and custom step counts, while also supporting CSV output for benchmark results and PyTorch profile data.

benchmark/training · high confidence

Introduce unified \`torch\_geometric.explain\` API with configuration and heterogeneous support

The \torch\_geometric.explain\ module is introduced, providing a structured API for GNN explainability. It adds \ExplainerConfig\, \ModelConfig\, and \ThresholdConfig\ to standardize explanation parameters, and introduces \Explanation\ and \HeteroExplanation\ classes to handle both homogeneous and heterogeneous graph attributions. The new \Explainer\ class orchestrates explanation algorithms, supporting binary classification modes and offering methods to visualize feature importance and graph subgraphs.

_torch\geometric/explain · high confidence

Introduction of PyTorch Lightning integration for graph data loading

This change introduces the \torch\_geometric.data.lightning\ module, providing \LightningDataset\, \LightningNodeData\, and \LightningLinkData\ classes that integrate PyTorch Geometric data structures with PyTorch Lightning's \LightningDataModule\. The implementation supports various loading strategies including full-batch, neighbor sampling, and link neighbor sampling, while handling configuration for validation and test dataloaders. It also includes logic to automatically detect and use either the \lightning\ or \pytorch\_lightning\ package, ensuring compatibility with different installation environments.

_torch\geometric/data/lightning · high confidence

Introduction of empty contrib subpackages for datasets and transforms

The \torch\geometric.contrib\ package now includes \datasets\ and \transforms\ subpackages. Currently, these are empty stubs (containing only empty \\\all\\_\ lists) that serve as placeholders for future contributions, meaning no new functionality or behavioral changes are available to users at this time.

_torch\_geometric/contrib/datasets, torch\geometric/contrib/transforms · high confidence

Introduction of experimental torch\_geometric.contrib subpackage

A new \torch\_geometric.contrib\ subpackage has been added to expose experimental components, including transforms, datasets, neural network modules, explainability tools, and utilities. Because this code is experimental and subject to change, importing it triggers a runtime warning advising users to proceed with caution.

_torch\geometric/contrib · high confidence

Introduction of new GNN model implementations and model registry

The \torch\_geometric.nn.models\ package now exposes a comprehensive set of new graph neural network architectures, including \AttentiveFP\, \ARLinkPredictor\ (Attract-Repel), \BasicGNN\ (with concrete implementations like \GCN\, \GraphSAGE\, \GIN\, \GAT\, \PNA\, and \EdgeCNN\), \CorrectAndSmooth\, \DeepGCNLayer\, \DimeNet\ (and \DimeNetPlusPlus\), \GPSE\, \MetaPath2Vec\, \RENet\, \SignedGCN\, \TGNMemory\, and \ViSNet\. These models are now available for direct import from the \torch\geometric.nn\ namespace via the updated \\\init\\_.py\ registry, enabling users to easily instantiate these specific architectures for molecular representation learning, link prediction, and semi-supervised node classification tasks.

_torch\geometric/nn/models · high confidence

Introduction of new PyTorch Geometric neural network modules and utilities

This release introduces several new components to the \torch\_geometric.nn\ package. Users can now use \PositionalEncoding\ and \TemporalEncoding\ for adding positional and temporal information to graph data. A new \Sequential\ module allows for the definition of complex GNN pipelines with explicit input/output signatures, while \DataParallel\ provides a wrapper for multi-GPU training with a deprecation warning recommending \DistributedDataParallel\. The package also adds \PyGModelHubMixin\ for saving and loading models to the Hugging Face Hub, \summary\ for model introspection, and custom \ModuleDict\ and \ParameterDict\ classes that support dots and tuples in keys. Additionally, new learning rate schedulers (\ConstantWithWarmupLR\, \CosineWithWarmupLR\, etc.) and resolvers for activations, normalizations, aggregations, optimizers, and schedulers are available to simplify model configuration.

_torch\geometric/nn · high confidence

Introduction of new data loading and sampling infrastructure

The \torch\_geometric.loader\ module has been restructured to introduce a new, unified data loading architecture. This includes new base classes like \DataLoaderIterator\ for efficient post-processing in the main process, and specialized loaders such as \CachedLoader\ for batching output caching, \DynamicBatchSampler\ for memory-aware dynamic batching, and \DataListLoader\ for list-based batching. The module now exposes a comprehensive suite of samplers and loaders including \NeighborLoader\, \LinkLoader\, \HGTLoader\, \ClusterData\, \GraphSAINTSampler\, \ImbalancedSampler\, and \ZipLoader\, alongside utility mixins like \AffinityMixin\ for CPU affinity control. Additionally, \RandomNodeSampler\ is deprecated in favor of \RandomNodeLoader\.

_torch\geometric/loader · high confidence

Introduction of new pooling layers and functions in \`torch\_geometric.nn.pool\`

The \torch\_geometric.nn.pool\ package has been restructured and expanded with several new graph pooling operators and utility functions. New pooling layers include \ClusterPooling\ (edge-based graph component pooling), \MemPooling\ (memory-based graph networks), \ASAPooling\ (adaptive structure-aware pooling), and \EdgePooling\ (edge contraction pooling). Additionally, the package now exports \KNNIndex\ classes for fast k-nearest neighbor search via FAISS, approximate k-NN functions (\approx\_knn\, \approx\_knn\_graph\) using pynndescent, and a \decimation\_indices\ utility for point cloud downsampling. Existing global pooling functions (\global\_add\_pool\, \global\_mean\_pool\, \global\_max\_pool\) and functional pooling utilities (\avg\_pool\, \max\_pool\, \avg\_pool\_x\, \max\_pool\_x\) are also included in this module.

_torch\geometric/nn/pool · high confidence

Introduction of the PyG Benchmark Suite

A new benchmark suite has been added to provide evaluation scripts for semi-supervised node classification, graph classification, point cloud classification, and runtime comparisons. Users can now install the suite via pip to compare various methods in homogeneous evaluation scenarios, with a specific focus on avoiding hyperparameter and model selection on the test set by utilizing an additional validation set.

benchmark · high confidence

Introduction of the \`torch\_geometric.nn.aggr\` package for graph aggregation

The library now provides a dedicated \torch\_geometric.nn.aggr\ package that centralizes graph aggregation operations. This change introduces a base \Aggregation\ class and a comprehensive suite of aggregation layers—including basic operators (Sum, Mean, Max, Min, Mul, Var, Std, Softmax, PowerMean), sequence-based models (LSTM, GRU), attention-based methods (Attentional, Set2Set, GraphMultisetTransformer, SetTransformer), and specialized layers (Equilibrium, Sort, LCM, DeepSets, MLPAggregation, VariancePreserving, PatchTransformer). These layers are unified under a common interface, allowing users to easily compose multiple aggregations via \MultiAggregation\ and \FusedAggregation\ for optimized performance.

_torch\geometric/nn/aggr · high confidence

New Docker and Singularity container definitions for NVIDIA and Intel GPUs

Added new container build files to support running PyTorch Geometric on NVIDIA GPUs (via \docker/Dockerfile\ based on the NGC PyG 24.09 image with CUDA 12.6) and Intel GPUs (via \docker/Dockerfile.xpu\ based on Intel Extension for PyTorch 2.1.30). The update also includes a \docker/README.md\ with usage instructions for these new images and a \docker/singularity\ file for Singularity container support.

docker · high confidence

New GraphGym batch execution and configuration generation tools

Users can now generate hyperparameter grid configurations and run experiments in parallel using new shell scripts and Python utilities. The \configs\_gen.py\ script creates YAML configuration files based on grid search definitions, while \run\_batch.sh\ and \parallel.sh\ automate the execution of these configurations across multiple jobs. Additionally, \agg\_batch.py\ provides a command-line interface to aggregate results from batch runs, and \run\_single.sh\ offers a quick way to test individual experiment setups.

graphgym · high confidence

New LLM and GNN+LLM example scripts and documentation

The \examples/llm\ directory now includes a comprehensive set of runnable examples for integrating Large Language Models with Graph Neural Networks. The new \README.md\ documents these examples, which include \g\_retriever.py\ for G-retriever integration (with Neo4j support), \git\_mol.py\ for the GIT-Mol multimodal molecular model, \glem.py\ for GLEM co-training on text-attributed graphs, \molecule\_gpt.py\ for instruction-following molecular property prediction, \protein\_mpnn.py\ for protein sequence design, \relbench\_gretriever.py\ for RelBench dataset integration, and \txt2kg\_rag.py\ and \txt2qa.py\ for end-to-end RAG pipelines using TXT2KG and synthetic QA generation.

examples/llm · high confidence

New PyTorch Lightning examples for graph classification and heterogeneous node classification

Added new example scripts (\gin.py\, \graph\_sage.py\, \relational\_gnn.py\) and a \README.md\ to the \examples/pytorch\_lightning\ directory. These examples demonstrate integrating PyG with PyTorch Lightning using \LightningDataset\ for graph classification (GIN on TUDataset) and \LightningNodeData\ for node classification (GraphSAGE on Reddit and Relational GNN on OGB\_MAG). The examples showcase features like heterogeneous graph support via \to\_hetero\, neighbor sampling loaders, and standard Lightning training loops with accuracy logging and checkpointing.

_examples/pytorch\lightning · high confidence

New and updated explanation examples for PyG's explain package

The \examples/explain\ directory now provides a comprehensive set of scripts demonstrating the \torch\_geometric.explain\ API. New examples include \captum\_explainer.py\ and \captum\_explainer\_hetero\_link.py\ for node and heterogeneous link prediction using Captum, \graphmask\_explainer.py\ for GraphMaskExplainer, and \mgnan\_graph\_mutagenicity.py\ for graph-level explanations with M-GNAN. Existing scripts like \gnn\_explainer\_link\_pred.py\ have been updated to support MPS devices and demonstrate both model and phenomenon explanation types for link prediction. A README has been added to document all available examples.

examples/explain · high confidence

New attention modules added to torch\_geometric.nn.attention

The \torch\geometric.nn.attention\ package now exposes four new attention mechanisms: \PerformerAttention\ (linear-time attention via random feature projection), \PolynormerAttention\ (polynomial-expressive linear attention), \SGFormerAttention\ (global attention with normalization and epsilon handling), and \QFormer\ (a simplified querying transformer encoder). These classes are registered in the \\\init\\_.py\ module and can be imported directly from \torch\_geometric.nn.attention\.

_torch\geometric/nn/attention · high confidence

New benchmark utility module for model and dataset handling

A new \benchmark/utils\ package has been added to centralize benchmarking infrastructure. It provides helper functions to load datasets (OGB-MAG, OGB-products, Reddit) with optional sparse tensor and bfloat16 support, instantiate various GNN models (GAT, GCN, PNA, GraphSAGE, and their heterogeneous variants), and manage data splits. The module also includes utilities for saving benchmark results to CSV files and a generic evaluation function for both homogeneous and heterogeneous graph data.

benchmark/utils · high confidence

New citation benchmark suite with profiling, bf16, and compile support

Added a new benchmark suite for semi-supervised node classification on Cora, CiteSeer, and PubMed datasets. The suite includes scripts for GCN, GAT, Chebyshev, SGC, ARMA, and APPNP models, along with shell scripts to run training and inference evaluations. The benchmark infrastructure now supports profiling via torch\_geometric.profile, mixed-precision inference with bf16, and model compilation via torch.compile, allowing users to measure performance and optimize model execution.

benchmark/citation · high confidence

New dataset cheatsheet utility functions

Added a new \torch\_geometric.datasets.utils\ module containing helper functions (\paper\_link\, \has\_stats\, \get\_stat\, \get\_children\, \get\_type\) that parse dataset class docstrings to extract paper links, dataset types, and statistical information. This enables automated generation of dataset cheatsheets and documentation by exposing structured data from existing dataset definitions.

_torch\geometric/datasets/utils · high confidence

New datasets and synthetic data generators added to torch\_geometric.datasets

The \torch\geometric.datasets\ module has been expanded with a large number of new dataset classes and synthetic data generators. New real-world datasets include \Actor\, \AirfRANS\, \Airports\, \AmazonBook\, \AMiner\, \AQSOL\, \AttributedGraphDataset\, \AmazonProducts\, and many others covering domains such as social networks, molecular graphs, point clouds, and heterogeneous graphs. Additionally, new synthetic datasets for explainability and benchmarking have been added, including \BA2MotifDataset\, \ExplainerDataset\, \InfectionDataset\, \MixHopSyntheticDataset\, \BAMultiShapesDataset\, and \BAShapes\. These additions are exposed via the \\\init\\_.py\ module, making them immediately available for import and use.

_torch\geometric/datasets · high confidence

New dense graph neural network layers and pooling operators

The \torch\_geometric.nn.dense\ package now provides dense-tensor equivalents for several graph neural network components, allowing operations on fixed-size adjacency matrices. This includes convolution layers \DenseGCNConv\, \DenseGATConv\, \DenseSAGEConv\, \DenseGraphConv\, and \DenseGINConv\, as well as pooling operators \dense\_diff\_pool\, \dense\_mincut\_pool\, and \DMoNPooling\. These modules support batched inputs with masking and are exported from the new \torch\_geometric.nn.dense\ namespace.

_torch\geometric/nn/dense · high confidence

The examples directory now includes new scripts demonstrating advanced graph neural network techniques. \ar\_link\_pred.py\ introduces an Attract-Repel embedding approach for link prediction, which can significantly boost accuracy compared to traditional methods. \agnn.py\ provides a simple example of an Attention-based Graph Neural Network (AGNN) for node classification. Additionally, \argva\_node\_clustering.py\ demonstrates variational graph autoencoders for node clustering, including visualization with t-SNE.

examples · high confidence

New explainability algorithm package with heterogeneous graph support

The \torch\_geometric.explain.algorithm\ module introduces a unified framework for graph explanation algorithms, including \GNNExplainer\, \CaptumExplainer\, \PGExplainer\, \AttentionExplainer\, \GraphMaskExplainer\, and \DummyExplainer\. This update adds native support for heterogeneous graphs across these explainers, allowing them to generate \HeteroExplanation\ objects. The \CaptumExplainer\ now supports binary classification tasks and allows specifying attribution methods via string names. Additionally, \PGExplainer\ has a new default bias of 0.01, and \GNNExplainer\ now handles link prediction tasks and heterogeneous graphs.

_torch\geometric/explain/algorithm · high confidence

New explainability evaluation metrics for fidelity, faithfulness, and ground truth

The \torch\_geometric.explain.metric\ module now provides new functions to evaluate explanation quality. Users can assess how well an explanation aligns with the underlying model logic using \fidelity\ (positive and negative fidelity scores) and \characterization\_score\, or measure consistency with the model's predictions via \unfaithfulness\. Additionally, \groundtruth\_metrics\ allows comparing explanation masks against ground-truth masks using standard classification metrics like accuracy, recall, precision, F1-score, and AUROC. These tools enable more rigorous evaluation of GNN explainability methods.

_torch\geometric/explain/metric · high confidence

New filesystem abstraction and dedicated IO module for graph data

The \torch\_geometric.io\ package has been introduced to centralize file reading and writing, backed by a new \fsspec\-based filesystem abstraction (\fs.py\) that enables reading from local, cloud (S3, GCS), and memory paths. This module provides dedicated readers for various formats: \read\_planetoid\_data\ for Planetoid datasets, \read\_tu\_data\ for TUDataset graphs, \read\_npz\ for NPZ files, \read\_obj\ for Wavefront OBJ meshes, \read\_off\ for OFF meshes, \read\_ply\ for PLY meshes (requiring the \openmesh\ library), and \read\_sdf\ for chemical SDF files. The filesystem layer also exposes utilities like \cp\, \mv\, \ls\, and \exists\ to handle data downloads and caching uniformly across different storage backends.

_torch\geometric/io · high confidence

New graph and influence visualization utilities

The library now includes a dedicated visualization package offering \visualize\_graph\ for rendering directed graphs with optional node labels and edge-weight-based styling, \visualize\_hetero\_graph\ for visualizing heterogeneous graphs with configurable node/edge sizing and opacity, and \influence\ for computing and returning normalized gradient-based influence scores for model inputs.

_torch\geometric/visualization · high confidence

New graph coarsening operator for filtering edges in pooling

A new \FilterEdges\ operator has been added to the graph pooling framework to handle edge connections during graph coarsening. This operator filters out edges where incident nodes are not assigned to any cluster, ensuring that only valid connections between supernodes are retained in the coarsened graph structure. The implementation includes base classes (\Connect\, \ConnectOutput\) and the specific \FilterEdges\ logic, which integrates with the existing \Select\ output to produce the final pooled graph representation.

_torch\geometric/nn/pool/connect · high confidence

New graph generator dataset classes for synthetic graph generation

The \torch\_geometric.datasets.graph\_generator\ module now provides a set of dataset classes for generating synthetic graphs, including \BAGraph\ (Barabási–Albert), \ERGraph\ (Erdős–Rényi), \GridGraph\, and \TreeGraph\. These classes inherit from a new \GraphGenerator\ base class, which offers a \resolve\ method for instantiation by name, allowing users to easily create and access these standard graph structures for benchmarking and explainability tasks.

_torch\_geometric/datasets/graph\generator · high confidence

New heterogeneous graph learning examples and documentation

The \examples/hetero\ directory now includes a comprehensive set of reference implementations for heterogeneous graph neural networks, accompanied by a new README.md cataloging the available examples. The collection covers diverse tasks and architectures, including node classification on DBLP and IMDB using HeteroConv and HAN, link prediction on MovieLens using GraphSAGE and temporal models, and unsupervised embedding learning via MetaPath2Vec and DMGI. New examples also demonstrate hierarchical sampling on OGB-MAG, loading heterogeneous graphs from raw CSV data, and building a temporal GNN-based recommender system on MovieLens with k-NN search.

examples/hetero · high confidence

New multi-GPU distributed training examples added

The \examples/multi\_gpu\ directory now includes a comprehensive set of new training scripts demonstrating distributed GNN training on NVIDIA GPUs and Intel XPUs. These examples cover single-node and multi-node scaling using \DistributedDataParallel\ and \NeighborLoader\, including graph-level batching (\distributed\_batching.py\), node-level sampling (\distributed\_sampling.py\), large-scale homogeneous graphs (\papers100m\_gcn.py\), heterogeneous graphs (\mag240m\_graphsage.py\, \taobao.py\), and model parallelism (\model\_parallel.py\). A new README documents these examples and links to cuGraph-accelerated variants, while an \sbatch\ script is provided for multi-node Slurm deployment.

_examples/multi\gpu · high confidence

New normalization layers and modules in torch\_geometric.nn.norm

The \torch\_geometric.nn.norm\ package now exposes a comprehensive suite of normalization layers for graph neural networks. Users can import and use \BatchNorm\, \HeteroBatchNorm\, \InstanceNorm\, \LayerNorm\, \HeteroLayerNorm\, \GraphNorm\, \GraphSizeNorm\, \PairNorm\, \MeanSubtractionNorm\, \MessageNorm\, and \DiffGroupNorm\. These modules provide specialized normalization strategies—such as graph-wise, node-wise, heterogeneous, and message-based normalization—to help accelerate training and mitigate issues like oversmoothing in deep GNN architectures.

_torch\geometric/nn/norm · high confidence

New utility module for inspecting GNN layer capabilities

A new \torch\_geometric.nn.conv.utils\ package has been added, providing a set of helper functions to programmatically inspect the capabilities of GNN convolution layers. Users can now check whether a specific layer supports features such as sparse tensors, edge weights, edge features, bipartite graphs, static graphs, lazy initialization, heterogeneous graphs, hypergraphs, or point clouds by calling functions like \supports\_sparse\_tensor\ or \processes\_heterogeneous\_graphs\. This enables dynamic discovery of layer properties based on their method signatures and documentation.

_torch\geometric/nn/conv/utils · high confidence

PyG 2.9 introduces core tensor subclasses and configuration infrastructure

PyTorch Geometric version 2.9.0 introduces three new tensor subclasses—Index, EdgeIndex, and HashTensor—exposed at the top level, which provide optimized, metadata-aware representations for graph indices and key-value mappings. The release also adds a new configuration system via ConfigMixin and ConfigStore, enabling classes to serialize and deserialize their state to/from dataclasses, and introduces a deprecated torch\_geometric.compile wrapper that now explicitly warns users to use torch.compile directly. Additionally, the package adds utilities for safe ONNX export, device detection (MPS/XPU), and experimental mode flags.

_torch\geometric · high confidence

Architecture

Introduction of the \`torch\_geometric.nn.conv\` package

The convolutional layers and the core \MessagePassing\ base class have been reorganized into a new \torch\_geometric.nn.conv\ sub-package. This change introduces a dedicated module structure for all graph neural network operators (such as \GCNConv\, \GATConv\, and \SAGEConv\), making them accessible via \torch\_geometric.nn.conv\ and updating the internal imports accordingly.

_torch\geometric/nn/conv · high confidence

Behavioural changes

Deprecate \`torch\_geometric.compile()\` and \`MessagePassing.jittable()\`

The \torch\_geometric.compile()\ function and the \MessagePassing.jittable()\ method are now deprecated in favor of native PyTorch compilation and TorchScript support. Users should migrate to \torch.compile()\ and standard TorchScript workflows to ensure compatibility with future PyTorch versions and to remove reliance on PyG-specific compilation wrappers.

(repo-wide) · high confidence

Deprecation of GraphMaskExplainer in contrib and addition of PGMExplainer

The \GraphMaskExplainer\ class in \torch\_geometric.contrib.explain\ is now deprecated and redirects users to \torch\_geometric.explain.algorithm.GraphMaskExplainer\. Additionally, the \PGMExplainer\ algorithm has been added to this module, providing probabilistic graphical model explanations for graph neural networks.

_torch\geometric/contrib/explain · high confidence

GraphGym training workflow now uses PyTorch Lightning

The GraphGym training entry point has been refactored to use PyTorch Lightning. The \train\ function now instantiates a \pl.Trainer\ to manage the training loop, validation, and testing, replacing the previous manual iteration logic. A new \GraphGymDataModule\ handles data loading, and the \GraphGymModule\ model class inherits from \LightningModule\ to integrate with the Lightning ecosystem. This change introduces a dependency on \pytorch\_lightning\ (or the \lightning\ package) and alters how training is executed, including checkpointing and logging via Lightning callbacks.

_torch\geometric/graphgym · high confidence

Introduce SelectTopK node selection with TorchScript support

The \TopKPooling\ component has been refactored into a new \SelectTopK\ class within the \torch\_geometric.nn.pool.select\ package, providing a structured node-selection mechanism that returns a \SelectOutput\ dataclass containing node indices, cluster assignments, and weights. This change introduces a new base \Select\ class and specific implementation \SelectTopK\, which supports both ratio-based and minimum-score-based node selection, and includes fixes to ensure full TorchScript compatibility for the selection logic.

_torch\geometric/nn/pool/select · high confidence

Introduction of deprecated distributed training components

The \torch\_geometric.distributed\ package has been introduced, providing a suite of classes for distributed graph training including \DistContext\, \LocalFeatureStore\, \LocalGraphStore\, \Partitioner\, \DistNeighborSampler\, \DistLoader\, \DistNeighborLoader\, and \DistLinkNeighborLoader\. These components enable node-level and edge-level sampling across multiple processes via RPC. However, this module is deprecated as of version 2.7.0 and will no longer be maintained; users are directed to the official distributed training tutorials or cuGraph examples for supported alternatives.

_torch\geometric/distributed · high confidence

Refactor of torch\_geometric.utils into modular sub-packages

The \torch\_geometric.utils\ module has been reorganized from a single monolithic file into a structured package with dedicated sub-modules (e.g., \\_scatter.py\, \\_coalesce.py\, \\_softmax.py\, \\_negative\sampling.py\). This change improves code maintainability and performance by allowing specialized implementations for core operations like scatter, coalesce, and softmax, while exposing the same public API through the main \\\init\\_.py\.

_torch\geometric/utils · high confidence

Restructure data module and deprecate loader imports

The \torch\_geometric.data\ package has been reorganized to explicitly expose core data classes (\Data\, \HeteroData\, \Batch\, \TemporalData\), storage abstractions (\FeatureStore\, \GraphStore\), and persistence utilities (\Database\, \OnDiskDataset\). As part of this restructuring, direct imports of data loaders (e.g., \DataLoader\, \NeighborSampler\) from \torch\_geometric.data\ are now deprecated in favor of importing them from \torch\_geometric.loader\. Additionally, safe serialization globals are registered for PyTorch 2.4+ to ensure compatibility when pickling these data objects.

_torch\geometric/data · high confidence

Test coverage

Added comprehensive test coverage for graph convolution layers; Added comprehensive test coverage for graph data loaders; Added comprehensive test coverage for graph utility functions; Added distributed test suite for neighbor sampling, loaders, and stores; Added test coverage for CuGraph GAT, RGCN, and SAGE convolution layers; Added test coverage for dense graph neural network layers and pooling modules; Added test coverage for explainability algorithms; Added test coverage for new and existing graph datasets; Added test coverage for profiling and benchmarking utilities; Added test coverage for the nn.aggr package; Added test coverage for torch\_geometric.llm models and RAG utilities; Added test suite for GraphGym configuration, registration, and Lightning integration; Added tests for GNN convolution utility functions; Added tests for IO filesystem and OFF format utilities; Added tests for MGNAN model and RBCD attack methods; Added tests for NeighborSampler base functionality and output merging; Added tests for PGMExplainer; Added tests for Performer, Polynormer, and QFormer attention modules; Added tests for PyTorch Lightning 2.0 integration in LightningDataModule; Added tests for SelectTopK pooling and topk selection logic; Added tests for TransE, ComplEx, DistMult, and RotatE KGE models; Added tests for bro and gini functional operations; Added tests for explainability evaluation metrics; Added tests for graph generator datasets; Added tests for graph neural network normalization layers; Added tests for graph pooling and KNN index utilities; Added tests for graph visualization and influence functions; Added tests for knn\_interpolate; Added tests for link prediction metrics; Added tests for motif generator classes; Added tests for new and existing GNN models; Added tests for new and existing graph transforms; Added tests for the FilterEdges graph coarsening operator; Added tests for the from\_relbench utility; Added tests for the graph explanation module; Added tests for the new torch\_geometric.llm module; Expanded test coverage for PyG data objects and storage backends; Expanded test coverage for PyTorch 2.x compilation, heterogeneous graph transformers, and model utilities; New centralized testing utilities package; New test infrastructure and coverage for core PyTorch Geometric components.

Dependencies

Migrate build system to pyproject.toml and update documentation dependencies

The project has migrated its package configuration from legacy formats to a modern \pyproject.toml\ file, establishing \flit\_core\ as the build backend and explicitly defining core dependencies (such as \numpy\, \aiohttp\, and \fsspec\) alongside optional dependency groups for features like GraphGym, Hugging Face Hub integration, and RAG. This change also introduces a new \docs/requirements.txt\ to pin specific versions for documentation builds, including a direct URL to a CPU-only PyTorch wheel and the \pyg\_sphinx\_theme\ repository, ensuring consistent and reproducible documentation generation.

(dependencies) · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

Baseline

  • First survey — no prior run to compare against. CAI 70.

Lenses

  • Code Health 84
  • Architecture 96
  • Maturity 69
  • Readiness 72
  • Security 65

Changes since last survey

  • 300 commits — 229 feature/other, 71 fixes

By area

  • (root) — 67 commits
  • .github/workflows — 44 commits
  • torch_geometric/nn — 23 commits
  • examples/llm — 20 commits
  • test/test_hash_tensor.py — 14 commits
  • docs/source — 12 commits
  • test/nn — 11 commits
  • torch_geometric/datasets — 11 commits
  • torch_geometric/metrics — 9 commits
  • examples/multi_gpu — 7 commits
  • .github/actions — 5 commits
  • test/llm — 5 commits
  • torch_geometric/llm — 5 commits
  • examples/README.md — 4 commits
  • test/datasets — 4 commits
  • torch_geometric/utils — 4 commits
  • .github/CODEOWNERS — 3 commits
  • .github/dependabot.yml — 3 commits
  • test/explain — 3 commits
  • test/loader — 3 commits

Notable commits

  • fix: Fix runtime crash due to Half/Float mismatch in GNN+LLM training and inference. (#10595)
  • fix: Fix B006 lint errors: using mutable structure in default argument (#10259)
  • fix: Fix B007 lint errors: unused loop control variable (#10303)
  • fix: Fix CI (#10345)
  • fix: Fix CI Test for PGM Explainer (#10165)
  • fix: Fix CI failing with PyTorch 2.12+ (#10647)
  • fix: Fix MovieLens dataset incompatibility with SentenceTransformers =5.x (#10668)
  • fix: Fix NumPy deprecations and enforce CI failure on warnings (#10283)
  • fix: Fix SGFormer to not attend across batch (#10045)
  • fix: Fix .llm import trigger of .distributed depr warning (#10512)
  • fix: Fix EdgeIndex docstring (#10008)
  • fix: Fix GPSE test times (#10174)
  • fix: Fix HashMap import from pyg-lib (#10035)
  • fix: Fix HashTensor.slice (#10061)
  • fix: Fix MD17 dataset labels (#10282)
  • fix: Fix MoleculeGPTDataset paper link (#10457)
  • fix: Fix captum overriding PyTorch version (#10071)
  • fix: Fix deal_nan() to Preserve Differentiability and Prevent Autograd Errors. (#10336)
  • fix: Fix glem.py edge case (#10492)
  • fix: Fix mypy --install-types in CI (#10763)
  • …and 280 more

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

pyg-team/pytorch_geometric was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 19 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit 79d33965a40b7fa83616a9f598a0f8619f25d939 — the exact code this score is about.
  • Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-5d04157a340d.