Pith. sign in

REVIEW 4 major objections 5 minor 11 cited by

PyG 2.0: Scalable Learning on Real World Graphs

T0 review · 4 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read PyG 2.0 modularizes graph learning so a single stack scales from small benchmarks to billion-node graphs, with compilation and pruning cutting runtime by up to 5x.

desk verdict PyG 2.0 is a genuine system paper backed by open-source code, but the speedup tables are internally inconsistent and some scalability claims outrun the evidence. read the letter →

arxiv 2507.16991 v2 pith:UIQCBVW6 submitted 2025-07-22 cs.LG cs.AI

classification cs.LGcs.AI
keywords graphneuralnetworksPyTorchGeometricscalablelearningheterogeneousgraphstemporalfeaturestoremodelcompilation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that one graph-learning framework can cover the whole spectrum from small benchmark graphs to billion-scale, heterogeneous, temporal real-world data without sacrificing the familiar training loop. Its central design move is to separate the data-loading pipeline into a feature store, a graph store, and a sampler behind two interfaces, so users can swap in-memory storage for external databases or distributed GPU storage by implementing only a few methods. On the compute side, it claims that a metadata-carrying EdgeIndex tensor, first-class aggregation, model compilation, and layer-wise pruning of sampled subgraphs produce 2–3x and 4–5x runtime improvements across common GNN architectures while maintaining accuracy. If these claims hold, the framework makes billion-node graph learning, relational deep learning on databases, and graph-augmented retrieval for language models practical in one toolchain.

What carries the argument

The load-bearing mechanism is the pair of remote-backend interfaces, FeatureStore and GraphStore, together with the EdgeIndex tensor. FeatureStore and GraphStore define where node features and edge indices live and how sampling is performed, so a custom backend needs only implement retrieval methods and the rest of the training loop stays unchanged. EdgeIndex is a tensor subclass that carries metadata such as sort order and undirectedness and caches CSR/CSC conversions on demand, letting message passing pick sparse-matrix multiplication and segmented aggregation when possible and skip redundant transposes on undirected graphs. The third piece is layer-wise progressive pruning: because a BFS-sampled subgraph arranges nodes by hop distance, slicing the adjacency and feature matrices in that order removes computation for nodes that cannot influence seed-node representations in later layers, with zero copying. That pruning step is what combines with compilation to reach the reported 4–5x improvement.

What would settle it

Run the paper's open benchmark suite on a fixed GPU and dataset and measure whether compilation gives 2–3x and compilation-plus-pruning gives 4–5x across the listed architectures; a standard configuration falling outside those ranges would refute the generality of the speedup claim. For the scaling claim, run the GPU-accelerated pipeline on 1, 2, 4, and 8 GPUs on a billion-edge graph and check whether data-loading throughput doubles at each step.

Watch

Extended reading notes

Core claim

The central claim is architectural: PyG 2.0 reorganizes graph learning into three swappable components—graph infrastructure, a neural framework, and post-processing—with data loading split into a feature store, a graph store, and a sampler. This split is what allows the same training loop to run on in-memory graphs and on graphs of roughly ten billion nodes stored in external systems. The paper further claims that message passing can be made compilation-friendly: an EdgeIndex tensor caches sparse formats and records sort order so the right compute path is chosen automatically, aggregations are treated as first-class plug-in operations, and a BFS-ordered multi-hop subgraph can be progressively trimmed layer by layer because later-hop nodes no longer contribute to seed-node representations. Combined, the paper reports 2–3x speedups from compilation alone and 4–5x when pruning is added, with predictive accuracy maintained.

Load-bearing premise

The headline speedup and linear-scaling numbers come from benchmark measurements whose hardware, dataset sizes, and repetitions are not stated, so the claims may hold only under the authors' own test conditions.

Editorial extensions

If this is right

  • A user can implement only the retrieval methods of a custom feature or graph backend and keep the identical training loop, whether data sits in RAM, a database, or a distributed store.
  • Compilation alone yields 2–3x forward-plus-backward speedups across GIN, GraphSAGE, GCN, GAT, and EdgeCNN with no accuracy loss.
  • Adding layer-wise pruning on top of compilation yields 4–5x speedups on the same architectures.
  • Any homogeneous message-passing GNN can be turned into a heterogeneous variant automatically, with layers replicated per edge type.
  • GPU-accelerated sampling and distributed feature storage achieve linear scaling when additional GPUs are stacked, and 2–8x faster data loading even on a single GPU.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • An unstated consequence of the storage-interface split is that the speedups become partly a property of the backend: a database or key-value store that implements FeatureStore well could make end-to-end training faster than the in-memory baseline without touching the model.
  • The benchmark tables report per-step runtimes on relatively small workloads; whether the 2–3x and 4–5x numbers survive end-to-end on billion-scale graphs, where sampling and feature transfer dominate, is a separate question the paper does not fully settle.
  • The reported GraphRAG accuracy jump from 16% to 32% suggests that encoding retrieved subgraphs with a GNN, not just retrieving text, is what drives the gain; ablating the GNN encoder against retrieval-only context would isolate that effect.
  • If the pruning argument is general, similar progressive slicing could apply to any sampling order that admits a topological ordering, not only BFS subgraphs.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. This paper presents PyG 2.0, a major update of the PyTorch Geometric library for graph neural networks. The authors describe the architecture and design principles, including heterogeneous and temporal graph support, FeatureStore and GraphStore storage abstractions, an EdgeIndex tensor with caching, model compilation via torch.compile, progressive pruning for subgraph sampling, explainability interfaces, and integrations such as cuGraph and GraphRAG. The paper claims 2–3× speedups from compilation (Table 1) and 4–5× speedups from compilation plus trimming (Table 2), as well as linear GPU scaling through cuGraph integration. It also reviews application areas including relational deep learning and LLM integration, and reports a 2× accuracy improvement for GraphRAG. The paper is primarily a system/architecture description rather than a conventional empirical study, and its quantitative claims are concentrated in two runtime tables and a third-party accuracy figure.

Significance. PyG is one of the most widely used GNN frameworks, and the architectural changes described here—EdgeIndex, FeatureStore/GraphStore, first-class aggregations, torch.compile support, and heterogeneous/temporal processing—are likely to affect a large user community. If the speedup and scaling claims are accurate, they would be practically valuable. The paper is also useful as a survey of application areas and the surrounding ecosystem, and the open-sourced code and benchmark protocols are a strength. However, the central quantitative evidence is currently too under-specified and partially inconsistent with the claims, so the measured performance numbers cannot yet be considered validated.

major comments (4)
  1. [Sec. 2.2, Table 1] The claim 'On average, we observe 2–3× speedup... cf. Table 1' is contradicted by the table's own numbers. Computing the ratio of eager to compile times gives 3.34× (GIN), 3.39× (GraphSAGE), 1.69× (EdgeCNN), 4.27× (GCN), and 3.57× (GAT), so the stated interval does not represent the data. In addition, the table reports runtime in milliseconds without stating the hardware, dataset, graph size, or number of repetitions; the pointer to a GitHub repository is not a substitute for reporting these details in the paper.
  2. [Sec. 2.3, Table 2] The claim that 'with both compilation and trimming enabled, runtimes can be improved by 4–5×' is not supported for all listed models: the speedup for GCN is 19.73/5.86 ≈ 3.37× and for GAT is 29.72/7.93 ≈ 3.75×, both below the claimed range. The table also does not clearly label the baseline (presumably eager with trimming disabled) and lacks the same experimental details as Table 1. Please either correct the speedup claim to match the data or add measurements that justify the 4–5× range.
  3. [Sec. 2.3, 'cuGraph Integration'] The statement that 'it is possible to achieve linear scaling when stacking additional GPUs' is made without presenting any scaling measurements, a reference to a study that measured it, or a description of the conditions under which linear scaling holds. Because this is one of the core scalability claims in the paper, it needs either an experiment (e.g., throughput vs. number of GPUs) or a citation to a published evaluation.
  4. [Sec. 3.2, 'Integration in Large Language Models'] The quantitative claim that GNN+LLM GraphRAG improves accuracy from 16% to 32% (a 2× increase) is attributed to an external NVIDIA blog [78] and is not measured in this work. Since this is presented as a concrete benefit of PyG's GraphRAG pipeline, the authors should either reproduce the experiment in the paper, clearly mark the number as a third-party result with the full experimental setup, or remove the quantitative comparison.
minor comments (5)
  1. [Page 1 footer] 'Published as a worshopt paper at KDD 2025' should read 'workshop paper.'
  2. [Table 2] The column header is garbled: 'Run Trim GIN Graph Edge GCN GATMode SAGE CNN' is not legible and should be reformatted to clearly identify the run mode and trimming columns.
  3. [Sec. 3.3] The heading 'Others Application Areas' should be 'Other Application Areas.'
  4. [Sec. 2.2 and Sec. 2.3] The paper states a '2–3× speedup' for compilation in Sec. 2.2 and a '4–5× speedup' for compilation plus trimming in Sec. 2.3, but it does not explain whether the latter is cumulative or separate; please clarify the relationship.
  5. [Sec. 2.3] The FeatureStore and GraphStore API is described only in words; a short code example of the required interface methods would help reproducibility and clarity.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: PyG 2.0 is an engineering/software description, not a derivation; its claims rest on open-sourced benchmarks and external references rather than on self-referential fits.

full rationale

This paper is a systems/software report on PyG 2.0. There is no derivation chain in which an output quantity is defined in terms of, or fitted to, the quantity it is claimed to predict. The speedup claims (Tables 1 and 2) are benchmark measurements, not identities; even though the benchmark protocol is only referenced by URL rather than fully specified, that is an evidence/reproducibility concern, not circularity. The paper does cite prior work by overlapping authors (PyG [28], PyTorch Frame [41], Relational Deep Learning [27], RelBench [72], Temporal Graph Benchmark [43]), but these citations are used to establish the existence and design of components and prior benchmarks, not to justify the new quantitative claims, and the artifacts are open-source and externally checkable. The GraphRAG accuracy improvement (16% to 32%) is attributed to an external NVIDIA blog post [78], not derived in this paper, and therefore cannot be a circular reduction. No equation in the paper defines a fitted parameter that is then renamed as a prediction, and no uniqueness theorem from the authors' prior work is invoked to force a modeling choice. The assertion of "linear scaling when stacking additional GPUs" lacks supporting data, but unsupported assertion is a correctness/evidence problem, not a demonstration that the claim reduces to its own inputs. Under the review rules, no specific circular step can be quoted and exhibited, so the honest finding is no significant circularity (0).

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

The framework depends on standard software and hardware assumptions (PyTorch, torch.compile, GPUs) and on the representativeness of its benchmark measurements. No free parameters or invented entities appear.

assumptions (2)
  • domain assumption Benchmark measurements in Tables 1 and 2 are representative of real-world graph workloads.
    The paper reports 2-3x and 4-5x speedups without specifying hardware, datasets, graph sizes, or repetitions, yet these numbers are the primary quantitative support for the scalability claims.
  • domain assumption torch.compile fuses PyG message passing operations without changing predictive accuracy.
    Table 1 caption says accuracy is maintained, but no accuracy numbers are given; the speedup claims depend on this assumption.

how reviews work

0 comments
Cite this review

Pith. "Pith review of PyG 2.0: Scalable Learning on Real World Graphs." pith.science (2026). https://pith.science/paper/UIQCBVW6

@misc{pith2026250716991,
  author       = {Pith},
  title        = {Pith review of: PyG 2.0: Scalable Learning on Real World Graphs},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/UIQCBVW6}},
  note         = {Machine review of arXiv:2507.16991}
}
read the original abstract

PyG (PyTorch Geometric) has evolved significantly since its initial release, establishing itself as a leading framework for Graph Neural Networks. In this paper, we present Pyg 2.0 (and its subsequent minor versions), a comprehensive update that introduces substantial improvements in scalability and real-world application capabilities. We detail the framework's enhanced architecture, including support for heterogeneous and temporal graphs, scalable feature/graph stores, and various optimizations, enabling researchers and practitioners to tackle large-scale graph learning problems efficiently. Over the recent years, PyG has been supporting graph learning in a large variety of application areas, which we will summarize, while providing a deep dive into the important areas of relational deep learning and large language modeling.

Figures

Figures reproduced from arXiv: 2507.16991 by the authors.

Figure 1
Figure 1. Architectural overview of PyG 2.0: The system’s modular design allows to swap out any component without affecting other parts of the pipeline. For example, one can seamlessly change the FeatureStore from in-memory to distributed key-value storage without modifying the DataLoader or model parts. This plug-and-play approach extends throughout the framework— from storage implementations to sampling strategies and expla… view at source ↗
Figure 2
Figure 2. Exemplary illustration of a GNN explainer in PyG: [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 4
Figure 4. The general GraphRAG pipeline in PyG: A natu [PITH_FULL_IMAGE:figures/full_fig_p006_4.png] view at source ↗
Figures from the paper (1 more)
Figure 3
Figure 3. Figure 3: End-to-end Relational Deep Learning on multi [PITH_FULL_IMAGE:figures/full_fig_p006_3.png]

Discussion (0). Sign in to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Fully Inductive Cardinality Estimation

    cs.DB 2026-07 conditional novelty 7.0 of 10

    An encoder-decoder GNN over a factor-graph view of RDF KGs estimates BGP cardinalities on entirely unseen graphs without retraining, cutting median q-error roughly in half versus the best baseline.

  2. Graph Neural Networks Are Not Continuous Across Graph Resolutions

    cs.LG 2026-05 unverdicted novelty 6.0 of 10

    GNNs are shown to lack continuity under graph resolution changes due to message-passing schemes, with a derived modification enabling consistent multi-scale representations validated experimentally.

  3. Predictive Query Language: A Domain-Specific Language for Predictive Modeling on Relational Databases

    cs.DB 2026-02 conditional novelty 6.0 of 10

    A SQL-like domain-specific language that declares predictive tasks on relational databases and automatically generates leakage-free training labels, with batch and low-latency implementations.

  4. RankGraph-2: Lifecycle Co-Design for Billion-Node Graph Learning in Recommendation

    cs.IR 2026-06 conditional novelty 5.5 of 10

    Lifecycle co-design of construction, training, and serving lets a simple heterogeneous GNN beat stronger models on Meta-scale similarity retrieval and cut serving cost 83%.

  5. RankGraph-2: Lifecycle Co-Design for Billion-Node Graph Learning in Recommendation

    cs.IR 2026-06 conditional novelty 5.0 of 10

    Jointly co-designing construction, training, and serving for billion-node graph retrieval gives RankGraph-2 large quality gains and 83% lower serving cost.

  6. On Efficient Scaling of GNNs via IO-Aware Layers Implementations

    cs.LG 2026-05 unverdicted novelty 5.0 of 10

    IO-aware GPU kernels for SpMM convolutions, degree-aware reductions, and fused attention layers deliver median speedups of 1.6-2.6x (up to 10x) and memory reductions up to 76x over DGL/PyG baselines on realistic graphs.

  7. Unified Multi-Domain Graph Pre-training for Homogeneous and Heterogeneous Graphs via Domain-Specific Expert Encoding

    cs.LG 2026-02 conditional novelty 5.0 of 10

    GPH^2 pre-trains one expert per graph on edge-dropped or meta-path views and fuses frozen experts with class-wise attention, outperforming type-specific graph pre-training baselines.

  8. Latent Dynamics Graph Convolutional Networks for model order reduction of parameterized time-dependent PDEs

    cs.LG 2026-01 conditional novelty 5.0 of 10

    LD-GCN couples an encoder-free latent-space neural ODE with a graph convolutional decoder, achieving accurate reduced-order modeling of time-dependent parameterized PDEs and detecting bifurcations from the latent traj...

  9. CrediBench: Building Web-Scale Network Datasets for Information Integrity

    cs.SI 2025-09 reject novelty 5.0 of 10

    CrediBench presents a one-month, 1-billion-edge Common Crawl web graph with text and 11.5K expert credibility labels, while the abstract's promised 8-month dataset and 85%-accuracy classifier are absent from the paper.

  10. RankGraph-2: Lifecycle Co-Design for Billion-Node Graph Learning in Recommendation

    cs.IR 2026-06 unverdicted novelty 4.0 of 10

    RankGraph-2 jointly optimizes graph subsampling, pre-computed neighborhoods, and a co-learned cluster index for billion-node recommendation retrieval, reporting 3.8x recall gains and up to +0.96% CTR.

  11. RankGraph-2: Lifecycle Co-Design for Billion-Node Graph Learning in Recommendation

    cs.IR 2026-06 unverdicted novelty 4.0 of 10

    RankGraph-2 jointly optimizes graph construction, training, and serving for billion-node recommendation retrieval, reporting 3.8x recall gains and CTR/CVR improvements via subsampling, pre-computed neighborhoods, and ...

Reference graph

Works this paper leans on

111 extracted references · 61 canonical work pages · cited by 8 Pith papers

  1. [78]

    Brian Shi, Alfred Clemedtson, Zach Blumenfeld, and Rishi Puri. 2025. Boosting Q&A Accuracy with GraphRAG Using PyG and Graph Databases . https://developer.nvidia.com/blog/boosting-qa-accuracy-with-graphrag- using-pyg-and-graph-databases

  2. [1]

    Large-scale graph representation learning with very deep GNNs and self-supervision

    R. Addanki, P. W. Battaglia, D. Budden, A. Deac, J. Godwin, T. Keck, W. L. Sibon Li, A. Sanchez-Gonzalez, J. Stott, S. Thakoor, and P. Veličković. 2021. Large-scale Graph Representation Learning with Very Deep GNNs and Self- supervision. CoRR abs/2107.09422 (2021)

  3. [2]

    Agarwal, O

    C. Agarwal, O. Queen, H. Lakkaraju, and M. Zitnik. 2023. Evaluating Explain- ability for Graph Neural Networks. CoRR abs/2208.09339 (2023)

  4. [3]

    Amara, Z

    K. Amara, Z. Ying, Z. Zhang, Z. Han, Y. Zhao, Y. Shan, U. Brandes, S. Schemm, and C. Zhang. 2022. GraphFramEx: Towards Systematic Evaluation of Explain- ability Methods for Graph Neural Networks. In LOG

  5. [4]

    S. Ö. Arik and T. Pfister. 2021. TabNet: Attentive interpretable tabular learning. In AAAI

  6. [5]

    K. Atz, L. Cotos, C. Isert, M. Håkansson, D. Focht, M. Hilleke, D. F. Nippa, M. Iff, J. Ledergerber, C. C. G. Schiebroek, V. Romeo, J. A. Hiss, D. Merk, P. Schneider, B. Kuhn, U. Grether, and G. Schneider. 2024. Prospective De Novo Drug Design with Deep Interactome Learning. Nature Communications (2024)

  7. [6]

    Fuchs, and Timothy Lillicrap

    Sergey Bartunov, Fabian B. Fuchs, and Timothy Lillicrap. 2022. Equilibrium Aggregation: Encoding Sets via Optimization. In UAI

  8. [7]

    Peter Battaglia, Jessica Blake Chandler Hamrick, Victor Bapst, Alvaro Sanchez, Vinicius Zambaldi, Mateusz Malinowski, Andrea Tacchetti, David Raposo, Adam Santoro, Ryan Faulkner, Caglar Gulcehre, Francis Song, Andy Ballard, Justin Gilmer, George E. Dahl, Ashish Vaswani, Kelsey Allen, Charles Nash, Victo- ria Jayne Langston, Chris Dyer, Nicolas Heess, Daan...

Show all 111 references
  1. [8]

    Bechler-Speicher, B

    M. Bechler-Speicher, B. Finkelshtein, F. Frasca, L. Müller, J. Tönshoff, A. Siraudin, V. Zaverkin, M. M. Bronstein, M. Niepert, B. Perozzi, M. Galkin, and C. Morris

  2. [9]

    Brahmbhatt, C

    S. Brahmbhatt, C. Tang, C. D. Twigg, C. C. Kemp, and J. Hays. 2020. ContactPose: A Dataset of Grasps with Object Contact and Hand Pose. In ECCV

  3. [10]

    Bronstein, J

    M. Bronstein, J. Bruna, T. Cohen, and P. Veličković. 2021. Geometric Deep Learning: Grids, Groups, Graphs, Geodesics, and Gauges. CoRR abs/2104.13478 (2021)

  4. [11]

    Cappart, D

    Q. Cappart, D. Chételat, E. B. Khalil, A. Lodi, C. Morris, and P. Veličković. 2023. Combinatorial Optimization and Reasoning with Graph Neural Networks.JMLR (2023)

  5. [12]

    J. Chen, T. Ma, and C. Xiao. 2018. FastGCN: Fast Learning with Graph Convo- lutional Networks via Importance Sampling. In ICLR

  6. [13]

    J. Chen, J. Yan, D. Z. Chen, and J. Wu. 2023. ExcelFormer: A Neural Network Surpassing GBDTs on Tabular Data. CoRR abs/2301.02819 (2023)

  7. [14]

    J. Chen, J. Zhu, and L. Song. 2018. Stochastic Training of Graph Convolutional Networks with Variance Reduction. In ICML

  8. [15]

    K. Y. Chen, P. H. Chiang, H. R. Chou, T. W. Chen, and T. H. Chang. 2023. Trompt: Towards a Better Deep Neural Network for Tabular Data. In ICML

  9. [16]

    Z. Chen, H. Mao, J. Liu, Y. Song, B. Li, W. Jin, B. Fatemi, A. Tsitsulin, B. Perozzi, H. Liu, and J. Tang. 2024. Text-space Graph Foundation Models: Comprehensive Benchmarks and New Insights. CoRR 2406.10727 (2024)

  10. [17]

    W. L. Chiang, X. Liu, S. Si, Y. Li, S. Bengio, and C. J. Hsieh. 2019. Cluster-GCN: An Efficient Algorithm for Training Deep and Large Graph Convolutional Networks. In KDD

  11. [18]

    Corso, L

    G. Corso, L. Cavalleri, D. Beaini, P. Liò, and P. Veličković. 2020. Principal Neighbourhood Aggregation for Graph Nets. In NeurIPS

  12. [19]

    A. R. Costa and C. G. Ralha. 2023. AC2CD: An actor–critic architecture for community detection in dynamic social networks. Knowledge-Based Systems (2023)

  13. [20]

    Defferrard, X

    M. Defferrard, X. Bresson, and P. Vandergheynst. 2016. Convolutional Neural Networks on Graphs with Fast Localized Spectral Filtering. In NIPS

  14. [21]

    C. Deng, Z. Yue, and Z. Zhang. 2024. Polynormer: Polynomial-Expressive Graph Transformer in Linear Time. In ICLR

  15. [22]

    L. Deng. 2012. The MNIST Database of Handwritten Digit Images for Machine Learning Research. IEEE Signal Processing Magazine (2012)

  16. [23]

    N. S. Detlefsen, J. Borovec, J. Schock, A. Harsh, T. Koker, L. D. Liello, D. Stancl, C. Quan, M. Grechkin, and W. Falcon. 2022. TorchMetrics: Machine Learning Metrics for PyTorch

  17. [24]

    Matthijs Douze, Alexandr Guzhva, Chengqi Deng, Jeff Johnson, Gergely Szilvasy, Pierre-Emmanuel Mazaré, Maria Lomeli, Lucas Hosseini, and Hervé Jégou. 2024. The FAISS Library. CoRR abs/2401.08281 (2024)

  18. [25]

    Darren Edge, Ha Trinh, Newman Cheng, Joshua Bradley, Alex Chao, Apurva Mody, Steven Truitt, Dasha Metropolitansky, Robert Osazuwa Ness, and Jonathan Larson. 2024. From Local to Global: A Graph RAG Approach to Query-focused Summarization. CoRR (2024)

  19. [26]

    Shangbin Feng, Zhaoxuan Tan, Herun Wan, Ningnan Wang, Zilong Chen, Binchi Zhang, Qinghua Zheng, Wenqian Zhang, Zhenyu Lei, Shujie Yang, Xinshun Feng, Qingyue Zhang, Hongrui Wang, Yuhan Liu, Yuyang Bai, Heng Wang, Zijian Cai, Yanbo Wang, Lijing Zheng, Zihan Ma, Jundong Li, and ...

  20. [27]

    Matthias Fey, Weihua Hu, Kexin Huang, Jan Eric Lenssen, Rishabh Ranjan, Joshua Robinson, Rex Ying, Jiaxuan You, and Jure Leskovec. 2024. Position: Re- lational Deep Learning: Graph Representation Learning on Relational Databases. In ICML

  21. [28]

    Matthias Fey and Jan Eric Lenssen. 2019. Fast graph representation learning with PyTorch Geometric. ICLR (RLGM Workshop) (2019)

  22. [29]

    M. Fey, J. E. Lenssen, C. Morris, J. Masci, and N. M. Kriege. 2020. Deep Graph Matching Consensus. In ICLR

  23. [30]

    M. Fey, J. E. Lenssen, F. Weichert, and J. Leskovec. 2021. GNNAutoScale: Scalable and Expressive Graph Neural Networks via Historical Embeddings. In ICML

  24. [31]

    Matthias Fey, Jan Eric Lenssen, Frank Weichert, and Heinrich Müller. 2018. SplineCNN: Fast Geometric Deep Learning With Continuous B-Spline Kernels. In CVPR

  25. [32]

    Victor Fung, Jiaxin Zhang, Eric Juarez, and Bobby G. Sumpter. 2021. Bench- marking Graph Neural Networks for Materials Chemistry. npj Computational Materials (2021). 7 Fey et al

  26. [33]

    Ziqi Gao, Chenran Jiang, Jiawen Zhang, Xiaosen Jiang, Lanqing Li, Peilin Zhao, Huanming Yang, Yong Huang, and Jia Li. 2023. Hierarchical Graph Learning for Protein–Protein Interaction. Nature Communications (2023)

  27. [34]

    Schoenholz, Patrick F

    Justin Gilmer, Samuel S. Schoenholz, Patrick F. Riley, Oriol Vinyals, and George E. Dahl. 2017. Neural Message Passing for Quantum Chemistry. In ICML

  28. [35]

    Yury Gorishniy, Ivan Rubachev, Valentin Khrulkov, and Artem Babenko. 2021. Revisiting Deep Learning Models for Tabular Data. In NeurIPS

  29. [36]

    Chaoyu Guan, Ziwei Zhang, Haoyang Li, Heng Chang, Zeyang Zhang, Yijian Qin, Jiyan Jiang, Xin Wang, and Wenwu Zhu. 2021. AutoGL: A Library for Automated Graph Learning. In ICLR (GTRL Workshop)

  30. [37]

    Will Hamilton, Zhitao Ying, and Jure Leskovec. 2017. Inductive Representation Learning on Large Graphs. In NeurIPS

  31. [38]

    Chaoyang He, Songze Li, Jinhyun So, Xiao Zeng, Mi Zhang, Hongyi Wang, Xi- aoyang Wang, Praneeth Vepakomma, Abhishek Singh, Hang Qiu, Xinghua Zhu, Jianzong Wang, Li Shen, Peilin Zhao, Yan Kang, Yang Liu, Ramesh Raskar, Qiang Yang, Murali Annavaram, and Salman Avestimehr. 2020. ...

  32. [39]

    X. He, X. Bresson, T. Laurent, A. Perold, Y. LeCun, and B. Hooi. 2024. Harnessing Explanations: LLM-to-LM Interpreter for Enhanced Text-Attributed Graph Representation Learning. In ICLR

  33. [40]

    Chawla, Thomas Laurent, Yann Le- Cun, Xavier Bresson, and Bryan Hooi

    Xiaoxin He, Yijun Tian, Yifei Sun, Nitesh V. Chawla, Thomas Laurent, Yann Le- Cun, Xavier Bresson, and Bryan Hooi. 2024. G-Retriever: Retrieval-Augmented Generation for Textual Graph Understanding and Question Answering. CoRR abs/2402.07630 (2024)

  34. [41]

    Weihua Hu, Yiwen Yuan, Zecheng Zhang, Akihiro Nitta, Kaidi Cao, Vid Kocijan, Jinu Sunil, Jure Leskovec, and Matthias Fey. 2024. PyTorch Frame: A Modular Framework for Multi-Modal Tabular Learning. CoRR abs/2404.00776 (2024)

  35. [42]

    Ziniu Hu, Yuxiao Dong, Kuansan Wang, and Yizhou Sun. 2020. Heterogeneous Graph Transformer. In WWW

  36. [43]

    Shenyang Huang, Farimah Poursafaei, Jacob Danovitch, Matthias Fey, Weihua Hu, Emanuele Rossi, Jure Leskovec, Michael Bronstein, Guillaume Rabusseau, and Reihaneh Rabbany. 2023. Temporal Graph Benchmark for Machine Learning on Temporal Graphs. In NeurIPS

  37. [44]

    Huang, T

    W. Huang, T. Zhang, Y. Rong, and J. Huang. 2018. Adaptive Sampling Towards Fast Graph Representation Learning. In NeurIPS

  38. [45]

    Xin Huang, Ashish Khetan, Milan Cvitkovic, and Zohar Karnin. 2020. Tab- Transformer: Tabular Data Modeling using Contextual Embeddings. CoRR abs/2012.06678 (2020)

  39. [46]

    Maria Huegle, Gabriel Kalweit, Moritz Werling, and Joschka Boedecker. 2020. Dynamic Interaction-Aware Scene Understanding for Reinforcement Learning in Autonomous Driving. In ICRA

  40. [47]

    Nikolaos Karalias and Andreas Loukas. 2020. Erdos Goes Neural: an Unsu- pervised Learning Framework for Combinatorial Optimization on Graphs. In NeurIPS

  41. [48]

    Jinwoo Kim, Tien Dat Nguyen, Seonwoo Min, Sungjun Cho, Moontae Lee, Honglak Lee, and Seunghoon Hong. 2022. Pure Transformers are Powerful Graph Learners. In NeurIPS

  42. [49]

    T. N. Kipf and M. Welling. 2017. Semi-supervised Classification with Graph Convolutional Networks. In ICLR

  43. [50]

    Marvin Klimke, Benjamin Völz, and Michael Buchholz. 2022. Cooperative Behavior Planning for Automated Driving Using Graph Neural Networks. In IV

  44. [51]

    Narine Kokhlikyan, Vivek Miglani, Miguel Martin, Edward Wang, Bilal Alsal- lakh, Jonathan Reynolds, Alexander Melnikov, Natalia Kliushkina, Carlos Araya, Siqi Yan, and Orion Reblitz-Richardson. 2020. Captum: A Unified and Generic Model Interpretability Library for PyTorch. CoR...

  45. [52]

    Simon Lang, Mihai Alexe, Matthew Chantry, Jesper Dramsch, Florian Pinault, Baudouin Raoult, Mariana Clare, Christian Lessig, Michael Maier-Gerber, Linus Magnusson, Zied Ben Bouallègue, Ana Nemesio, Peter Düben, Andrew Brown, Florian Pappenberger, and Florence Rabier. 2024. AIF...

  46. [53]

    Jan Eric Lenssen, Christian Osendorfer, and Jonathan Masci. 2020. Deep Iterative Surface Normal Estimation. In CVPR

  47. [54]

    Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2020. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. CoRR abs/2005.1...

  48. [55]

    Guohao Li, Chenxin Xiong, Ali Thabet, and Bernard Ghanem. 2020. DeeperGCN: All You Need to Train Deeper GCNs. CoRR abs/2006.07739 (2020)

  49. [56]

    Kay Liu, Yingtong Dou, Xueying Ding, Xiyang Hu, Ruitong Zhang, Hao Peng, Lichao Sun, and Philip S. Yu. 2024. PyGOD: A Python Library for Graph Outlier Detection. JMLR (2024)

  50. [57]

    Meng Liu, Youzhi Luo, Limei Wang, Yaochen Xie, Hao Yuan, Shurui Gui, Haiyang Yu, Zhao Xu, Jingtun Zhang, Yi Liu, Keqiang Yan, Haoran Liu, Cong Fu, Bora M Oztekin, Xuan Zhang, and Shuiwang Ji. 2021. DIG: A Turnkey Library for Diving into Graph Deep Learning Research. JMLR (2021)

  51. [58]

    Dongsheng Luo, Wei Cheng, Dongkuan Xu, Wenchao Yu, Bo Zong, Haifeng Chen, and Xiang Zhang. 2020. Parameterized Explainer for Graph Neural Network. CoRR abs/2011.04573 (2020)

  52. [59]

    Markowitz, K

    E. Markowitz, K. Balasubramanian, M. Mirtaheri, S. Abu-El-Haija, B. Perozzi, G. Ver Steeg, and A. Galstyan. 2021. Graph Traversal with Tensor Functionals: A Meta-Algorithm for Scalable Learning. In ICLR

  53. [60]

    Marco Maurizi, Chao Gao, and Filippo Berto. 2022. Predicting Stress, Strain and Deformation Fields in Materials and Structures with Graph Neural Networks. Scientific Reports (2022)

  54. [61]

    Dimitrios Michail, Nikos Kanakaris, and Iraklis Varlamis. 2022. Detection of Fake News Campaigns using Graph Convolutional Networks.IJIM Data Insights (2022)

  55. [62]

    Xiaoyu Mo, Yang Xing, and Chen Lv. 2021. Heterogeneous Edge-Enhanced Graph Attention Network For Multi-Agent Trajectory Prediction. CoRR abs/2106.07161 (2021)

  56. [63]

    NVIDIA Corporation. 2025. RAPIDS cuGraph. https://docs.rapids.ai/api/ cugraph

  57. [64]

    NVIDIA Corporation. 2025. RAPIDS: GPU Accelerated Data Science. https: //rapids.ai

  58. [65]

    Joel Oskarsson, Tomas Landelius, Marc Peter Deisenroth, and Fredrik Lind- sten. 2024. Probabilistic Weather Forecasting with Hierarchical Graph Neural Networks. In NeurIPS

  59. [66]

    Joel Oskarsson, Tomas Landelius, and Fredrik Lindsten. 2023. Graph-based Neural Weather Prediction for Limited Area Modeling. In NeurIPS (Workshop on Tackling Climate Change with Machine Learning)

  60. [67]

    Stern, and Artem Cherkasov

    Mohit Pandey, Michael Fernandez, Francesco Gentile, Olexandr Isayev, Alexan- der Tropsha, Abraham C. Stern, and Artem Cherkasov. 2022. The Transforma- tional Role of GPU Computing and Deep Learning in Drug Discovery. Nature Machine Intelligence (2022)

  61. [68]

    Paszke, S

    A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. Kopf, E. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala

  62. [69]

    Ruizhong Qiu, Zhiqing Sun, and Yiming Yang. 2022. DIMES: A Differentiable Meta Solver for Combinatorial Optimization Problems. In NeurIPS

  63. [70]

    Rampášek, M

    L. Rampášek, M. Galkin, V. Prakash, A. T. Luu, G. Wold, and D. Beaini. 2022. Recipe for a General, Powerful, Scalable Graph Transformer. In NeurIPS

  64. [71]

    Reed, Zachary DeVito, Horace He, Ansley Ussery, and Jason Ansel

    James K. Reed, Zachary DeVito, Horace He, Ansley Ussery, and Jason Ansel

  65. [72]

    Lenssen, Yiwen Yuan, Zecheng Zhang, Xinwei He, and Jure Leskovec

    Joshua Robinson, Rishabh Ranjan, Weihua Hu, Kexin Huang, Jiaqi Han, Alejan- dro Dobles, Matthias Fey, Jan E. Lenssen, Yiwen Yuan, Zecheng Zhang, Xinwei He, and Jure Leskovec. 2024. RelBench: A Benchmark for Deep Learning on Relational Databases

  66. [73]

    Benedek Rozemberczki, Paul Scherer, Yixuan He, George Panagopoulos, Alexan- der Riedel, Maria Astefanoaei, Oliver Kiss, Ferenc Beres, Guzman Lopez, Nicolas Collignon, and Rik Sarkar. 2021. PyTorch Geometric Temporal: Spatiotemporal Signal Processing with Neural Machine Learnin...

  67. [74]

    Kipf, Peter Bloem, Rianne van den Berg, Ivan Titov, and Max Welling

    Michael Schlichtkrull, Thomas N. Kipf, Peter Bloem, Rianne van den Berg, Ivan Titov, and Max Welling. 2018. Modeling Relational Data with Graph Convolutional Networks. In The Semantic Web

  68. [75]

    Michael Sejr Schlichtkrull, Nicola De Cao, and Ivan Titov. 2021. Interpreting Graph Neural Networks for NLP with Differentiable Edge Masking. In ICLR

  69. [76]

    Martin J. A. Schuetz, John Kyle Brubaker, and Helmut G. Katzgraber. 2021. Combinatorial Optimization with Physics-inspired Graph Neural Networks. Nature Machine Intelligence (2021)

  70. [77]

    Prithviraj Sen, Galileo Namata, Mustafa Bilgic, Lise Getoor, Brian Gallagher, and Tina Eliassi-Rad. 2008. Collective Classification in Network Data. AI Mag. (2008)

  71. [79]

    Hamed Shirzad, Ameya Velingker, Balaji Venkatachalam, Danica J Sutherland, and Ali Kemal Sinop. 2023. Exphormer: Sparse Transformers for Graphs. In ICML

  72. [80]

    Simonyan, A

    K. Simonyan, A. Vedaldi, and A. Zisserman. 2013. Deep Inside Convolutional Networks: Visualising Image Classification Models and Saliency Maps. CoRR 1312.6034 (2013)

  73. [81]

    Springenberg, A

    J.T. Springenberg, A. Dosovitskiy, T. Brox, and M. Riedmiller. 2015. Striving for Simplicity: The All Convolutional Net. In ICLR (Workshop Track)

  74. [82]

    Sundararajan, A

    M. Sundararajan, A. Taly, and Q. Yan. 2024. Axiomatic Attribution for Deep Networks. In ICML

  75. [83]

    Tailor, Felix L

    Shyam A. Tailor, Felix L. Opolka, Pietro Liò, and Nicholas D. Lane. 2022. Do We Need Anisotropic Graph Neural Networks?. In ICLR

  76. [84]

    Zeyuan Tan, Xiulong Yuan, Congjie He, Man-Kit Sit, Guo Li, Xiaoze Liu, Baole Ai, Kai Zeng, Peter Pietzuch, and Luo Mai. 2023. Quiver: Supporting GPUs 8 PyG 2 . 0: Scalable Learning on Real World Graphs for Low-Latency, High-Throughput GNN Serving with Workload Awareness. CoRR ...

  77. [85]

    Vijay Thakkar, Pradeep Ramani, Cris Cecka, Aniket Shivam, Honghao Lu, Ethan Yan, Jack Kosaian, Mark Hoemmen, Haicheng Wu, Andrew Kerr, Matt Nicely, Duane Merrill, Dustyn Blasig, Fengqi Qiao, Piotr Majcher, Paul Springer, Markus Hohnerbach, Jin Wang, and Manish Gupta. 2023. CUT...

  78. [86]

    Veličković, G

    P. Veličković, G. Cucurull, A. Casanova, A. Romero, P. Liò, and Y. Bengio. 2018. Graph Attention Networks. In ICLR

  79. [87]

    M. Wang, D. Zheng, Z. Ye, Q. Gan, M. Li, X. Song, J. Zhou, C. Ma, L. Yu, Y. Gai, T. Xiao, T. He, G. Karypis, J. Li, and Z. Zhang. 2019. Deep Graph Library: A Graph-Centric, Highly-Performant Package for Graph Neural Networks. CoRR abs/1909.01315 (2019)

  80. [88]

    S. Wang, J. Huang, Z. Chen, Y. Song, W. Tang, H. Mao, W. Fan, H. Liu, X. Liu, D. Yin, and Q. Li. 2025. Graph Machine Learning in the Era of Large Language Models (LLMs). TIST (2025)

  81. [89]

    Xiao Wang, Houye Ji, Chuan Shi, Bai Wang, Yanfang Ye, Peng Cui, and Philip S Yu. 2019. Heterogeneous Graph Attention Network. In WWW

  82. [90]

    Yiwei Wang, Yujun Cai, Yuxuan Liang, Henghui Ding, Changhu Wang, and Bryan Hooi. 2021. Time-Aware Neighbor Sampling for Temporal Graph Net- works. CoRR abs/2112.09845 (2021)

  83. [91]

    Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement De- langue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Ma...

  84. [92]

    Qitian Wu, Wentao Zhao, Chenxiao Yang, Hengrui Zhang, Fan Nie, Haitian Jiang, Yatao Bian, and Junchi Yan. 2024. SGFormer: Simplifying and Empowering Transformers for Large-Graph Representations. CoRR abs/2306.10759 (2024)

  85. [93]

    Keyulu Xu, Weihua Hu, Jure Leskovec, and Stefanie Jegelka. 2019. How Powerful are Graph Neural Networks?. In ICLR

  86. [94]

    Carl Yang, Aydın Buluç, and John D Owens. 2018. Design Principles for Sparse Matrix Multiplication on the GPU. In Euro-PAR

  87. [95]

    Dongxu Yang. 2024. Optimizing Memory and Retrieval for Graph Neural Net- works with WholeGraph (Part 1). https://developer.nvidia.com/blog/optimizing- memory-and-retrieval-for-graph-neural-networks-with-wholegraph-part-1

  88. [96]

    Dongxu Yang. 2024. Optimizing Memory and Retrieval for Graph Neural Net- works with WholeGraph (Part 2). https://developer.nvidia.com/blog/optimizing- memory-and-retrieval-for-graph-neural-networks-with-wholegraph-part-2

  89. [97]

    Yingguang Yang, Renyu Yang, Yangyang Li, Kai Cui, Zhiqin Yang, Yue Wang, Jie Xu, and Haiyong Xie. 2023. RoSGAS: Adaptive Social Bot Detection with Reinforced Self-supervised GNN Architecture Search. ACM Trans. Web (2023)

  90. [98]

    Rex Ying, Dylan Bourgeois, Jiaxuan You, Marinka Zitnik, and Jure Leskovec

  91. [99]

    J. You, R. Ying, and J. Leskovec. 2020. Design Space for Graph Neural Networks. In NeurIPS

  92. [100]

    Khar- gonekar, and Mohammad Abdullah Al Faruque

    Shih-Yuan Yu, Arnav Vaibhav Malawade, Deepan Muthirayan, Pramod P. Khar- gonekar, and Mohammad Abdullah Al Faruque. 2022. Scene-Graph Augmented Data-Driven Risk Assessment of Autonomous Vehicle Decisions. IEEE Transac- tions on Intelligent Transportation Systems (2022)

  93. [101]

    Zeiler and Rob Fergus

    Matthew D. Zeiler and Rob Fergus. 2014. Visualizing and Understanding Con- volutional Networks. In ECCV

  94. [102]

    Hanqing Zeng, Muhan Zhang, Yinglong Xia, Ajitesh Srivastava, Andrey Male- vich, Rajgopal Kannan, Viktor Prasanna, Long Jin, and Ren Chen. 2025. Decou- pling the Depth and Scope of Graph Neural Networks. In NeurIPS

  95. [103]

    In NeurIPS

    GNNExplainer: Generating Explanations for Graph Neural Networks. In NeurIPS

  96. [104]

    Chaoyi Zhang, Jianhui Yu, Yang Song, and Weidong Cai. 2021. Exploiting Edge-Oriented Reasoning for 3D Point-Based Scene Graph Analysis. In CVPR

  97. [105]

    Yang Zhang, Zixiang Zhou, Philip David, Xiangyu Yue, Zerong Xi, Boqing Gong, and Hassan Foroosh. 2020. PolarNet: An Improved Grid Representation for Online LiDAR Point Clouds Semantic Segmentation. In CVPR

  98. [106]

    D. Zou, Z. Hu, Y. Wang, S. Jiang, Y. Sun, and Q. Gu. 2019. Layer-Dependent Im- portance Sampling for Training Deep and Large Graph Convolutional Networks. In NeurIPS. 9

  99. [108]

    H. Zeng, H. Zhou, A. Srivastava, R. Kannan, and V. Prasanna. 2020. GraphSAINT: Graph Sampling Based Inductive Learning Method. In ICLR

  100. [2019]

    In NeurIPS

    PyTorch: An Imperative Style, High-Performance Deep Learning Library. In NeurIPS

  101. [2021]

    CoRR abs/2112.08429 (2021)

    torch.fx: Practical Program Capture and Transformation for Deep Learn- ing in Python. CoRR abs/2112.08429 (2021)

  102. [2022]

    In NeurIPS

    TwiBot-22: Towards Graph-Based Twitter Bot Detection. In NeurIPS

  103. [2025]

    Position: Graph Learning Will Lose Relevance Due To Poor Benchmarks. In ICML

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.