REVIEW 4 major objections 6 minor 72 references
Scaling Large-scale GNN Training to Thousands of Processors on CPU-based Supercomputers
T0 review · 4 major / 6 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read The paper claims that finding the optimal communication schedule for a distributed GCN layer reduces exactly to solving minimum vertex cover on a bipartite graph, yielding the smallest possible boundary-node traffic and a measured 1.5x…
desk verdict Solid systems paper with a real algorithmic contribution, but it overclaims generality to attention GNNs and its accuracy abstract outruns the data. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the minimum vertex cover of a bipartite graph: the smallest set of vertices such that every edge touches at least one chosen vertex, computable in polynomial time by the classical duality with maximum matching. Here the vertices are boundary nodes of a partition, the edges are the cut edges between two workers, and the chosen vertices are exactly the nodes whose feature vectors must be communicated. The cover decides the assignment of each cut edge to pre-aggregation (aggregate before sending) or post-aggregation (aggregate after receiving); because every edge is incident to a cover node, the features of the cover vertices suffice to supply all boundary messages. This is what converts a communication scheduling problem into a graph-theoretic optimization, and the paper's analysis shows the optimal communication volume equals the cover size.
What would settle it
Run one GCN layer with an attention-based aggregator (edge weights depending on both endpoint features) on a partitioned graph, and compare the hybrid pre/post aggregation output against the same layer computed without partitioning; any difference in the aggregated vectors shows the lossless claim does not extend beyond sum/mean aggregators.
Extended reading notes
Core claim
The paper's central claim is that the optimal lossless communication schedule for a distributed GCN layer can be computed exactly. After partitioning, each worker sees a remote graph whose edges are cut edges to boundary nodes; treating this graph as bipartite, the worker marks a minimum vertex cover. Edges whose source is in the cover are aggregated after communication (post), and the rest are aggregated before communication (pre). Every cut edge is incident to a cover vertex, so sending the feature vectors of cover vertices alone accounts for all boundary traffic, and by the classical theorem equating vertex cover with maximum matching, the cover is as small as possible. The paper states that this makes the communication volume per layer the minimal number of boundary-node feature vectors required, and measures 1.5x lower volume than pre-only or post-only aggregation on a large graph; with Int2 quantization the data volume drops by an additional factor of about 15. On top of this, the paper argues that masked label propagation restores the accuracy lost by aggressive quantization, giving convergence at the same rate as full precision with a bounded error neighborhood.
Load-bearing premise
The hybrid split is lossless only if aggregation is separable — adding part of the neighbors' contributions before communication and the rest after gives the same result as full aggregation — which holds for the sum and mean aggregators tested but is not verified for attention-based or other nonlinear aggregators the paper claims to support.
Editorial extensions
If this is right
- Every distributed full-batch GNN layer that uses a sum-like (separable) aggregator can exchange only the minimal boundary-node features, so the communication cost is bounded by the partition's vertex-cover size rather than its edge-cut size.
- The communication advantage grows with process count: the measured speedup over the compared CPU baseline rises from under 1x at small scales to 6x at large scales, and the framework reaches 8,192 MPI ranks on the largest public datasets.
- Aggressive Int2 quantization can be applied uniformly without adaptive bit selection, reducing communicated data by roughly 15x on top of the hybrid split while keeping final accuracy within a small range of full precision.
- Because convergence is preserved with quantized communication, full-batch training can retain the original graph structure and still scale, avoiding the accuracy degradation associated with sampling-based mini-batch methods.
- On the largest datasets tested, the system's best epoch times are shorter than those reported for distributed GPU full-batch frameworks, suggesting CPU supercomputers can be a competitive platform for very large graphs.
Reading between the lines
- The minimum vertex cover equivalence is stated for the remote graph of a single partition; applying it per worker independently may not yield a globally minimal all-to-all schedule, since boundary nodes appear in multiple remote graphs, and coordinating cover choices across workers is a natural extension.
- For attention-based aggregators, the pre/post split is not automatically lossless because each neighbor's contribution weight depends on both endpoint features; a modified scheme would need to communicate attention coefficients or restrict the split to the message values after weights are fixed.
- The same bipartite-cover construction should transfer to other distributed graph computations with separable reductions, such as sparse matrix-vector products or PageRank-style iterations, where the edge-cut overhead has the same structure.
- A testable prediction is that the measured 1.5x communication reduction should vary with partition quality: partitions with many disjoint remote components should approach the vertex-cover lower bound more closely, while star-like boundary structures may leave less room for hybrid splitting.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. SuperGCN is a distributed full-batch GCN training framework for CPU-based supercomputers. The paper makes three contributions: (i) cache- and vectorization-aware aggregation operators for x86 and A64FX CPUs, (ii) a hybrid pre-/post-aggregation scheme that formulates communication-volume minimization as minimum vertex cover on bipartite graphs, and (iii) a communication-aware Int2 quantization scheme augmented with masked label propagation and LayerNorm. The system is evaluated on ABCI (Intel Xeon) and Fugaku (Arm A64FX), scaling to 8,192 MPI ranks on datasets including Ogbn-papers100M, Ogb-lsc-mag240M, and IGB260M, and it is compared against DistGNN and several GPU-based full-batch GNN training systems.
Significance. If the claims hold, the MVC-based communication reduction is an elegant and practically validated contribution for linear/order-independent aggregators, and the scale of the evaluation (thousands of MPI ranks on the largest public graph datasets) is substantially beyond most prior full-batch GNN systems. The paper presents a broad experimental campaign, and Table 5 directly confirms the predicted ~1.5x communication-volume reduction over pre-only or post-only aggregation. The main caveats are that the claimed generality to attention-based GNNs is not established, and the accuracy claims need statistical qualification; both are addressable in revision.
major comments (4)
- [Sec. 3.2 and Secs. 5.2-5.3] The paper states that SuperGCN can be 'seamlessly applied to the distributed training of these message-passing-based GNN models' and names GAT. For GAT, the aggregation weight alpha_{vu} for edge (u,v) depends on the destination feature h_v (via the softmax over N(v)), which is not available at the sender when a pre-aggregation partial message is constructed. The hybrid pre/post construction in Algorithm 1 is therefore exact only for aggregators that are linear/order-independent in neighbor features (sum, and mean when normalization is applied after summation). All experiments in Sec. 8 use GraphSAGE. Please either provide a concrete protocol for attention-based aggregators or restrict the applicability claim to separable aggregators and move GAT-class models to future work.
- [Abstract and Table 3] The abstract's claim 'without sacrificing model convergence and accuracy' is not supported by the Int2 ablation. In Table 3, SuperGCN (Int2, w/o LP) on ogbn-papers100M reaches 60.19-62.73, about 3 points below the corresponding FP32 w/o LP values (63.33-63.62); Int2 matches FP32 only when masked label propagation is enabled (e.g., 65.71 vs 65.62 at 1024 procs). The claim should be qualified to the full system. In addition, all accuracy numbers appear to be single runs; please report multiple seeds with means and standard deviations so that 'matches FP32' is statistically meaningful.
- [Sec. 6.3, Lemma 1] The proof of Lemma 1 is delegated to AdapQ [56] with only 'adapted to our setting'. The setting differs in material ways: SuperGCN uses fixed Int2 quantization, injects masked label propagation into the features, and applies LayerNorm before each layer, none of which appears in [56]. Please provide a self-contained proof of the unbiasedness and bounded-variance assumptions for this specific pipeline, or state precisely which arguments of [56] carry over unchanged.
- [Sec. 6.3, Lemma 2] Equation (10) in Lemma 2 asserts H^{(l+1)} = A H^{(l)} W^{(l)} = A^l (X + Y_embed) (W^{(0)} ... W^{(l)}). This equality is not valid for the GCNs used in the experiments, which include nonlinear activation functions between layers. As written, the lemma and the subsequent Proposition 1 overstate the theoretical support. Please reformulate the statement as an approximation or restrict it to linear GCNs, and adjust Proposition 1 accordingly.
minor comments (6)
- [Sec. 8.5 / Table 4] The accuracy shown for SuperGCN on ogbn-products in Table 4 (80.24) is higher than any value in Table 3 (max 79.68); the text says the epoch count was increased for this comparison, but the exact setting should be stated so readers can reconcile the two tables.
- [Eq. (2)] In the text below Eq. (2), the variable V_{i,j}^{comm} is used, while the equation defines C_{i,j}^{comm}; please unify the notation.
- [Fig. 10] The panel labeled 'UK-2007-02' appears to refer to 'UK-2007-05' in Table 2; please correct the label.
- [Table 3] The entry 'N/A refers to the abnormal accuracy' for DistGNN on Reddit is vague; please specify the failure mode (e.g., divergence or numerical issues) and whether the comparison with DistGNN on Reddit is therefore omitted.
- [Sec. 7.2] The phrase 'we optimize the implementation of the NetworkX library' is ambiguous; it likely means an optimized variant or a faster custom implementation, and should be clarified.
- [Sec. 8.5] The baseline name 'AdaptQ' should be 'AdapQ' for consistency with Table 1 and reference [56].
Circularity Check
No significant circularity: the MVC communication-volume reduction is a genuine combinatorial equivalence, the performance model is first-principles, and the accuracy argument rests on external analyses; the only self-citation is non-load-bearing.
full rationale
The central derivation in Sec. 5.3.2 (Eq. 1) is not circular: the paper independently defines communication volume as the number of vertex features transmitted and the hybrid pre/post construction, then proves that any feasible transmission scheme corresponds to a vertex cover of the bipartite remote graph, so minimizing volume is exactly the minimum vertex cover problem. This is a genuine reduction, valid for the sum/mean aggregators used in the GraphSAGE experiments. The performance model in Eq. (8) is derived from stated hardware ratios (alpha, beta, gamma, delta), not fitted to the measured times; Sec. 8.3 uses it only as a qualitative consistency check. The quantization convergence bound (Lemma 1) follows the external AdapQ analysis [56], Lemma 2 is an algebraic identity for label embedding, and Proposition 1 relies on external masked-label-propagation work [51]; none of these are self-citations. The only self-citation is [69] in Related Work, which merely lists the authors' prior workshop paper among quantization-based methods and is not load-bearing for any claim. One scope limitation, not circularity, should be noted: Sec. 3.2 claims seamless applicability to all message-passing GNNs including GAT, but pre-aggregation is exact only for order-independent (separable) aggregators; attention weights depending on the destination feature prevent correct pre-aggregation, and all experiments use GraphSAGE. This is an overclaim rather than a circular derivation, so it does not raise the circularity score.
Assumptions & free parameters
assumptions (5)
- standard math König's theorem and Hopcroft-Karp provide an efficient exact solution for minimum vertex cover on bipartite graphs.
- domain assumption The stochastic gradient is unbiased and has bounded variance, and the loss gradient is rho-Lipschitz.
- domain assumption Adding label embeddings to features acts as label propagation and improves separation for nodes of the same label.
- domain assumption METIS partitions the graph with balanced node weights and locality, so intra-node communication is cheaper than inter-node.
- ad hoc to paper The aggregation operators are linear (sum/mean) so pre- and post-aggregation commute with communication.
Cite this review
Pith. "Pith review of Scaling Large-scale GNN Training to Thousands of Processors on CPU-based Supercomputers." pith.science (2026). https://pith.science/paper/L4ZZWOJM
@misc{pith2026241116025,
author = {Pith},
title = {Pith review of: Scaling Large-scale GNN Training to Thousands of Processors on CPU-based Supercomputers},
year = {2026},
howpublished = {\url{https://pith.science/paper/L4ZZWOJM}},
note = {Machine review of arXiv:2411.16025}
}
abstract
Graph Convolutional Networks (GCNs), particularly for large-scale graphs, are crucial across numerous domains. However, training distributed full-batch GCNs on large-scale graphs suffers from inefficient memory access patterns and high communication overhead. To address these challenges, we introduce \method{}, an efficient and scalable distributed GCN training framework tailored for CPU-powered supercomputers. Our contributions are threefold: (1) we develop general and efficient aggregation operators designed for irregular memory access, (2) we propose a hierarchical aggregation scheme that reduces communication costs without altering the graph structure, and (3) we present a communication-aware quantization scheme to enhance performance. Experimental results demonstrate that \method{} achieves a speedup of up to 6$\times$ compared with the SoTA implementations, and scales to 1000s of HPC-grade CPUs on the largest publicly available datasets, without sacrificing model convergence and accuracy. Moreover, due to the effective strong scaling of \method{}, we outperform SoTA GPU-based and CPU-based distributed full-batch GCN training frameworks, in absolute performance, for large-scale graphs.
Figures
Figures from the paper (8 more)
Reference graph
Works this paper leans on
-
[56]
Borui Wan, Juntao Zhao, and Chuan Wu. 2023. Adaptive Message Quantization and Parallelization for Distributed Full-graph GNN Train- ing. Proceedings of Machine Learning and Systems 5 (2023)
work page 2023
-
[1]
AIST. 2023. ABCI Supercomputer
work page 2023
-
[2]
Ariful Azad, Georgios A Pavlopoulos, Christos A Ouzounis, Nikos C Kyrpides, and Aydin Buluç. 2018. HipMCL: a high-performance parallel implementation of the Markov clustering algorithm for large-scale networks. Nucleic acids research 46, 6 (2018), e33–e33
2018
-
[3]
Filipe De Avila Belbute-Peres, Thomas Economon, and Zico Kolter
-
[4]
Open Graph Benchmark. 2024. OGB-Leaderboards for Node Property Prediction. https://ogb.stanford.edu/docs/leader_nodeprop/
work page 2024
-
[5]
Paolo Boldi, Marco Rosa, Massimo Santini, and Sebastiano Vigna. 2011. Layered Label Propagation: A MultiResolution Coordinate-Free Or- dering for Compressing Social Networks. In Proceedings of the 20th international conference on World Wide Web , Sadagopan Srinivasan, Krithi Ramamritham, Arun Kumar, M. P. Ravindra, Elisa Bertino, and Ravi Kumar (Eds.). AC...
work page 2011
-
[6]
Paolo Boldi and Sebastiano Vigna. 2004. The WebGraph Framework I: Compression Techniques. In Proc. of the Thirteenth International World Wide Web Conference (WWW 2004) . ACM Press, Manhattan, USA, 595–601
work page 2004
-
[7]
Zhenkun Cai, Xiao Yan, Yidi Wu, Kaihao Ma, James Cheng, and Fan Yu. 2021. DGCL: an efficient communication library for distributed GNN training. In Proceedings of the Sixteenth European Conference on Computer Systems. 130–144
work page 2021
Show all 72 references
-
[8]
Yadi Cao, Menglei Chai, Minchen Li, and Chenfanfu Jiang. 2023. Effi- cient learning of mesh-based physical simulation with bi-stride multi- scale graph neural network. In International Conference on Machine Learning. PMLR, 3541–3558
2023
-
[9]
Jie Chen, Tengfei Ma, and Cao Xiao. 2018. Fastgcn: fast learning with graph convolutional networks via importance sampling.arXiv preprint arXiv:1801.10247 (2018)
2018 arXiv
-
[10]
Jianfei Chen, Lianmin Zheng, Zhewei Yao, Dequan Wang, Ion Sto- ica, Michael Mahoney, and Joseph Gonzalez. 2021. Actnn: Reducing training memory footprint via 2-bit activation compressed training. In International Conference on Machine Learning . PMLR, 1803–1813
2021
-
[11]
Jianfei Chen, Jun Zhu, and Le Song. 2017. Stochastic training of graph convolutional networks with variance reduction. arXiv preprint arXiv:1710.10568 (2017)
2017 arXiv
-
[12]
Jianfei Chen, Jun Zhu, and Le Song. 2018. Stochastic Training of Graph Convolutional Networks with Variance Reduction. In International Conference on Machine Learning . PMLR, 942–950
2018
-
[13]
Tianqi Chen, Thierry Moreau, Ziheng Jiang, Lianmin Zheng, Eddie Yan, Haichen Shen, Meghan Cowan, Leyuan Wang, Yuwei Hu, Luis Ceze, et al. 2018. {TVM}: An automated{End-to-End} optimizing compiler for deep learning. In 13th USENIX Symposium on Operating Systems Design and Imple...
2018
-
[14]
Wei-Lin Chiang, Xuanqing Liu, Si Si, Yang Li, Samy Bengio, and Cho- Jui Hsieh. 2019. Cluster-gcn: An efficient algorithm for training deep and large graph convolutional networks. In Proceedings of the 25th ACM SIGKDD international conference on knowledge discovery & data minin...
2019
-
[15]
Wei Dai, Yi Zhou, Nanqing Dong, Hao Zhang, and Eric P Xing. 2018. Toward understanding the impact of staleness in distributed machine learning. arXiv preprint arXiv:1810.03264 (2018)
2018 arXiv
-
[16]
Reinhard Diestel. 2017. Graph Theory (5th ed.). Graduate Texts in Mathematics, Vol. 173. Springer
2017
-
[17]
Boyuan Feng, Yuke Wang, Xu Li, Shu Yang, Xueqiao Peng, and Yufei Ding. 2020. Sgquant: Squeezing the last bit on graph neural networks with specialized quantization. In 2020 IEEE 32nd International Confer- ence on Tools with Artificial Intelligence (ICTAI) . IEEE, 1044–1052
2020
-
[18]
Matthias Fey and Jan Eric Lenssen. 2019. Fast graph representation learning with PyTorch Geometric. arXiv preprint arXiv:1903.02428 (2019)
2019 arXiv
-
[19]
Qiang Fu, Yuede Ji, and H Howie Huang. 2022. TLPGNN: A lightweight two-level parallelism paradigm for graph neural network computation on GPU. In Proceedings of the 31st International Symposium on High- Performance Parallel and Distributed Computing . 122–134
2022
-
[20]
Fujitsu. 2023. A64FX Microarchitecture Manual. https://github.com/ fujitsu/A64FX/blob/master/doc/A64FX_Microarchitecture_Manual_ en_1.8.1.pdf Accessed: 2025-02-21
2023
-
[21]
Trevor Gale, Matei Zaharia, Cliff Young, and Erich Elsen. 2020. Sparse gpu kernels for deep learning. In SC20: International Conference for High Performance Computing, Networking, Storage and Analysis . IEEE, 1–14
2020
-
[22]
Swapnil Gandhi and Anand Padmanabha Iyer. 2021. P3: Distributed Deep Graph Learning at Scale.. In OSDI. 551–568
2021
-
[23]
Zhijiang Guo, Yan Zhang, and Wei Lu. 2019. Attention guided graph convolutional networks for relation extraction. arXiv preprint arXiv:1906.07510 (2019)
2019 arXiv
-
[24]
Aric Hagberg, Pieter J Swart, and Daniel A Schult. 2008. Exploring network structure, dynamics, and function using NetworkX . Technical Report. Los Alamos National Laboratory (LANL), Los Alamos, NM (United States)
2008
-
[25]
Will Hamilton, Zhitao Ying, and Jure Leskovec. 2017. Inductive rep- resentation learning on large graphs. Advances in neural information processing systems 30 (2017)
2017
-
[26]
Alexander Heinecke, Greg Henry, Maxwell Hutchinson, and Hans Pabst. 2016. LIBXSMM: accelerating small matrix multiplications by runtime code generation. In SC’16: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis. IEEE...
2016
-
[27]
John E Hopcroft and Richard M Karp. 1973. An nˆ5/2 algorithm for maximum matchings in bipartite graphs. SIAM Journal on computing 2, 4 (1973), 225–231
1973
-
[28]
Weihua Hu, Matthias Fey, Hongyu Ren, Maho Nakata, Yuxiao Dong, and Jure Leskovec. 2021. Ogb-lsc: A large-scale challenge for machine learning on graphs. arXiv preprint arXiv:2103.09430 (2021)
2021 arXiv
-
[29]
Weihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong, Hongyu Ren, Bowen Liu, Michele Catasta, and Jure Leskovec. 2020. Open graph benchmark: Datasets for machine learning on graphs. Advances in neural information processing systems 33 (2020), 22118–22133
2020
-
[30]
Yuwei Hu, Zihao Ye, Minjie Wang, Jiali Yu, Da Zheng, Mu Li, Zheng Zhang, Zhiru Zhang, and Yida Wang. 2020. Featgraph: A flexible and efficient backend for graph neural network systems. In SC20: International Conference for High Performance Computing, Networking, Storage and fA...
2020
-
[31]
Guyue Huang, Guohao Dai, Yu Wang, and Huazhong Yang. 2020. Ge- spmm: General-purpose sparse matrix-matrix multiplication on gpus for graph neural networks. In SC20: International Conference for High Performance Computing, Networking, Storage and Analysis. IEEE, 1–12
2020
-
[32]
Benoit Jacob, Skirmantas Kligys, Bo Chen, Menglong Zhu, Matthew Tang, Andrew Howard, Hartwig Adam, and Dmitry Kalenichenko
-
[33]
Zhihao Jia, Sina Lin, Mingyu Gao, Matei Zaharia, and Alex Aiken. 2020. Improving the accuracy, scalability, and performance of graph neural networks with roc. Proceedings of Machine Learning and Systems 2 (2020), 187–198. Scaling Large-scale GNN Training to Thousands of Proces...
2020
-
[34]
Hao Jiang, Peng Cao, MingYi Xu, Jinzhu Yang, and Osmar Zaiane
-
[35]
Badia, and Mohamed Wahib
Albert Njoroge Kahira, Truong Thao Nguyen, Leonardo Bautista- Gomez, Ryousei Takano, Rosa M. Badia, and Mohamed Wahib. 2021. An Oracle for Guiding Large-Scale Model/Hybrid Parallel Training of Convolutional Neural Networks. In HPDC. ACM, 161–173
2021
-
[36]
Tim Kaler, Alexandros Iliopoulos, Philip Murzynowski, Tao Schardl, Charles E Leiserson, and Jie Chen. 2023. Communication-efficient graph neural networks with probabilistic neighborhood expansion analysis and caching. Proceedings of Machine Learning and Systems 5 (2023)
2023
-
[37]
Computers in Biology and Medicine 127 (2020), 104096
Hi-GCN: A hierarchical graph convolution network for graph embedding learning of brain network and brain disorders prediction. Computers in Biology and Medicine 127 (2020), 104096
2020
-
[38]
George Karypis, Kirk Schloegel, and Vipin Kumar. 1997. Parmetis: Parallel graph partitioning and sparse matrix ordering library. (1997)
1997
-
[39]
Arpandeep Khatua, Vikram Sharma Mailthody, Bhagyashree Taleka, Tengfei Ma, Xiang Song, and Wen-mei Hwu. 2023. Igb: Addressing the gaps in labeling, features, heterogeneity, and size of public graph datasets for deep learning research. In Proceedings of the 29th ACM SIGKDD Conf...
2023
-
[40]
Tim Kaler, Nickolas Stathas, Anne Ouyang, Alexandros-Stavros Il- iopoulos, Tao Schardl, Charles E Leiserson, and Jie Chen. 2022. Ac- celerating training and inference of graph neural networks with fast sampling and pipelining. Proceedings of Machine Learning and Systems 4 (202...
2022
-
[41]
Dénes Kőnig. 1931. Gráfok és mátrixok. Matematikai és Fizikai Lapok 38 (1931), 116–119
1931
-
[42]
Remi Lam, Alvaro Sanchez-Gonzalez, Matthew Willson, Peter Wirns- berger, Meire Fortunato, Ferran Alet, Suman Ravuri, Timo Ewalds, Zach Eaton-Rosen, Weihua Hu, et al. 2023. Learning skillful medium- range global weather forecasting. Science 382, 6677 (2023), 1416–1421
2023
-
[43]
Thomas N Kipf and Max Welling. 2016. Semi-supervised classification with graph convolutional networks. arXiv preprint arXiv:1609.02907 (2016)
2016 arXiv
-
[44]
Vasimuddin Md, Sanchit Misra, Guixiang Ma, Ramanarayan Mohanty, Evangelos Georganas, Alexander Heinecke, Dhiraj Kalamkar, Nes- reen K Ahmed, and Sasikanth Avancha. 2021. Distgnn: Scalable dis- tributed training for large-scale graph neural networks. In Proceedings of the Inter...
2021
-
[45]
Jintao Meng, Peng Chen, Mohamed Wahib, Mingjun Yang, Liangzhen Zheng, Yanjie Wei, Shengzhong Feng, and Wei Liu. 2022. Boosting the predictive performance with aqueous solubility dataset curation. Scientific Data 9, 1 (2022), 71
2022
-
[46]
Lingxiao Ma, Zhi Yang, Youshan Miao, Jilong Xue, Ming Wu, Lidong Zhou, and Yafei Dai. 2019. {NeuGraph}: Parallel deep neural net- work computation on large graphs. In 2019 USENIX Annual Technical Conference (USENIX ATC 19). 443–458
2019
-
[47]
Jingshu Peng, Zhao Chen, Yingxia Shao, Yanyan Shen, Lei Chen, and Jiannong Cao. 2022. Sancus: sta le n ess-aware c omm u nication- avoiding full-graph decentralized training in large-scale graph neural networks. Proceedings of the VLDB Endowment 15, 9 (2022), 1937–1950
2022
-
[48]
Tobias Pfaff, Meire Fortunato, Alvaro Sanchez-Gonzalez, and Peter W Battaglia. 2020. Learning mesh-based simulation with graph networks. arXiv preprint arXiv:2010.03409 (2020)
2020 arXiv
-
[49]
Hesham Mostafa. 2022. Sequential aggregation and rematerialization: Distributed full-batch training of graph neural networks on large graphs. Proceedings of Machine Learning and Systems 4 (2022), 265– 275
2022
-
[50]
Mitsuhisa Sato, Yutaka Ishikawa, Hirofumi Tomita, Yuetsu Kodama, Tetsuya Odajima, Miwako Tsuji, Hisashi Yashiro, Masaki Aoki, Naoyuki Shida, Ikuo Miyoshi, et al. 2020. Co-design for a64fx many- core processor and” fugaku”. InSC20: International Conference for High Performance ...
2020
-
[51]
Yunsheng Shi, Zhengjie Huang, Shikun Feng, Hui Zhong, Wenjin Wang, and Yu Sun. 2020. Masked label prediction: Unified mes- sage passing model for semi-supervised classification. arXiv preprint arXiv:2009.03509 (2020)
2020 arXiv
-
[52]
Seongok Ryu, Yongchan Kwon, and Woo Youn Kim. 2019. A Bayesian graph convolutional network for reliable prediction of molecular prop- erties with uncertainty quantification. Chemical science 10, 36 (2019), 8438–8446
2019
-
[53]
John Thorpe, Yifan Qiao, Jonathan Eyolfson, Shen Teng, Guanzhou Hu, Zhihao Jia, Jinliang Wei, Keval Vora, Ravi Netravali, Miryung Kim, et al. 2021. Dorylus: Affordable, scalable, and accurate{GNN} training with distributed{CPU} servers and serverless threads. In15th USENIX Sym...
2021
-
[54]
Alok Tripathy, Katherine Yelick, and Aydın Buluç. 2020. Reducing communication in graph neural network training. In SC20: Interna- tional Conference for High Performance Computing, Networking, Storage and Analysis. IEEE, 1–14
2020
-
[55]
Mengying Sun, Sendong Zhao, Coryandar Gilvary, Olivier Elemento, Jiayu Zhou, and Fei Wang. 2020. Graph convolutional networks for computational drug development and discovery. Briefings in bioinfor- matics 21, 3 (2020), 919–935
2020
-
[57]
Cheng Wan, Youjie Li, Ang Li, Nam Sung Kim, and Yingyan Lin. 2022. BNS-GCN: Efficient full-graph training of graph convolutional net- works with partition-parallelism and random boundary node sampling. Proceedings of Machine Learning and Systems 4 (2022), 673–693
2022
-
[58]
Petar Velickovic, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Lio, Yoshua Bengio, et al. 2017. Graph attention net- works. stat 1050, 20 (2017), 10–48550
2017
-
[59]
Hongwei Wang and Jure Leskovec. 2020. Unifying graph convolutional neural networks and label propagation.arXiv preprint arXiv:2002.06755 (2020)
2020 arXiv
-
[60]
Tianyi Wang, Yang Chen, Zengbin Zhang, Tianyin Xu, Long Jin, Pan Hui, Beixing Deng, and Xing Li. 2011. Understanding graph sam- pling algorithms for social network analysis. In 2011 31st international conference on distributed computing systems workshops . IEEE, 123–128
2011
-
[61]
Wolfe, Anastasios Kyrillidis, Nam Sung Kim, and Yingyan Lin
Cheng Wan, Youjie Li, Cameron R. Wolfe, Anastasios Kyrillidis, Nam Sung Kim, and Yingyan Lin. 2022. PipeGCN: Efficient Full-Graph Training of Graph Convolutional Networks with Pipelined Feature Communication. In International Conference on Learning Representa- tions. https://o...
2022
-
[62]
Hongxia Yang. 2019. Aligraph: A comprehensive graph neural net- work platform. In Proceedings of the 25th ACM SIGKDD international conference on knowledge discovery & data mining . 3165–3166
2019
-
[63]
Zihao Ye, Ruihang Lai, Junru Shao, Tianqi Chen, and Luis Ceze. 2023. SparseTIR: Composable abstractions for sparse compilation in deep learning. In Proceedings of the 28th ACM International Conference on Ar- chitectural Support for Programming Languages and Operating Systems, ...
2023
-
[64]
Keyulu Xu, Weihua Hu, Jure Leskovec, and Stefanie Jegelka. 2018. How powerful are graph neural networks? arXiv preprint arXiv:1810.00826 (2018)
2018 arXiv
-
[65]
Ye Yuan and Ziv Bar-Joseph. 2020. GCNG: graph convolutional net- works for inferring gene interaction from spatial transcriptomics data. Genome biology 21, 1 (2020), 1–16
2020
-
[66]
Meng Zhang, Qinghao Hu, Cheng Wan, Haozhao Wang, Peng Sun, Yonggang Wen, and Tianwei Zhang. 2024. Sylvie: 3d-adaptive and universal system for large-scale graph neural network training. In2024 IEEE 40th International Conference on Data Engineering (ICDE) . IEEE, 3823–3836
2024
-
[67]
Rex Ying, Ruining He, Kaifeng Chen, Pong Eksombatchai, William L Hamilton, and Jure Leskovec. 2018. Graph convolutional neural net- works for web-scale recommender systems. In Proceedings of the 24th ACM SIGKDD international conference on knowledge discovery & data mining. 974–983
2018
-
[68]
Feng Zhu, Ruihao Gong, Fengwei Yu, Xianglong Liu, Yanfei Wang, Zhe- long Li, Xiuqi Yang, and Junjie Yan. 2020. Towards unified int8 training for convolutional neural network. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition . 1969–1979
2020
-
[69]
Chen Zhuang, Peng Chen, Xin Liu, Toshio Endo, Satoshi Matsuoka, and Mohamed Wahib. 2024. Communication Optimization for Distributed GCN Training on ABCI Supercomputer. In 2024 IEEE International Conference on Cluster Computing Workshops (CLUSTER Workshops) . IEEE, 160–161. ,
2024
-
[70]
Da Zheng, Chao Ma, Minjie Wang, Jinjing Zhou, Qidong Su, Xiang Song, Quan Gan, Zheng Zhang, and George Karypis. 2020. Distdgl: distributed graph neural network training for billion-scale graphs. In 2020 IEEE/ACM 10th Workshop on Irregular Applications: Architectures and Algori...
2020
-
[2018]
In Proceedings of the IEEE conference on computer vision and pattern recognition
Quantization and training of neural networks for efficient integer- arithmetic-only inference. In Proceedings of the IEEE conference on computer vision and pattern recognition . 2704–2713
-
[2020]
Ininternational conference on machine learning
Combining differentiable PDE solvers and graph neural networks for fluid flow prediction. Ininternational conference on machine learning. PMLR, 2402–2411
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.