REVIEW 2 major objections 6 minor 37 references
Joint optimization of sampling, memory I/O, and sparse operators yields up to 4.7× faster TGNN training without accuracy loss.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-11 08:53 UTC pith:VDTHHZO4
load-bearing objection Solid single-GPU systems paper that co-designs the three real TGNN bottlenecks and delivers measured 2× wall-clock gains with open code; the pre-sampling stationarity assumption is real but not load-bearing. the 2 major comments →
FAST: A Holistic Framework for Optimizing Memory-I/O, Computation, and Sampling in Temporal GNN Training
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper establishes that the three dominant bottlenecks of TGNN training—memory I/O, irregular graph operators, and temporal sampling—share measurable redundancy and locality patterns that can be harvested once by a lightweight pre-sampling pass and then exploited jointly, producing average 2.1× and peak 4.7× end-to-end speedups over prior systems with no accuracy loss.
What carries the argument
SlimCache (within-batch ID compression plus greedy cross-batch placement of hot node/edge features under a fixed GPU budget), together with edge-centric aggregation, thread-loop edge-softmax, and topology-aware CPU thread binding derived from the same pre-sampling statistics.
Load-bearing premise
A single pre-sampling pass that looks only at root nodes under a fixed batch size and top-k sampling remains representative enough for the whole subsequent training epoch.
What would settle it
Re-run the full training suite while regenerating hot-ID lists and affinity matrices every few batches instead of once; if end-to-end speedup collapses or accuracy drifts, the static-pre-sampling premise fails.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents FAST, a single-GPU framework that jointly optimizes the three main stages of continuous-time temporal GNN training: temporal neighbor sampling, host–device feature movement, and sparse graph operators (aggregation and edge-softmax). SlimCache combines within-batch ID compression with a greedy cross-batch cache that treats nodes and edges differently under a fixed GPU budget (Algorithm 1). Thread-efficient CUDA kernels replace node-parallel aggregation with a COO edge-centric scheme and replace warp-shuffle edge-softmax with a CSR thread-loop reduction that enlarges the per-block working set. A topology-aware sampler builds a thread-affinity matrix from a lightweight pre-sampling pass and binds high-affinity threads to shared L2/L3 domains via Blossom matching (Algorithm 2). On four public dynamic graphs and three models (TGN, TGAT, DySAT) the system reports average 2.1 imes (up to 4.7 imes) end-to-end speedup over TGL, ETC and SIMPLE while preserving average-precision accuracy (Table 4), with component ablations (Figures 8–11) and kernel counters (Table 5) supporting the claimed sources of gain.
Significance. If the reported speedups hold, FAST supplies a practical, modular co-design that simultaneously attacks the three dominant bottlenecks of large-scale CTDG training—an area where prior systems have largely optimized stages in isolation. The public code release, the explicit accuracy tables, and the kernel-level counters (L1/L2 hit rates, active warps/threads) make the claims falsifiable and reusable by other frameworks. The work is therefore of clear engineering value to the systems-for-GNN community and to practitioners training TGNNs on graphs with tens to hundreds of millions of edges under realistic single-GPU memory budgets.
major comments (2)
- Table 4 and §6.2: several baseline entries are OOM (ETC/SIMPLE on GDELT and BITCOIN). While the paper correctly notes that FAST still runs, the relative speedups versus those systems become undefined; the abstract’s “average 2.1× over state-of-the-art systems” should be recomputed only over the configurations where every baseline finishes, or the OOM cases should be reported separately so that the headline number is not inflated by incomplete runs.
- §6.1 and Table 4: results are stated to be averages of five independent runs, yet no standard deviations or confidence intervals appear for either wall-clock time or AP. For the largest claimed gains (4.7× on WIKITALK-TGN, 4.2× on GDELT-TGN) the absence of variance makes it impossible to judge whether the differences are statistically stable; adding error bars or a short variance table is load-bearing for the central performance claim.
minor comments (6)
- Abstract vs. Table 4: abstract claims “average of 2.1×”; the body text in §6.2 quotes 2.6× over TGL. Align the two numbers or clarify the exact averaging set.
- Figure 1 caption and §3.1: the 78 % I/O figure is given for WIKITALK; a one-sentence note that the same breakdown holds (or does not) for the other three datasets would strengthen the motivation.
- Algorithm 2 and §4.4: the complexity argument uses hop=0 (roots only) yet still writes O(S^K); a brief remark that K is set to 0 in the implementation would remove the apparent discrepancy.
- Table 1 “Thread efficiency” column is derived from average degree; a short formula or footnote would make the 53–76 % numbers reproducible.
- §5: the compression engine is said to emit “CSR-style ind_ptr”; a one-line clarification that this is used only for the ESM kernel (not for the COO AGG) would avoid reader confusion.
- Typographical: “FASTintroduces”, “FASTemploys”, missing spaces after system names appear repeatedly in the abstract and introduction; also “edgeSoftmax” vs. “edge-softmax” inconsistency.
Circularity Check
No significant circularity: empirical systems speedups measured against external baselines, not derived by construction from fitted inputs or self-citations.
full rationale
FAST is a systems paper whose central claim (average 2.1× / up to 4.7× end-to-end training speedup on public CTDG datasets without accuracy loss) is established by direct wall-clock comparison to independent baselines (TGL, ETC, SIMPLE, TASER, dGNN, FastGL) under identical training settings (Table 4, Figures 8–11). SlimCache’s greedy hot-ID placement, the COO/CSR operator redesigns, and topology-aware thread binding are offline configuration steps whose effectiveness is measured, not assumed; pre-sampling statistics (Algorithms 1–2) supply placement hints only and do not appear inside any equation that defines the reported speedups. Accuracy is reported as held-out average precision, identical to baselines. No self-definitional equation, fitted-parameter-as-prediction, load-bearing self-citation uniqueness claim, or renaming of a known result is present. The derivation chain is therefore self-contained and non-circular.
Axiom & Free-Parameter Ledger
free parameters (4)
- top-k neighbor sampling =
10
- batch size =
2000
- cache budget ratio α =
variable (0.1–1.0)
- number of sampling threads =
8
axioms (4)
- domain assumption Sampled temporal subgraphs exhibit substantial within-batch ID repetition and cross-batch node/edge overlap that can be exploited by compression+caching.
- domain assumption Dynamic-graph neighborhoods are small-degree (average degree ≤ 7.56 under k=10), so atomic contention in edge-parallel aggregation remains negligible.
- ad hoc to paper A single pre-sampling pass with hop=0 (roots only) yields affinity and hot-ID statistics that remain valid for the whole epoch.
- domain assumption Training is performed on a single machine with one GPU under a main-memory (not disk) regime.
invented entities (3)
-
SlimCache (joint within-batch compression + greedy cross-batch node/edge cache)
independent evidence
-
Thread-efficient AGG (COO edge-parallel) and ESM (CSR thread-loop) operators
independent evidence
-
Topology-aware sampler with Blossom affinity binding
independent evidence
read the original abstract
Temporal Graph Neural Networks (TGNNs) are widely used for learning from dynamic graphs in applications such as recommendation, social network analysis, and traffic forecasting. However, scaling TGNN training to large dynamic graphs remains challenging due to three intertwined bottlenecks: memory I/O, irregular computation, and temporal neighbor sampling. Existing systems often optimize these stages in isolation, leaving substantial performance headroom on the table. We present FAST, a holistic framework that accelerates end-to-end TGNN training by jointly optimizing sampling, memory I/O, and computation. FAST introduces SlimCache, which exploits within-batch compression and cross-batch caching to reduce host-device data movement under limited GPU memory budgets. It further designs thread-efficient graph operators tailored to sparse temporal subgraphs, improving GPU cache locality and reducing the latency of aggregation and edge softmax. In addition, FAST employs a topology-aware sampling strategy that improves CPU cache locality and accelerates temporal neighbor sampling. Extensive experiments on real-world large dynamic graphs show that FAST achieves an average of 2.1x (up to 4.7x) speedup over state-of-the-art systems without sacrificing model accuracy.
Figures
Reference graph
Works this paper leans on
-
[1]
Chaoyi Chen, Dechao Gao, Yanfeng Zhang, Qiange Wang, Zhenbo Fu, Xuecang Zhang, Junhua Zhu, Yu Gu, and Ge Yu. 2023. NeutronStream: A Dynamic GNN Training Framework with Sliding Window for Graph Streams.Proc. VLDB Endow. 17, 3 (Nov. 2023), 455–468. doi:10.14778/3632093.3632108
-
[2]
Zhaodong Chen, Mingyu Yan, Maohua Zhu, Lei Deng, Guoqi Li, Shuangchen Li, and Yuan Xie. 2020. fuseGNN: accelerating graph convolutional neural network training on GPGPU. InProceedings of the 39th International Conference on Computer-Aided Design(Virtual Event, USA)(ICCAD ’20). Association for Computing Machinery, New York, NY, USA, Article 60, 9 pages. do...
arXiv 2020
-
[3]
Gangda Deng, Hongkuan Zhou, Hanqing Zeng, Yinglong Xia, Christopher Leung, Jianbo Li, Rajgopal Kannan, and Viktor Prasanna. 2024. TASER: Temporal Adap- tive Sampling for Fast and Accurate Dynamic Graph Representation Learning. In2024 IEEE International Parallel and Distributed Processing Symposium (IPDPS). 926–937. doi:10.1109/IPDPS57955.2024.00087
-
[4]
Shihong Gao, Yiming Li, Yanyan Shen, Yingxia Shao, and Lei Chen. 2024. ETC: Efficient Training of Temporal Graph Neural Networks over Large-Scale Dynamic Graphs.Proc. VLDB Endow.17, 5 (Jan. 2024), 1060–1072. doi:10.14778/3641204. 3641215
-
[5]
Shihong Gao, Yiming Li, Xin Zhang, Yanyan Shen, Yingxia Shao, and Lei Chen
-
[6]
SIMPLE: Efficient Temporal Graph Neural Network Training at Scale with Dynamic Data Placement.Proc. ACM Manag. Data2, 3, Article 174 (May 2024), 25 pages. doi:10.1145/3654977
-
[7]
Yidong Gong and Pradeep Kumar. 2024. GNNOne: A Unified System Op- timizations for GNN Kernels. InProceedings of the 33rd International Sym- posium on High-Performance Parallel and Distributed Computing(Pisa, Italy) (HPDC ’24). Association for Computing Machinery, New York, NY, USA, 15–27. doi:10.1145/3625549.3658655
-
[8]
Rui Guo, Zezhong Ding, Xike Xie, and Jianliang Xu. 2025. SWIFT: Enabling Large-Scale Temporal Graph Learning on a Single Machine.Proc. ACM Manag. Data3, 4, Article 266 (Sept. 2025), 27 pages. doi:10.1145/3749184
-
[9]
Abhinav Jangda, Sandeep Polisetty, Arjun Guha, and Marco Serafini. 2021. Accel- erating graph sampling for graph machine learning using GPUs. InProceedings of the Sixteenth European Conference on Computer Systems(Online Event, United Kingdom)(EuroSys ’21). Association for Computing Machinery, New York, NY, USA, 311–326. doi:10.1145/3447786.3456244
-
[10]
Jin, Lingbo Liu, Fuxian Li, and Jincai Huang
G. Jin, Lingbo Liu, Fuxian Li, and Jincai Huang. 2023. Spatio-Temporal Graph Neu- ral Point Process for Traffic Congestion Event Prediction.AAAIabs/2311.08635, 14268–14276
Pith/arXiv arXiv 2023
-
[11]
Dániel Kondor, Márton Pósfai, István Csabai, and Gábor Vattay. 2014. Do the Rich Get Richer? An Empirical Analysis of the Bitcoin Transaction Network. PLOS ONE9, 2 (Feb. 2014), 1–10. doi:10.1371/journal.pone.0086197
-
[12]
Srijan Kumar, Xikun Zhang, and Jure Leskovec. 2019. Predicting dynamic em- bedding trajectory in temporal interaction networks. InProceedings of the 25th ACM SIGKDD international conference on knowledge discovery & data mining. 1269–1278
2019
-
[13]
Kalev Leetaru and Philip A Schrodt. 2013. Gdelt: Global data on events, location, and tone, 1979–2012. InISA annual convention, Vol. 2. Citeseer, 1–49
2013
-
[14]
Yiming Li, Yanyan Shen, Lei Chen, and Mingxuan Yuan. 2023. Orca: Scalable Temporal Graph Neural Network Training with Theoretical Guarantees.Proc. ACM Manag. Data1, 1, Article 52 (May 2023), 27 pages. doi:10.1145/3588737
-
[15]
Yuwen Liu, Lianyong Qi, Weiming Liu, Xiaolong Xu, Xuyun Zhang, and Wanchun Dou. 2024. GraphSAGE-based POI Recommendation via Continuous-Time Mod- eling. InCompanion Proceedings of the ACM Web Conference 2024(Singapore, Singapore)(WWW ’24). Association for Computing Machinery, New York, NY, USA, 585–588. doi:10.1145/3589335.3651515
-
[16]
Karavanic
Konstantin Macarenco, Kristina Frye, Benjamin Hamlin, and Karen L. Karavanic
-
[17]
In45th International Conference on Parallel Processing Workshops, ICPP Workshops 2016, Philadelphia, PA, USA, August 16-19,
The Effects of System Management Interrupts on Multithreaded, Hyper- threaded, and MPI Applications. In45th International Conference on Parallel Processing Workshops, ICPP Workshops 2016, Philadelphia, PA, USA, August 16-19,
2016
-
[18]
IEEE Computer Society, 338–345. doi:10.1109/ICPPW.2016.55
-
[19]
Maxim Milakov and Natalia Gimelshein. 2018. Online Normalizer Calculation for Softmax. (2018). arXiv:1805.02867
Pith/arXiv arXiv 2018
-
[20]
Ashwin Paranjape, Austin R Benson, and Jure Leskovec. 2017. Motifs in temporal networks. InProceedings of the tenth ACM international conference on web search and data mining. 601–610
2017
-
[21]
Emanuele Rossi, Ben Chamberlain, Fabrizio Frasca, Davide Eynard, Federico Monti, and Michael Bronstein. 2020. Temporal Graph Networks for Deep Learn- ing on Dynamic Graphs. InProceedings of the ICML 2020 Workshop on Graph Representation Learning
2020
-
[22]
Ryan Rossi and Nesreen Ahmed. 2015. The network data repository with inter- active graph analytics and visualization. InProceedings of the AAAI conference on artificial intelligence, Vol. 29
2015
-
[23]
Aravind Sankar, Yanhong Wu, Liang Gou, Wei Zhang, and Hao Yang. 2020. DySAT: Deep Neural Representation Learning on Dynamic Graphs via Self- Attention Networks. InProceedings of the 13th International Conference on Web Search and Data Mining(Houston, TX, USA)(WSDM ’20). Association for Com- puting Machinery, New York, NY, USA, 519–527. doi:10.1145/3336191.3371845
-
[24]
Guangming Sheng, Junwei Su, Chao Huang, and Chuan Wu. 2024. MSPipe: Efficient Temporal GNN Training via Staleness-Aware Pipeline. InProceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (Barcelona, Spain)(KDD ’24). Association for Computing Machinery, New York, NY, USA, 2651–2662. doi:10.1145/3637528.3671844
-
[25]
Amy Shoemaker and Sagar Vare. 2016. Edmonds’ blossom algorithm.CME18 (2016)
2016
-
[26]
Minjie Yu Wang. 2019. Deep graph library: Towards efficient and scalable deep learning on graphs. InICLR workshop on representation learning on graphs and manifolds. doi:https://doi.org/10.48550/arXiv.1909.01315
-
[27]
Yufeng Wang and Charith Mendis. 2023. TGOpt: Redundancy-Aware Optimiza- tions for Temporal Graph Attention Networks. InProceedings of the 28th ACM SIGPLAN Annual Symposium on Principles and Practice of Parallel Programming (Montreal, QC, Canada)(PPoPP ’23). Association for Computing Machinery, New York, NY, USA, 354–368. doi:10.1145/3572848.3577490
-
[28]
Da Xu, Chuanwei Ruan, Evren Körpeoğlu, Sushant Kumar, and Kannan Achan
-
[29]
InInternational Conference on Learning Representations (ICLR)
Inductive representation learning on temporal graphs. InInternational Conference on Learning Representations (ICLR)
-
[30]
Jianbang Yang, Dahai Tang, Xiaoniu Song, Lei Wang, Qiang Yin, Rong Chen, Wenyuan Yu, and Jingren Zhou. 2022. GNNLab: a factored system for sample- based GNN training over GPUs. InProceedings of the Seventeenth European Conference on Computer Systems(Rennes, France)(EuroSys ’22). Association 10 for Computing Machinery, New York, NY, USA, 417–434. doi:10.11...
doi:10.1145/3492321 2022
-
[31]
Chenle Yu, Sara Royuela, and Eduardo Quiñones. 2024. Enhancing Hetero- geneous Computing Through OpenMP and GPU Graph. InProceedings of the 53rd International Conference on Parallel Processing(Gotland, Sweden)(ICPP ’24). Association for Computing Machinery, New York, NY, USA, 534–543. doi:10.1145/3673038.3673050
-
[32]
Hengrui Zhang, Zhongming Yu, Guohao Dai, Guyue Huang, Yufei Ding, Yuan Xie, and Yu Wang. 2022. Understanding GNN Computational Graph: A Coordinated Computation, IO, and Memory Perspective. InProceedings of Machine Learning and Systems (MLSys), Vol. 4. 467–484. doi:10.48550/arXiv.2110.09524
-
[33]
Mengqi Zhang, Shu Wu, Xueli Yu, Qiang Liu, and Liang Wang. 2023. Dynamic Graph Neural Networks for Sequential Recommendation.IEEE Trans. on Knowl. and Data Eng.35, 5 (May 2023), 4741–4753. doi:10.1109/TKDE.2022.3151618
-
[34]
Yuchen Zhong, Guangming Sheng, Tianzuo Qin, Minjie Wang, Quan Gan, and Chuan Wu. 2023. GNNFlow: A Distributed Framework for Continuous Temporal GNN Learning on Dynamic Graphs. (2023). arXiv:2311.17410 [cs.DC]
Pith/arXiv arXiv 2023
-
[35]
Hongkuan Zhou, Da Zheng, Israt Nisa, Vasileios Ioannidis, Xiang Song, and George Karypis. 2022. TGL: A General Framework for Temporal GNN Training onBillion-Scale Graphs.Proc. VLDB Endow.15, 8 (2022), 1572–1580
2022
-
[36]
Zeyu Zhu, Peisong Wang, Qinghao Hu, Gang Li, Xiaoyao Liang, and Jian Cheng
-
[37]
FastGL: A GPU-Efficient Framework for Accelerating Sampling-Based GNN Training at Large Scale. InProceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 4(Hilton La Jolla Torrey Pines, La Jolla, CA, USA)(ASPLOS ’24). Association for Computing Machinery, New York, NY, USA, 94–110...
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.