Pith. sign in

REVIEW 2 major objections 6 minor 37 references

Joint optimization of sampling, memory I/O, and sparse operators yields up to 4.7× faster TGNN training without accuracy loss.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-11 08:53 UTC pith:VDTHHZO4

load-bearing objection Solid single-GPU systems paper that co-designs the three real TGNN bottlenecks and delivers measured 2× wall-clock gains with open code; the pre-sampling stationarity assumption is real but not load-bearing. the 2 major comments →

arxiv 2607.05095 v1 pith:VDTHHZO4 submitted 2026-07-06 cs.LG

FAST: A Holistic Framework for Optimizing Memory-I/O, Computation, and Sampling in Temporal GNN Training

classification cs.LG
keywords temporal graph neural networksdynamic graphsmemory I/O optimizationgraph operatorsneighbor samplingGPU cachingthread affinity
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Temporal Graph Neural Networks learn from timestamped interactions, but training them on large dynamic graphs is slowed by three bottlenecks that compound: repeated host-to-GPU feature transfers, load-imbalanced aggregation and edge-softmax on sparse subgraphs, and cache-unfriendly neighbor sampling on the CPU. Existing systems attack these stages separately and leave large performance gaps. FAST shows that a single pre-sampling pass can expose both within-batch repetition and cross-batch overlap, allowing a unified SlimCache to compress transfers while greedily placing the hottest node and edge features in limited GPU memory; the same structural observations let the authors redesign aggregation as edge-parallel work and edge-softmax as thread-efficient reduction, and bind sampling threads to CPU cores that share cache. On four real dynamic graphs the combined design delivers an average 2.1× (peak 4.7×) end-to-end speedup while preserving model accuracy, demonstrating that co-design across the three stages is both necessary and sufficient for practical large-scale TGNN training.

Core claim

The paper establishes that the three dominant bottlenecks of TGNN training—memory I/O, irregular graph operators, and temporal sampling—share measurable redundancy and locality patterns that can be harvested once by a lightweight pre-sampling pass and then exploited jointly, producing average 2.1× and peak 4.7× end-to-end speedups over prior systems with no accuracy loss.

What carries the argument

SlimCache (within-batch ID compression plus greedy cross-batch placement of hot node/edge features under a fixed GPU budget), together with edge-centric aggregation, thread-loop edge-softmax, and topology-aware CPU thread binding derived from the same pre-sampling statistics.

Load-bearing premise

A single pre-sampling pass that looks only at root nodes under a fixed batch size and top-k sampling remains representative enough for the whole subsequent training epoch.

What would settle it

Re-run the full training suite while regenerating hot-ID lists and affinity matrices every few batches instead of once; if end-to-end speedup collapses or accuracy drifts, the static-pre-sampling premise fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 6 minor

Summary. The paper presents FAST, a single-GPU framework that jointly optimizes the three main stages of continuous-time temporal GNN training: temporal neighbor sampling, host–device feature movement, and sparse graph operators (aggregation and edge-softmax). SlimCache combines within-batch ID compression with a greedy cross-batch cache that treats nodes and edges differently under a fixed GPU budget (Algorithm 1). Thread-efficient CUDA kernels replace node-parallel aggregation with a COO edge-centric scheme and replace warp-shuffle edge-softmax with a CSR thread-loop reduction that enlarges the per-block working set. A topology-aware sampler builds a thread-affinity matrix from a lightweight pre-sampling pass and binds high-affinity threads to shared L2/L3 domains via Blossom matching (Algorithm 2). On four public dynamic graphs and three models (TGN, TGAT, DySAT) the system reports average 2.1 imes (up to 4.7 imes) end-to-end speedup over TGL, ETC and SIMPLE while preserving average-precision accuracy (Table 4), with component ablations (Figures 8–11) and kernel counters (Table 5) supporting the claimed sources of gain.

Significance. If the reported speedups hold, FAST supplies a practical, modular co-design that simultaneously attacks the three dominant bottlenecks of large-scale CTDG training—an area where prior systems have largely optimized stages in isolation. The public code release, the explicit accuracy tables, and the kernel-level counters (L1/L2 hit rates, active warps/threads) make the claims falsifiable and reusable by other frameworks. The work is therefore of clear engineering value to the systems-for-GNN community and to practitioners training TGNNs on graphs with tens to hundreds of millions of edges under realistic single-GPU memory budgets.

major comments (2)
  1. Table 4 and §6.2: several baseline entries are OOM (ETC/SIMPLE on GDELT and BITCOIN). While the paper correctly notes that FAST still runs, the relative speedups versus those systems become undefined; the abstract’s “average 2.1× over state-of-the-art systems” should be recomputed only over the configurations where every baseline finishes, or the OOM cases should be reported separately so that the headline number is not inflated by incomplete runs.
  2. §6.1 and Table 4: results are stated to be averages of five independent runs, yet no standard deviations or confidence intervals appear for either wall-clock time or AP. For the largest claimed gains (4.7× on WIKITALK-TGN, 4.2× on GDELT-TGN) the absence of variance makes it impossible to judge whether the differences are statistically stable; adding error bars or a short variance table is load-bearing for the central performance claim.
minor comments (6)
  1. Abstract vs. Table 4: abstract claims “average of 2.1×”; the body text in §6.2 quotes 2.6× over TGL. Align the two numbers or clarify the exact averaging set.
  2. Figure 1 caption and §3.1: the 78 % I/O figure is given for WIKITALK; a one-sentence note that the same breakdown holds (or does not) for the other three datasets would strengthen the motivation.
  3. Algorithm 2 and §4.4: the complexity argument uses hop=0 (roots only) yet still writes O(S^K); a brief remark that K is set to 0 in the implementation would remove the apparent discrepancy.
  4. Table 1 “Thread efficiency” column is derived from average degree; a short formula or footnote would make the 53–76 % numbers reproducible.
  5. §5: the compression engine is said to emit “CSR-style ind_ptr”; a one-line clarification that this is used only for the ESM kernel (not for the COO AGG) would avoid reader confusion.
  6. Typographical: “FASTintroduces”, “FASTemploys”, missing spaces after system names appear repeatedly in the abstract and introduction; also “edgeSoftmax” vs. “edge-softmax” inconsistency.

Circularity Check

0 steps flagged

No significant circularity: empirical systems speedups measured against external baselines, not derived by construction from fitted inputs or self-citations.

full rationale

FAST is a systems paper whose central claim (average 2.1× / up to 4.7× end-to-end training speedup on public CTDG datasets without accuracy loss) is established by direct wall-clock comparison to independent baselines (TGL, ETC, SIMPLE, TASER, dGNN, FastGL) under identical training settings (Table 4, Figures 8–11). SlimCache’s greedy hot-ID placement, the COO/CSR operator redesigns, and topology-aware thread binding are offline configuration steps whose effectiveness is measured, not assumed; pre-sampling statistics (Algorithms 1–2) supply placement hints only and do not appear inside any equation that defines the reported speedups. Accuracy is reported as held-out average precision, identical to baselines. No self-definitional equation, fitted-parameter-as-prediction, load-bearing self-citation uniqueness claim, or renaming of a known result is present. The derivation chain is therefore self-contained and non-circular.

Axiom & Free-Parameter Ledger

4 free parameters · 4 axioms · 3 invented entities

The central speedup claim rests on empirical measurements under a fixed set of training hyper-parameters and on the observed redundancy/imbalance statistics of the four evaluation graphs. No free parameters are fitted to produce the speedup numbers themselves; the free parameters listed are the conventional knobs of the TGNN training pipeline that the authors chose once and held constant across all systems.

free parameters (4)
  • top-k neighbor sampling = 10
    Fixed to k=10 for all models and baselines; controls subgraph density and therefore both redundancy and load imbalance.
  • batch size = 2000
    Fixed to 2000; directly affects within-batch repetition rate and cache pressure.
  • cache budget ratio α = variable (0.1–1.0)
    User-chosen fraction of GPU memory allocated to SlimCache; experiments sweep 0–1 but default comparisons use values that fit the A100 40 GB.
  • number of sampling threads = 8
    Set to 8 following baseline defaults; affinity matrix size and binding quality depend on it.
axioms (4)
  • domain assumption Sampled temporal subgraphs exhibit substantial within-batch ID repetition and cross-batch node/edge overlap that can be exploited by compression+caching.
    Quantified in Table 1 via pre-sampling; the entire SlimCache design rests on this empirical regularity holding for the target graphs.
  • domain assumption Dynamic-graph neighborhoods are small-degree (average degree ≤ 7.56 under k=10), so atomic contention in edge-parallel aggregation remains negligible.
    Stated in Section 4.3 and used to justify the COO-based AGG design over CSR balancing.
  • ad hoc to paper A single pre-sampling pass with hop=0 (roots only) yields affinity and hot-ID statistics that remain valid for the whole epoch.
    Explicit design choice in Section 4.4 and Algorithm 2; if the statistics drift, both SlimCache hit rates and topology-aware binding degrade.
  • domain assumption Training is performed on a single machine with one GPU under a main-memory (not disk) regime.
    Stated in Section 4.1; excludes multi-GPU and out-of-core settings where the relative gains may change.
invented entities (3)
  • SlimCache (joint within-batch compression + greedy cross-batch node/edge cache) independent evidence
    purpose: Reduce host–device traffic under limited GPU memory by exploiting heterogeneous node/edge redundancy.
    New system component; independent evidence is the measured I/O reduction in Figure 8 and the open-source implementation.
  • Thread-efficient AGG (COO edge-parallel) and ESM (CSR thread-loop) operators independent evidence
    purpose: Raise warp utilization and cache locality on sparse temporal subgraphs.
    Custom CUDA kernels specialized for TGNN sparsity; evidence is the kernel counters in Table 5 and operator speedups in Figure 9.
  • Topology-aware sampler with Blossom affinity binding independent evidence
    purpose: Map high-overlap sampling threads onto shared CPU cache domains.
    New scheduling policy; evidence is the L2/L3 hit-rate improvements in Figure 10.

pith-pipeline@v1.1.0-grok45 · 28844 in / 3064 out tokens · 26723 ms · 2026-07-11T08:53:00.845896+00:00 · methodology

0 comments
read the original abstract

Temporal Graph Neural Networks (TGNNs) are widely used for learning from dynamic graphs in applications such as recommendation, social network analysis, and traffic forecasting. However, scaling TGNN training to large dynamic graphs remains challenging due to three intertwined bottlenecks: memory I/O, irregular computation, and temporal neighbor sampling. Existing systems often optimize these stages in isolation, leaving substantial performance headroom on the table. We present FAST, a holistic framework that accelerates end-to-end TGNN training by jointly optimizing sampling, memory I/O, and computation. FAST introduces SlimCache, which exploits within-batch compression and cross-batch caching to reduce host-device data movement under limited GPU memory budgets. It further designs thread-efficient graph operators tailored to sparse temporal subgraphs, improving GPU cache locality and reducing the latency of aggregation and edge softmax. In addition, FAST employs a topology-aware sampling strategy that improves CPU cache locality and accelerates temporal neighbor sampling. Extensive experiments on real-world large dynamic graphs show that FAST achieves an average of 2.1x (up to 4.7x) speedup over state-of-the-art systems without sacrificing model accuracy.

Figures

Figures reproduced from arXiv: 2607.05095 by Hao Chen, Kai Sheng, Lei Liu, Qingrui Zhu, Xin He, Yushu Cai.

Figure 1
Figure 1. Figure 1: The execution time breakdown analysis on WIK [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: (a) Forward pass time breakdown on WIKITALK. [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Distributions of root node degree. unbalance rate is: UR(𝑑) = US(𝑑) USworst = 𝑅𝑑 (𝐷max − 𝑑) 𝐷max − 1 . (2) A higher UR(𝑑) indicates more severe load imbalance. We report the results in [PITH_FULL_IMAGE:figures/full_fig_p003_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Overall architecture of FAST. (a) slimming process. 1 2 2 3 3 3 3 1 1 2 3 4 1 2 3 [ 0, 1, 1, 2, 2, 2, 2 ] 1 2 3 4 [ 0, 0, 1, 2, 3 ] (b)Avg. cache process. 0 1 2 3 4 1 5 2 6 10 2 0 7 8 9 ... Iter. Node Feature 0 1 2 3 4 1 1 2 3 6 2 1 2 3 7 ... Iter. Edge Feature (c) Greedy cache process. 0 1 2 3 4 1 5 2 6 10 2 0 7 8 9 ... Iter. Node Feature 0 1 2 3 4 1 1 2 3 6 2 1 2 3 7 ... Iter. Edge Feature Node Feat. Nod… view at source ↗
Figure 5
Figure 5. Figure 5: Process of slimming and cache [PITH_FULL_IMAGE:figures/full_fig_p004_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: (a) Invalid thread in shuffler.(b) Reduction using [PITH_FULL_IMAGE:figures/full_fig_p005_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Topology-aware binding and affinity matrix. [PITH_FULL_IMAGE:figures/full_fig_p006_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: The time spent on the memory IO comparison be [PITH_FULL_IMAGE:figures/full_fig_p008_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: The time spend on graph operator in TGAT. (a) TGL, [PITH_FULL_IMAGE:figures/full_fig_p008_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: (a) The sampling time of TGL and FAST.(b) The [PITH_FULL_IMAGE:figures/full_fig_p009_10.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

37 extracted references · 9 canonical work pages

  1. [1]

    Chaoyi Chen, Dechao Gao, Yanfeng Zhang, Qiange Wang, Zhenbo Fu, Xuecang Zhang, Junhua Zhu, Yu Gu, and Ge Yu. 2023. NeutronStream: A Dynamic GNN Training Framework with Sliding Window for Graph Streams.Proc. VLDB Endow. 17, 3 (Nov. 2023), 455–468. doi:10.14778/3632093.3632108

  2. [2]

    Zhaodong Chen, Mingyu Yan, Maohua Zhu, Lei Deng, Guoqi Li, Shuangchen Li, and Yuan Xie. 2020. fuseGNN: accelerating graph convolutional neural network training on GPGPU. InProceedings of the 39th International Conference on Computer-Aided Design(Virtual Event, USA)(ICCAD ’20). Association for Computing Machinery, New York, NY, USA, Article 60, 9 pages. do...

  3. [3]

    Gangda Deng, Hongkuan Zhou, Hanqing Zeng, Yinglong Xia, Christopher Leung, Jianbo Li, Rajgopal Kannan, and Viktor Prasanna. 2024. TASER: Temporal Adap- tive Sampling for Fast and Accurate Dynamic Graph Representation Learning. In2024 IEEE International Parallel and Distributed Processing Symposium (IPDPS). 926–937. doi:10.1109/IPDPS57955.2024.00087

  4. [4]

    Shihong Gao, Yiming Li, Yanyan Shen, Yingxia Shao, and Lei Chen. 2024. ETC: Efficient Training of Temporal Graph Neural Networks over Large-Scale Dynamic Graphs.Proc. VLDB Endow.17, 5 (Jan. 2024), 1060–1072. doi:10.14778/3641204. 3641215

  5. [5]

    Shihong Gao, Yiming Li, Xin Zhang, Yanyan Shen, Yingxia Shao, and Lei Chen

  6. [6]

    ACM Manag

    SIMPLE: Efficient Temporal Graph Neural Network Training at Scale with Dynamic Data Placement.Proc. ACM Manag. Data2, 3, Article 174 (May 2024), 25 pages. doi:10.1145/3654977

  7. [7]

    Yidong Gong and Pradeep Kumar. 2024. GNNOne: A Unified System Op- timizations for GNN Kernels. InProceedings of the 33rd International Sym- posium on High-Performance Parallel and Distributed Computing(Pisa, Italy) (HPDC ’24). Association for Computing Machinery, New York, NY, USA, 15–27. doi:10.1145/3625549.3658655

  8. [8]

    Rui Guo, Zezhong Ding, Xike Xie, and Jianliang Xu. 2025. SWIFT: Enabling Large-Scale Temporal Graph Learning on a Single Machine.Proc. ACM Manag. Data3, 4, Article 266 (Sept. 2025), 27 pages. doi:10.1145/3749184

  9. [9]

    Abhinav Jangda, Sandeep Polisetty, Arjun Guha, and Marco Serafini. 2021. Accel- erating graph sampling for graph machine learning using GPUs. InProceedings of the Sixteenth European Conference on Computer Systems(Online Event, United Kingdom)(EuroSys ’21). Association for Computing Machinery, New York, NY, USA, 311–326. doi:10.1145/3447786.3456244

  10. [10]

    Jin, Lingbo Liu, Fuxian Li, and Jincai Huang

    G. Jin, Lingbo Liu, Fuxian Li, and Jincai Huang. 2023. Spatio-Temporal Graph Neu- ral Point Process for Traffic Congestion Event Prediction.AAAIabs/2311.08635, 14268–14276

  11. [11]

    Dániel Kondor, Márton Pósfai, István Csabai, and Gábor Vattay. 2014. Do the Rich Get Richer? An Empirical Analysis of the Bitcoin Transaction Network. PLOS ONE9, 2 (Feb. 2014), 1–10. doi:10.1371/journal.pone.0086197

  12. [12]

    Srijan Kumar, Xikun Zhang, and Jure Leskovec. 2019. Predicting dynamic em- bedding trajectory in temporal interaction networks. InProceedings of the 25th ACM SIGKDD international conference on knowledge discovery & data mining. 1269–1278

  13. [13]

    Kalev Leetaru and Philip A Schrodt. 2013. Gdelt: Global data on events, location, and tone, 1979–2012. InISA annual convention, Vol. 2. Citeseer, 1–49

  14. [14]

    Yiming Li, Yanyan Shen, Lei Chen, and Mingxuan Yuan. 2023. Orca: Scalable Temporal Graph Neural Network Training with Theoretical Guarantees.Proc. ACM Manag. Data1, 1, Article 52 (May 2023), 27 pages. doi:10.1145/3588737

  15. [15]

    Yuwen Liu, Lianyong Qi, Weiming Liu, Xiaolong Xu, Xuyun Zhang, and Wanchun Dou. 2024. GraphSAGE-based POI Recommendation via Continuous-Time Mod- eling. InCompanion Proceedings of the ACM Web Conference 2024(Singapore, Singapore)(WWW ’24). Association for Computing Machinery, New York, NY, USA, 585–588. doi:10.1145/3589335.3651515

  16. [16]

    Karavanic

    Konstantin Macarenco, Kristina Frye, Benjamin Hamlin, and Karen L. Karavanic

  17. [17]

    In45th International Conference on Parallel Processing Workshops, ICPP Workshops 2016, Philadelphia, PA, USA, August 16-19,

    The Effects of System Management Interrupts on Multithreaded, Hyper- threaded, and MPI Applications. In45th International Conference on Parallel Processing Workshops, ICPP Workshops 2016, Philadelphia, PA, USA, August 16-19,

  18. [18]

    doi:10.1109/ICPPW.2016.55

    IEEE Computer Society, 338–345. doi:10.1109/ICPPW.2016.55

  19. [19]

    Maxim Milakov and Natalia Gimelshein. 2018. Online Normalizer Calculation for Softmax. (2018). arXiv:1805.02867

  20. [20]

    Ashwin Paranjape, Austin R Benson, and Jure Leskovec. 2017. Motifs in temporal networks. InProceedings of the tenth ACM international conference on web search and data mining. 601–610

  21. [21]

    Emanuele Rossi, Ben Chamberlain, Fabrizio Frasca, Davide Eynard, Federico Monti, and Michael Bronstein. 2020. Temporal Graph Networks for Deep Learn- ing on Dynamic Graphs. InProceedings of the ICML 2020 Workshop on Graph Representation Learning

  22. [22]

    Ryan Rossi and Nesreen Ahmed. 2015. The network data repository with inter- active graph analytics and visualization. InProceedings of the AAAI conference on artificial intelligence, Vol. 29

  23. [23]

    Aravind Sankar, Yanhong Wu, Liang Gou, Wei Zhang, and Hao Yang. 2020. DySAT: Deep Neural Representation Learning on Dynamic Graphs via Self- Attention Networks. InProceedings of the 13th International Conference on Web Search and Data Mining(Houston, TX, USA)(WSDM ’20). Association for Com- puting Machinery, New York, NY, USA, 519–527. doi:10.1145/3336191.3371845

  24. [24]

    Guangming Sheng, Junwei Su, Chao Huang, and Chuan Wu. 2024. MSPipe: Efficient Temporal GNN Training via Staleness-Aware Pipeline. InProceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (Barcelona, Spain)(KDD ’24). Association for Computing Machinery, New York, NY, USA, 2651–2662. doi:10.1145/3637528.3671844

  25. [25]

    Amy Shoemaker and Sagar Vare. 2016. Edmonds’ blossom algorithm.CME18 (2016)

  26. [26]

    Minjie Yu Wang. 2019. Deep graph library: Towards efficient and scalable deep learning on graphs. InICLR workshop on representation learning on graphs and manifolds. doi:https://doi.org/10.48550/arXiv.1909.01315

  27. [27]

    Yufeng Wang and Charith Mendis. 2023. TGOpt: Redundancy-Aware Optimiza- tions for Temporal Graph Attention Networks. InProceedings of the 28th ACM SIGPLAN Annual Symposium on Principles and Practice of Parallel Programming (Montreal, QC, Canada)(PPoPP ’23). Association for Computing Machinery, New York, NY, USA, 354–368. doi:10.1145/3572848.3577490

  28. [28]

    Da Xu, Chuanwei Ruan, Evren Körpeoğlu, Sushant Kumar, and Kannan Achan

  29. [29]

    InInternational Conference on Learning Representations (ICLR)

    Inductive representation learning on temporal graphs. InInternational Conference on Learning Representations (ICLR)

  30. [30]

    Jianbang Yang, Dahai Tang, Xiaoniu Song, Lei Wang, Qiang Yin, Rong Chen, Wenyuan Yu, and Jingren Zhou. 2022. GNNLab: a factored system for sample- based GNN training over GPUs. InProceedings of the Seventeenth European Conference on Computer Systems(Rennes, France)(EuroSys ’22). Association 10 for Computing Machinery, New York, NY, USA, 417–434. doi:10.11...

  31. [31]

    Chenle Yu, Sara Royuela, and Eduardo Quiñones. 2024. Enhancing Hetero- geneous Computing Through OpenMP and GPU Graph. InProceedings of the 53rd International Conference on Parallel Processing(Gotland, Sweden)(ICPP ’24). Association for Computing Machinery, New York, NY, USA, 534–543. doi:10.1145/3673038.3673050

  32. [32]

    Hengrui Zhang, Zhongming Yu, Guohao Dai, Guyue Huang, Yufei Ding, Yuan Xie, and Yu Wang. 2022. Understanding GNN Computational Graph: A Coordinated Computation, IO, and Memory Perspective. InProceedings of Machine Learning and Systems (MLSys), Vol. 4. 467–484. doi:10.48550/arXiv.2110.09524

  33. [33]

    Mengqi Zhang, Shu Wu, Xueli Yu, Qiang Liu, and Liang Wang. 2023. Dynamic Graph Neural Networks for Sequential Recommendation.IEEE Trans. on Knowl. and Data Eng.35, 5 (May 2023), 4741–4753. doi:10.1109/TKDE.2022.3151618

  34. [34]

    Yuchen Zhong, Guangming Sheng, Tianzuo Qin, Minjie Wang, Quan Gan, and Chuan Wu. 2023. GNNFlow: A Distributed Framework for Continuous Temporal GNN Learning on Dynamic Graphs. (2023). arXiv:2311.17410 [cs.DC]

  35. [35]

    Hongkuan Zhou, Da Zheng, Israt Nisa, Vasileios Ioannidis, Xiang Song, and George Karypis. 2022. TGL: A General Framework for Temporal GNN Training onBillion-Scale Graphs.Proc. VLDB Endow.15, 8 (2022), 1572–1580

  36. [36]

    Zeyu Zhu, Peisong Wang, Qinghao Hu, Gang Li, Xiaoyao Liang, and Jian Cheng

  37. [37]

    FastGL: A GPU-Efficient Framework for Accelerating Sampling-Based GNN Training at Large Scale. InProceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 4(Hilton La Jolla Torrey Pines, La Jolla, CA, USA)(ASPLOS ’24). Association for Computing Machinery, New York, NY, USA, 94–110...