REVIEW 3 major objections 5 minor 69 references
By sharding communication one level deeper than existing shard-based overlap, this paper shows that data-dependent GPU communication can be overlapped with computation in an all-to-all pattern, delivering up to 1.6x speedup and a heuristic
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-03 17:12 UTC pith:LXHGMB5Q
load-bearing objection Finer-grain decomposition for compute–communication overlap is a solid experimental idea with real measured speedups on MI300X, but the schedule-picking heuristic is under-validated and sloppily specified. the 3 major comments →
Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper's central claim is that finer-grain decomposition of data-dependent communication, one level below shard granularity, converts peer-to-peer transfers into all-to-all transfers that keep direct-connection GPU interconnects busy. The resulting overheads, decomposition and contention inefficiencies, can be characterized using two static GEMM properties: operations per byte (OTB) and memory traffic (MT). From this characterization the paper derives four concrete schedules and a heuristic that selects among them, and offloading communication to GPU DMA engines reduces contention further. This combination is what delivers the reported speedups.
What carries the argument
FiCCO (Finer-grain Compute-Communication Overlap): communication is re-sharded by the number of GPUs, so in an eight-GPU system each transfer is one-eighth the size of a shard-level transfer. This turns communication into an all-to-all pattern and enables a design space of schedules distinguished by computation uniformity (uniform vs. heterogeneous), computation granularity (fused vs. unfused GEMM kernels), and communication shape (1D vs. 2D). The selection heuristic uses a GEMM's OTB and MT, forms their product, compares it against a machine-level threshold with a 5x multiplier, and picks among the schedules.
Load-bearing premise
The heuristic assumes that a GEMM's static arithmetic intensity (OTB) and memory traffic (MT), combined as a product and compared with a hand-chosen 5x machine-level threshold, reliably predicts which FiCCO schedule minimizes total inefficiency on unseen workloads and hardware.
What would settle it
Run FiCCO on a ring or torus topology, or with GEMM shapes outside the 16 synthetic scenarios, and measure whether the heuristic still picks the winning schedule in roughly 81% of cases; if accuracy drops well below that, the static OTB*MT signature is not portable across topologies.
If this is right
- FiCCO attains up to 1.6x speedup over serial execution across realistic GEMM and all-gather scenarios.
- The proposed heuristic picks the optimal schedule in 81% of unseen synthetic scenarios, with mispredictions losing about 14% of speedup.
- DMA-based communication offload reduces contention inefficiency versus GPU-core-driven communication in all measured cases.
- Shard-based overlap can be slower than serial execution on direct-connection topologies, while FiCCO avoids that degradation.
- The design-space framing gives frameworks and runtimes a concrete mechanism for choosing overlap schedules based on operation shapes.
Where Pith is reading between the lines
- The heuristic's 5x threshold and OTB*MT product are likely calibrated to the specific GPU and topology studied; on other hardware the demarcation may shift, so the heuristic may need recalibration rather than re-derivation.
- If DMA engines gain support for 2D copies, the emulated 2D schedules should be re-measured; real 2D DMA may close part of the remaining gap to ideal speedup.
- The same one-level-deeper decomposition idea could extend to reduce-scatter-based parallelism, such as tensor parallelism with gradients, once DMA engines support arithmetic operations, which the paper explicitly leaves out.
- The correlation between static GEMM properties and inefficiency signatures suggests that a similar OTB/MT-based selector could be applied to other overlap schemes, not just the four schedules studied here.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes FiCCO, a finer-granularity compute-communication overlap scheme that decomposes communication one level deeper than shard-based overlap (e.g., transfer sizes one-eighth of shard-based in an 8-GPU system). This decomposition is argued to unlock all-to-all communication, better utilize direct-connection topologies such as AMD MI300X, and enable a richer space of execution schedules. The paper characterizes two inefficiencies—decomposition inefficiency loss (DIL) and contention inefficiency loss (CIL)—and maps them to static GEMM features (OTB and MT). It then defines four FiCCO schedules, proposes a heuristic based on OTB×MT and a 5× machine-dependent threshold to select among them, and offloads communication to GPU DMA engines. Experiments on 16 real GEMM scenarios report up to 1.6× speedup over serial execution, and an additional 16 synthetic scenarios are used to claim 81% heuristic accuracy.
Significance. If the heuristic generalizes, the paper offers a practical, software-only method for improving dependent compute-communication overlap on full-mesh GPU systems, a setting where existing shard-based overlap degrades. The strengths are the concrete measurements on a real MI300X system, the use of DMA offload to reduce contention, and the attempt to ground schedule selection in static operator properties. The core speedup result is plausible and independent of the heuristic. However, the heuristic's central claim of 81% accuracy on unseen scenarios rests on a small, author-generated validation set, a hand-picked threshold, and a unit-inconsistent decision variable. The 2D communication schedule is also emulated rather than measured. These issues limit the current evidence for transferability, which is the main practical contribution.
major comments (3)
- [V-C and VI-C] The heuristic's decision boundary is not well-founded. Section V-C defines combined OTB×MT for the GEMM and compares it to 'machine-level' OTB×MT defined as 'peak compute FLOPs' (FLOPs/s). This is a unit mismatch: OTB (FLOPs/byte) times MT (bytes) yields FLOPs, so the comparison implicitly imposes a 1-second timescale. The 5× threshold is hand-chosen and appears tuned to the 15 real scenarios in Table I, on which the heuristic then achieves 100% accuracy. The only out-of-sample evaluation is 16 synthetic scenarios with no cross-validation, no sensitivity analysis for the 5× value, and no confidence intervals. With one free parameter and 16 test points, 13/16 correct is weak evidence for the '81% of unseen scenarios' claim. Please provide a principled derivation of the threshold, report accuracy as a function of the threshold, or validate on independently collected workloads/hardware.
- [VI-B and Figure 12b] The 2D schedule is emulated using 1D memory copies of the same size because '2D memory copies with DMAs are not supported today.' The paper still reports 'with emulated 2D schedules we attain as high as 1.7× speedup' and includes uniform-fused-2D in the headline results and heuristic evaluation. A 1D copy does not reproduce the buffer layout, gather/scatter behavior, or link-level traffic pattern of a true 2D communication shape. This means the 2D arm of the design space is not actually evaluated, and the 1.7× number is a best-case estimate, not a measured result. Please either implement true 2D transfers, use a simulator validated against the 1D results, or clearly label these as estimates throughout the abstract, Section V, and Section VI.
- [IV-B and VI] The paper reports speedups and DIL/CIL values as point averages of 5 runs (after 10 warmups) but gives no error bars, standard deviations, or statistical significance tests. This matters because the heuristic is selecting among schedules whose speedups can be close (e.g., in Figure 12b the difference between schedules is often small relative to the reported 6% operator variation). Without variance information, the reader cannot tell whether the 81% accuracy or the 1.6× speedup is robust, or whether the heuristic's choices are within measurement noise. Please report error bars or variance for the headline numbers and for the per-schedule speedups in Figure 12b.
minor comments (5)
- [VI-C] The 16 synthetic scenarios are described only as 'wide ranging OTB and MT combinations.' Please specify how they were generated (parameter ranges, GEMM shapes, whether they include the M<K case), and list them so the out-of-sample claim is reproducible.
- [IV-C] The phrase 'op-to-byte' is used inconsistently; consider using 'OTB' consistently after first definition. Also, in Figure 7 the data labels are small and hard to read; enlarge them or report the values in a table.
- [VI-A] The text says 'we observe 7× communication slowdown' for shard-based overlap on MI300X, but Figure 13 shows speedup, not communication slowdown. Clarify where the 7× comes from or add a direct measurement.
- [IV-B.2] The omitted scenarios (e.g., tensor parallelism with reduce-scatter) are excluded because DMA engines do not support math today. This is a reasonable limitation, but it should be stated earlier and in the abstract or conclusion, since it bounds the applicability of FiCCO.
- [VII] The related work discussion is brief for a design-space paper. In particular, the comparison to Triton-Distributed is reported as out-of-memory; consider adding a small-scale comparison or at least a qualitative discussion of how FiCCO differs in kernel-authoring burden.
Circularity Check
No significant circularity: FiCCO speedups are measured and the heuristic, though in-sample validated, is not defined by the predicted schedule labels.
full rationale
The paper's central speedup claim is an empirical result: FiCCO is implemented and measured on MI300X against serial execution, shard-overlap, RCCL-based FiCCO, and an ideal roofline, with the up-to-1.6x speedup coming from measured execution times (Sections VI-B and VI-D). The DIL/CIL characterization is also measured, and the heuristic in Section V-C selects schedules from static GEMM features (M vs K, OTB, MT) using a decision rule; the rule is not defined in terms of the measured winning schedule, so no predicted quantity is equal to an input by construction. The 5x machine-level threshold is hand-picked rather than derived, and evaluating the heuristic on the same 15 real scenarios is in-sample validation—a generalization limitation, not a definitional circularity. The 16-scenario synthetic test provides a separate, though small, out-of-sample check. Self-citations such as ConCCL [1], T3 [42], and Tale of Two Cs [41] provide background on DMA offload and motivation, but the paper implements DMA via hipMemcpyDtoDAsync and measures its effect directly, so the citations are not load-bearing. No step in the derivation chain reduces a predicted result to its own inputs.
Axiom & Free-Parameter Ledger
free parameters (2)
- 5x threshold in FiCCO heuristic =
5
- OTB*MT product as combined static metric
axioms (4)
- domain assumption 8-GPU MI300X full-mesh topology is representative of modern GPU systems
- domain assumption DMA-based async memory copies can overlap with GEMM kernels and reduce interference
- domain assumption Static GEMM dimensions (M,N,K) determine OTB and MT, which predict DIL and CIL
- domain assumption The 16 synthetic scenarios are representative of unseen real deployments
read the original abstract
Modern ML workloads demand distributing training and inference across multiple GPUs. However, these parallelization techniques often suffer from exposed critical-path communication, leaving a potential 1.7x speedup on the table through compute-communication overlap. Prior overlapping methods harness the fact that ML model state and inputs are already sharded into the number of GPUs, and overlap the compute and communication at shard granularity. However, such coarse-grained overlap suffers from limited network topology support, and suboptimal dataflows. In this work, we instead make a case for finer-grain compute-communication overlap which we term FiCCO. FiCCO operates one level deeper than traditional sharding, and unlocks overlap for a wider set of network topologies and enables finer-grain dataflow. We show that FiCCO opens up a wider design space of execution schedules than possible at shard-level alone. To walk the design space of schedules, we study and characterize the performance inefficiencies on doing overlap and overlay the schedules with the associated inefficiency signatures. Our characterization reveals decomposition and contention based slowdowns to be the major performance limiters, and we correlate the slowdown factors with the static compute/communication operator sizes. This helps us design heuristics (that frameworks and runtimes can harness) to select bespoke FiCCO schedules based on the nature of underlying ML operations. Finally, to further minimize contention inefficiencies inherent with operation overlap, we offload communication to GPU DMA engines. We evaluate several scenarios from realistic ML deployments and demonstrate that our proposed heuristics driven bespoke schedules deliver up to 1.6x speedup. Further, our heuristics provide accurate guidance to pick the optimal schedule in 81% of unseen scenarios.
Figures
Reference graph
Works this paper leans on
-
[1]
Conccl: Optimizing ML concurrent computation and communication with GPU DMA engines,
A. Agrawal, S. Aga, S. Pati, and M. Islam, “Conccl: Optimizing ML concurrent computation and communication with GPU DMA engines,” inIEEE International Symposium on Performance Analysis of Systems and Software, ISPASS 2025, Ghent, Belgium, May 11-13, 2025. IEEE, 2025, pp. 1–11. [Online]. Available: https: //doi.org/10.1109/ISPASS64960.2025.00018
arXiv 2025
-
[2]
[Distributed GEMM: A novel CUTLASS-based implementation of Tensor Parallelism for NVLink-enabled systems,
Ali Hassani, Michael Isaev, Nic McDonald, Jie Ren, Vijay Thakkar, Haicheng Wu, and Humphrey Shi, “[Distributed GEMM: A novel CUTLASS-based implementation of Tensor Parallelism for NVLink-enabled systems,” https://blog.shi-labs. com/distributed-gemm-88be6a481e2b, December 2024
2024
-
[3]
(2023) Amd instinct™ mi300x accelerators
AMD. (2023) Amd instinct™ mi300x accelerators. [On- line]. Available: https://www.amd.com/en/products/accelerators/instinct/ mi300/mi300x.html
2023
-
[4]
HIP: C++ Heterogeneous-Compute Interface for Portability,
AMD, “HIP: C++ Heterogeneous-Compute Interface for Portability,” https://github.com/ROCm/HIP, 2024
2024
-
[5]
ROCm Communication Collectives Library (RCCL),
——, “ROCm Communication Collectives Library (RCCL),” https: //github.com/ROCm/rccl, 2024
2024
-
[6]
ROCm: HIPStream,
——, “ROCm: HIPStream,” https://rocm.docs.amd.com/projects/HIP/ en/latest/reference/hip runtime api/modules/stream management.html, 2024
2024
-
[7]
ROCm/rocBLAS: Next generation BLAS implementation for ROCm platform,
——, “ROCm/rocBLAS: Next generation BLAS implementation for ROCm platform,” https://github.com/ROCm/rocBLAS, 2024
2024
-
[8]
(2025) Hip graphs
AMD. (2025) Hip graphs. [Online]. Avail- able: https://rocm.docs.amd.com/projects/HIP/en/docs-develop/how-to/ hip runtime api/hipgraph.html
2025
-
[9]
(2025) hipblaslt
——. (2025) hipblaslt. [Online]. Available: https://github.com/ROCm/ rocm-libraries
2025
-
[10]
FLUX: fast software-based communication overlap on gpus through kernel fusion,
L. Chang, W. Bao, Q. Hou, C. Jiang, N. Zheng, Y . Zhong, X. Zhang, Z. Song, Z. Jiang, H. Lin, X. Jin, and X. Liu, “FLUX: fast software-based communication overlap on gpus through kernel fusion,”CoRR, vol. abs/2406.06858, 2024. [Online]. Available: https://doi.org/10.48550/arXiv.2406.06858
-
[11]
C. Chen, X. Li, Q. Zhu, J. Duan, P. Sun, X. Zhang, and C. Yang, “Centauri: Enabling efficient scheduling for communication- computation overlap in large model training via communication partitioning,” inProceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS ’24). New York, NY ,...
arXiv 2024
-
[12]
Evaluating large language models trained on code,
M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. de Oliveira Pinto, J. Kaplan, H. Edwards, Y . Burda, N. Joseph, G. Brockman, A. Ray, R. Puri, G. Krueger, M. Petrov, H. Khlaaf, G. Sastry, P. Mishkin, B. Chan, S. Gray, N. Ryder, M. Pavlov, A. Power, L. Kaiser, M. Bavarian, C. Winter, P. Tillet, F. P. Such, D. Cummings, M. Plappert, F. Chantzis, E. Barnes, A. Her...
Pith/arXiv arXiv 2021
-
[13]
Revisiting scaling laws for language models: The role of data quality and training strategies,
Z. Chen, S. Wang, T. Xiao, Y . Wang, S. Chen, X. Cai, J. He, and J. Wang, “Revisiting scaling laws for language models: The role of data quality and training strategies,” in Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar, Eds. Vienna, Austr...
2025
-
[14]
Concerto: Automatic communication optimization and scheduling for large-scale deep learning,
S. Cheng, S. Lin, L. Diao, H. Wu, S. Wang, C. Si, Z. Liu, X. Zhao, J. Du, W. Lin, and Y . You, “Concerto: Automatic communication optimization and scheduling for large-scale deep learning,” inProceedings of the 30th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 1, ASPLOS 2025, Rotterdam, The ...
arXiv 2025
-
[15]
Scaling llama 3 training with efficient parallelism strategies,
W. Chu, X. Xie, J. Yu, J. Wang, A. Phanishayee, C. Tang, Y . Hao, J. Huang, M. Ozdal, J. Wang, V . Goswami, N. Goyal, A. Kadian, A. Gu, C. Cai, F. Tian, X. Wang, M. Si, P. Balaji, C.-H. Chu, and J. Park, “Scaling llama 3 training with efficient parallelism strategies,” inProceedings of the 52nd Annual International Symposium on Computer Architecture, ser....
arXiv 2025
-
[16]
Transformations to parallel codes for communication-computation overlap,
A. Danalis, K.-Y . Kim, L. Pollock, and M. Swany, “Transformations to parallel codes for communication-computation overlap,” inSC ’05: Proceedings of the 2005 ACM/IEEE Conference on Supercomputing, 2005, pp. 58–58
2005
-
[17]
Mpi-aware compiler optimizations for improving communication-computation overlap,
A. Danalis, L. Pollock, M. Swany, and J. Cavazos, “Mpi-aware compiler optimizations for improving communication-computation overlap,” inProceedings of the 23rd International Conference on Supercomputing, ser. ICS ’09. New York, NY , USA: Association for Computing Machinery, 2009, p. 316–325. [Online]. Available: https://doi.org/10.1145/1542275.1542321
arXiv 2009
-
[18]
Flashattention: Fast and memory-efficient exact attention with io-awareness,
T. Dao, D. Y . Fu, S. Ermon, A. Rudra, and C. R ´e, “Flashattention: Fast and memory-efficient exact attention with io-awareness,” inAdvances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022, New Orleans, LA, USA, November 28 - December 9, 2022, S. Koyejo, S. Mohamed, A. Agarwal, D....
2022
-
[19]
Deepseek-v3 technical report,
DeepSeek-AI, A. Liu, B. Feng, B. Xue, B. Wang, B. Wu, C. Lu, C. Zhao, C. Deng, C. Zhang, C. Ruan, D. Dai, D. Guo, D. Yang, D. Chen, D. Ji, E. Li, F. Lin, F. Dai, F. Luo, G. Hao, G. Chen, G. Li, H. Zhang, H. Bao, H. Xu, H. Wang, H. Zhang, H. Ding, H. Xin, H. Gao, H. Li, H. Qu, J. L. Cai, J. Liang, J. Guo, J. Ni, J. Li, J. Wang, J. Chen, J. Chen, J. Yuan, J...
-
[20]
The llama 3 herd of models,
A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al- Dahle, A. Letman, A. Mathur, A. Schelten, A. Vaughan, A. Yang, A. Fan, A. Goyal, A. Hartshorn, A. Yang, A. Mitra, A. Sravankumar, A. Korenev, A. Hinsvark, A. Rao, A. Zhang, A. Rodriguez, A. Gregerson, A. Spataru, B. Roziere, B. Biron, B. Tang, B. Chern, C. Caucheteux, C. Nayak, C. Bi, C. Ma...
-
[21]
Compiler-assisted overlapping of communication and computation in mpi applications,
J. Guo, Q. Yi, J. Meng, J. Zhang, and P. Balaji, “Compiler-assisted overlapping of communication and computation in mpi applications,” in 2016 IEEE International Conference on Cluster Computing (CLUSTER), 2016, pp. 60–69
2016
-
[22]
Bandwidth characterization of deepspeed on distributed large language model training,
B. Hanindhito, B. Patel, and L. K. John, “Bandwidth characterization of deepspeed on distributed large language model training,” in IEEE International Symposium on Performance Analysis of Systems and Software, ISPASS 2024, Indianapolis, IN, USA, May 5- 7, 2024. IEEE, 2024, pp. 241–256. [Online]. Available: https: //doi.org/10.1109/ISPASS61541.2024.00031
arXiv 2024
-
[23]
Efficient and adaptable overlapping for computation and communication via signaling and reordering,
K. Hong, X. Li, M. Liu, Q. Mao, T. Wu, Z. Huang, L. Chen, Z. Wang, Y . Zhang, Z. Zhu, G. Dai, and Y . Wang, “Efficient and adaptable overlapping for computation and communication via signaling and reordering,” 2025
2025
-
[24]
[Distributed w/ TorchTitan] Introducing Async Tensor Parallelism in PyTorch,
Horace He, Less Wright, Luca Wehrstedt, Tianyu Liu, Wanchao Liang, “[Distributed w/ TorchTitan] Introducing Async Tensor Parallelism in PyTorch,” https://discuss.pytorch.org/t/ distributed-w-torchtitan-introducing-async-tensor-parallelism-in-pytorch/ 209487, September 2024
2024
-
[25]
Demystifying nccl: An in-depth analysis of gpu communication protocols and algorithms,
Z. Hu, S. Shen, T. Bonato, S. Jeaugey, C. Alexander, E. Spada, J. Dinan, J. Hammond, and T. Hoefler, “Demystifying nccl: An in-depth analysis of gpu communication protocols and algorithms,”
-
[26]
Gpipe: Efficient training of giant neural networks using pipeline parallelism,
Y . Huang, Y . Cheng, A. Bapna, O. Firat, D. Chen, M. X. Chen, H. Lee, J. Ngiam, Q. V . Le, Y . Wu, and Z. Chen, “Gpipe: Efficient training of giant neural networks using pipeline parallelism,” inAdvances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouv...
2019
-
[27]
A loop transformation algorithm for communication overlapping,
K. Ishizaki, H. Komatsu, and T. Nakatani, “A loop transformation algorithm for communication overlapping,”Int. J. Parallel Program., vol. 28, no. 2, p. 135–154, Apr. 2000. [Online]. Available: https://doi.org/10.1023/A:1007554715418
-
[28]
A. Jangda, J. Huang, G. Liu, A. H. N. Sabet, S. Maleki, Y . Miao, M. Musuvathi, T. Mytkowicz, and O. Saarikivi, “Breaking the computation and communication abstraction barrier in distributed machine learning workloads,” inProceedings of the 27th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, ser. ASP...
arXiv 2022
-
[29]
Available: https://arxiv.org/abs/2507.04786
[Online]. Available: https://arxiv.org/abs/2507.04786
-
[30]
A survey on large language models for code generation,
J. Jiang, F. Wang, J. Shen, S. Kim, and S. Kim, “A survey on large language models for code generation,”CoRR, vol. abs/2406.00515,
-
[31]
Reducing activation recomputation in large transformer models,
V . A. Korthikanti, J. Casper, S. Lym, L. McAfee, M. Andersch, M. Shoeybi, and B. Catanzaro, “Reducing activation recomputation in large transformer models,” inProceedings of the Sixth Conference on Machine Learning and Systems, MLSys 2023, Miami, FL, USA, June 4-8, 2023, D. Song, M. Carbin, and T. Chen, Eds. mlsys.org, 2023. [On- line]. Available: https:...
arXiv 2023
-
[32]
Lit silicon: A case where thermal imbalance couples concurrent execution in multiple gpus,
M. Kurzynski, S. Aga, and D. Wu, “Lit silicon: A case where thermal imbalance couples concurrent execution in multiple gpus,”
-
[33]
A. Q. Jiang, A. Sablayrolles, A. Roux, A. Mensch, B. Savary, C. Bamford, D. S. Chaplot, D. de Las Casas, E. B. Hanna, F. Bressand, G. Lengyel, G. Bour, G. Lample, L. R. Lavaud, L. Saulnier, M. Lachaux, P. Stock, S. Subramanian, S. Yang, S. Antoniak, T. L. Scao, T. Gervet, T. Lavril, T. Wang, T. Lacroix, and W. E. Sayed, “Mixtral of experts,”CoRR, vol. abs...
-
[34]
Pytorch 11 distributed: Experiences on accelerating data parallel training,
S. Li, Y . Zhao, R. Varma, O. Salpekar, P. Noordhuis, T. Li, A. Paszke, J. Smith, B. Vaughan, P. Damania, and S. Chintala, “Pytorch 11 distributed: Experiences on accelerating data parallel training,”Proc. VLDB Endow., vol. 13, no. 12, pp. 3005–3018, 2020. [Online]. Available: http://www.vldb.org/pvldb/vol13/p3005-li.pdf
2020
-
[35]
Available: https://doi.org/10.48550/arXiv.2406.00515
[Online]. Available: https://doi.org/10.48550/arXiv.2406.00515
-
[36]
Ring attention with blockwise transformers for near-infinite context,
——, “Ring attention with blockwise transformers for near-infinite context,” 2023. [Online]. Available: https://arxiv.org/abs/2310.01889
Pith/arXiv arXiv 2023
-
[37]
MLPerf Inference Results v5.0,
MLCommons, “MLPerf Inference Results v5.0,” https://github.com/ mlcommons/inference results v5.0, 2025, accessed: 2025-11-16
2025
-
[38]
Available: https://arxiv.org/abs/2511.09861
[Online]. Available: https://arxiv.org/abs/2511.09861
-
[39]
Gshard: Scaling giant models with conditional computation and automatic sharding,
D. Lepikhin, H. Lee, Y . Xu, D. Chen, O. Firat, Y . Huang, M. Krikun, N. Shazeer, and Z. Chen, “Gshard: Scaling giant models with conditional computation and automatic sharding,” in9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net, 2021. [Online]. Available: https://openreview.net/fo...
2021
-
[40]
Pytorch: An imperative style, high-performance deep learning library,
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. K ¨opf, E. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala, “Pytorch: An imperative style, high-performance deep learning library,” 2019. [Online]. Available: https://arxiv.org...
Pith/arXiv arXiv 2019
-
[42]
T3: Transparent tracking & triggering for fine-grained overlap of compute & collectives,
——, “T3: Transparent tracking & triggering for fine-grained overlap of compute & collectives,” inProceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2, ser. ASPLOS ’24. New York, NY , USA: Association for Computing Machinery, 2024, p. 1146–1164. [Online]. Available: https://...
arXiv 2024
-
[43]
Exact dependence analysis for increased communication overlap,
S. Pellegrini, T. Hoefler, and T. Fahringer, “Exact dependence analysis for increased communication overlap,” inRecent Advances in the Message Passing Interface, J. L. Tr ¨aff, S. Benkner, and J. J. Dongarra, Eds. Berlin, Heidelberg: Springer Berlin Heidelberg, 2012, pp. 89–99
2012
-
[44]
Automatic gen- eration of software pipelines for heterogeneous parallel systems,
J. A. Pienaar, S. Chakradhar, and A. Raghunathan, “Automatic gen- eration of software pipelines for heterogeneous parallel systems,” in Proceedings of the International Conference on High Performance Computing, Networking, Storage and Analysis, ser. SC ’12. Washington, DC, USA: IEEE Computer Society Press, 2012
2012
-
[45]
Efficient large-scale language model training on gpu clusters using megatron-lm,
D. Narayanan, M. Shoeybi, J. Casper, P. LeGresley, M. Patwary, V . Korthikanti, D. Vainbrand, P. Kashinkunti, J. Bernauer, B. Catanzaro, A. Phanishayee, and M. Zaharia, “Efficient large-scale language model training on gpu clusters using megatron-lm,” inProceedings of the International Conference for High Performance Computing, Networking, Storage and Ana...
arXiv 2021
-
[46]
Stream-k: Work-centric parallel decomposition for dense matrix- matrix multiplication on the GPU,
M. Osama, D. Merrill, C. Cecka, M. Garland, and J. D. Owens, “Stream-k: Work-centric parallel decomposition for dense matrix- matrix multiplication on the GPU,” inProceedings of the 28th ACM SIGPLAN Annual Symposium on Principles and Practice of Parallel Programming, PPoPP 2023, Montreal, QC, Canada, 25 February 2023 - 1 March 2023, M. M. Dehnavi, M. Kulk...
arXiv 2023
-
[47]
Enabling compute-communication overlap in distributed deep learning training platforms,
S. Rashidi, M. Denton, S. Sridharan, S. Srinivasan, A. Suresh, J. Nie, and T. Krishna, “Enabling compute-communication overlap in distributed deep learning training platforms,” in48th ACM/IEEE Annual International Symposium on Computer Architecture, ISCA 2021, Virtual Event / Valencia, Spain, June 14-18, 2021. IEEE, 2021, pp. 540–553. [Online]. Available:...
arXiv 2021
-
[48]
Tale of two cs: Computation vs. communication scaling for future transformers on future hardware,
S. Pati, S. Aga, M. Islam, N. Jayasena, and M. D. Sinclair, “Tale of two cs: Computation vs. communication scaling for future transformers on future hardware,” inIEEE International Symposium on Workload Characterization, IISWC 2023, Ghent, Belgium, October 1-3, 2023. IEEE, 2023, pp. 140–153. [Online]. Available: https://doi.org/10.1109/IISWC59245.2023.00026
arXiv 2023
-
[49]
Slechta, N
B. Slechta, N. Comly, A. Eassa, J. DeLaere, and S. Raj. (2024, Aug) Nvidia nvlink and nvidia nvswitch supercharge large language model inference. NVIDIA Technical Blog. [Online]. Available: https://developer.nvidia.com/blog/ nvidia-nvlink-and-nvidia-nvswitch-supercharge-large-language-model-inference/ #nvswitch is critical for fast multi-gpu llm inference
2024
-
[50]
Triton: an intermediate language and compiler for tiled neural network computations,
P. Tillet, H. Kung, and D. D. Cox, “Triton: an intermediate language and compiler for tiled neural network computations,” in Proceedings of the 3rd ACM SIGPLAN International Workshop on Machine Learning and Programming Languages, MAPL@PLDI 2019, Phoenix, AZ, USA, June 22, 2019, T. Mattson, A. Muzahid, and A. Solar-Lezama, Eds. ACM, 2019, pp. 10–19. [Onlin...
arXiv 2019
-
[51]
Llama: Open and efficient foundation language models,
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozi `ere, N. Goyal, E. Hambro, F. Azhar, A. Rodriguez, A. Joulin, E. Grave, and G. Lample, “Llama: Open and efficient foundation language models,” 2023. [Online]. Available: https: //arxiv.org/abs/2302.13971
Pith/arXiv arXiv 2023
-
[52]
Optimizing distributed ML communication with fused computation-collective operations,
K. Punniyamurthy, K. Hamidouche, and B. M. Beckmann, “Optimizing distributed ML communication with fused computation-collective operations,” inProceedings of the International Conference for High Performance Computing, Networking, Storage, and Analysis, SC 2024, Atlanta, GA, USA, November 17-22, 2024. IEEE, 2024, p. 88. [Online]. Available: https://dl.acm...
Pith/arXiv arXiv 2024
-
[53]
Deepspeed-moe: Advancing mixture-of- experts inference and training to power next-generation AI scale,
S. Rajbhandari, C. Li, Z. Yao, M. Zhang, R. Y . Aminabadi, A. A. Awan, J. Rasley, and Y . He, “Deepspeed-moe: Advancing mixture-of- experts inference and training to power next-generation AI scale,” in International Conference on Machine Learning, ICML 2022, 17-23 July 2022, Baltimore, Maryland, USA, ser. Proceedings of Machine Learning Research, K. Chaud...
2022
-
[54]
Domino: Eliminating Communication in LLM Training via Generic Tensor Slicing and Overlapping
G. Wang, C. Zhang, Z. Shen, A. Li, and O. Ruwase, “Domino: Eliminating communication in LLM training via generic tensor slicing and overlapping,”CoRR, vol. abs/2409.15241, 2024. [Online]. Available: https://doi.org/10.48550/arXiv.2409.15241
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.2409.15241 2024
-
[55]
A hybrid tensor-expert-data parallelism approach to optimize mixture-of-experts training,
S. Singh, O. Ruwase, A. A. Awan, S. Rajbhandari, Y . He, and A. Bhatele, “A hybrid tensor-expert-data parallelism approach to optimize mixture-of-experts training,” inProceedings of the 37th International Conference on Supercomputing, ICS 2023, Orlando, FL, USA, June 21-23, 2023, K. A. Gallivan, E. Gallopoulos, D. S. Nikolopoulos, and R. Beivide, Eds. ACM...
arXiv 2023
-
[56]
Pytorch symmetricmemory: Harnessing nvlink programmability with ease,
Y . Wang, H. He, and L. Wehrstedt, “Pytorch symmetricmemory: Harnessing nvlink programmability with ease,” Feb 2025, pyTorch Developer Forum. [Online]. Available: https://dev-discuss.pytorch.org/t/ pytorch-symmetricmemory-harnessing-nvlink-programmability-with-ease/ 2798
2025
-
[57]
Petuum: A new platform for distributed machine learning on big data,
E. P. Xing, Q. Ho, W. Dai, J. K. Kim, J. Wei, S. Lee, X. Zheng, P. Xie, A. Kumar, and Y . Yu, “Petuum: A new platform for distributed machine learning on big data,”IEEE Trans. Big Data, vol. 1, no. 2, pp. 49–67, 2015. [Online]. Available: https://doi.org/10.1109/TBDATA.2015.2472014
arXiv 2015
-
[58]
Context parallelism for scalable million-token inference,
A. Yang, J. Yang, A. Ibrahim, X. Xie, B. Tang, G. Sizov, J. Reizenstein, J. Park, and J. Huang, “Context parallelism for scalable million-token inference,”CoRR, vol. abs/2411.01783, 2024. [Online]. Available: https://doi.org/10.48550/arXiv.2411.01783 12
-
[59]
Llama 2: Open foundation and fine-tuned chat models,
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y . Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, D. Bikel, L. Blecher, C. C. Ferrer, M. Chen, G. Cucurull, D. Esiobu, J. Fernandes, J. Fu, W. Fu, B. Fuller, C. Gao, V . Goswami, N. Goyal, A. Hartshorn, S. Hosseini, R. Hou, H. Inan, M. Kardas, V . Kerkez, M. Khabsa, I. Kloumann, A. Koren...
Pith/arXiv arXiv 2023
-
[60]
P. Vaithilingam, T. Zhang, and E. L. Glassman, “Expectation vs. experience: Evaluating the usability of code generation tools powered by large language models,” inCHI ’22: CHI Conference on Human Factors in Computing Systems, New Orleans, LA, USA, 29 April 2022 - 5 May 2022, Extended Abstracts, S. D. J. Barbosa, C. Lampe, C. Appert, and D. A. Shamma, Eds....
arXiv 2022
-
[61]
S. Zheng, W. Bao, Q. Hou, X. Zheng, J. Fang, C. Huang, T. Li, H. Duanmu, R. Chen, R. Xu, Y . Guo, N. Zheng, Z. Jiang, X. Di, D. Wang, J. Ye, H. Lin, L.-W. Chang, L. Lu, Y . Liang, J. Zhai, and X. Liu, “Triton-distributed: Programming overlapping kernels on distributed ai systems with the triton compiler,” 2025. [Online]. Available: https://arxiv.org/abs/2...
Pith/arXiv arXiv 2025
-
[62]
Overlap communication with dependent computation via decomposition in large deep learning models,
S. Wang, J. Wei, A. Sabne, A. Davis, B. Ilbeyi, B. Hechtman, D. Chen, K. S. Murthy, M. Maggioni, Q. Zhang, S. Kumar, T. Guo, Y . Xu, and Z. Zhou, “Overlap communication with dependent computation via decomposition in large deep learning models,” inProceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and O...
arXiv 2023
-
[66]
Comet: Fine-grained computation-communication overlapping for mixture- of-experts,
S. Zhang, N. Zheng, H. Lin, Z. Jiang, W. Bao, C. Jiang, Q. Hou, W. Cui, S. Zheng, L. Chang, Q. Chen, and X. Liu, “Comet: Fine-grained computation-communication overlapping for mixture- of-experts,”CoRR, vol. abs/2502.19811, 2025. [Online]. Available: https://doi.org/10.48550/arXiv.2502.19811
-
[67]
Pytorch FSDP: experiences on scaling fully sharded data parallel,
Y . Zhao, A. Gu, R. Varma, L. Luo, C. Huang, M. Xu, L. Wright, H. Shojanazeri, M. Ott, S. Shleifer, A. Desmaison, C. Balioglu, P. Damania, B. Nguyen, G. Chauhan, Y . Hao, A. Mathews, and S. Li, “Pytorch FSDP: experiences on scaling fully sharded data parallel,” Proc. VLDB Endow., vol. 16, no. 12, pp. 3848–3860, 2023. [Online]. Available: https://www.vldb....
2023
-
[69]
Tilelink: Generating efficient compute-communication overlapping kernels using tile-centric primitives,
S. Zheng, J. Fang, X. Zheng, Q. Hou, W. Bao, N. Zheng, Z. Jiang, D. Wang, J. Ye, H. Lin, L.-W. Chang, and X. Liu, “Tilelink: Generating efficient compute-communication overlapping kernels using tile-centric primitives,” inEighth Conference on Machine Learning and Systems,
-
[70]
Available: https://openreview.net/forum?id=ccjvBkTRRe 13
[Online]. Available: https://openreview.net/forum?id=ccjvBkTRRe 13
-
[2022]
Available: http://papers.nips.cc/paper files/paper/2022/ hash/67d57c32e20fd0a7a302cb81d36e40d5-Abstract-Conference.html
[Online]. Available: http://papers.nips.cc/paper files/paper/2022/ hash/67d57c32e20fd0a7a302cb81d36e40d5-Abstract-Conference.html
2022
-
[2023]
Available: https://doi.org/10.48550/arXiv.2310.01889
[Online]. Available: https://doi.org/10.48550/arXiv.2310.01889
-
[2024]
Available: https://arxiv.org/abs/2407.21783
[Online]. Available: https://arxiv.org/abs/2407.21783
-
[2025]
Available: https://arxiv.org/abs/2412.19437
[Online]. Available: https://arxiv.org/abs/2412.19437
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.